Property Based Testing
Write, review, and triage property-based tests — Hypothesis, fast-check, proptest, and Echidna or Medusa for Solidity invariants
- Type
- Plugin
- Repository
- trailofbits/skills
- GitHub stars
- 7.3k
- License
- CC-BY-SA-4.0
- Repo last updated
- Sep 25, 2026
- Version
- 1.2.2
- Author
- Henrik Brodin
What Property Based Testing is
Property Based Testing is a plugin published in the trailofbits/skills repository on GitHub, which has about 7.3k stars. The repository describes itself as: “Trail of Bits Claude Code skills for security research, vulnerability detection, and audit workflows”
A plugin is a package that bundles skills, slash commands, subagents, hooks, and MCP connectors so they install together. Plugins are plain files with a manifest at .claude-plugin/plugin.json, and they work in both Claude Code and Claude Cowork.
Installing Property Based Testing adds everything it ships in one step. Connectors inside a plugin still need to be connected separately, and hooks and subagents only run in Cowork and Claude Code, not in regular chat.
How to install Property Based Testing
Claude Code
- Add the repository as a plugin marketplace: claude plugin marketplace add trailofbits/skills
- Install the plugin: claude plugin install property-based-testing@<marketplace-name>, using the marketplace name from the repository's .claude-plugin/marketplace.json.
- Restart the session if the new skills or commands don't appear straight away.
Claude Cowork
- Open Customize → Plugins and choose Add marketplace.
- Enter trailofbits/skills (the owner/repo shorthand works for GitHub).
- Find Property Based Testing in the list, click Install, then connect any connectors it needs from its Connectors tab.
New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.
Inside the source file
An excerpt from plugins/property-based-testing/.claude-plugin/plugin.json, shared under the repository's CC-BY-SA-4.0 license. Read the full file on GitHub.
Write, review, and triage property-based tests — Hypothesis, fast-check, proptest, and Echidna or Medusa for Solidity invariants.
Installation
This plugin is part of the Trail of Bits Skills marketplace.
Via Marketplace (Recommended)
/plugin marketplace add trailofbits/skills
/plugin menuThen select the property-based-testing plugin to install.
Manual Installation
/plugin install trailofbits/skills/plugins/property-based-testingWhat's Included
This plugin provides a skill covering three jobs: writing property tests, reviewing existing ones for tests that assert nothing, and triaging a shrunk counterexample into a wrong property, an ambiguous spec, or a real bug. It recognises these shapes:
- Serialization pairs: encode/decode, serialize/deserialize, toJSON/fromJSON
- Parsers: URL parsing, config parsing, protocol parsing
- Normalization: normalize, sanitize, clean, canonicalize
- Validators: is_valid, validate, check_*
- Data structures: Custom collections with add/remove/get operations
- Mathematical/algorithmic: Pure functions, sorting, ordering
- Smart contracts: Solidity/Vyper contracts, token operations, state invariants
Supported Languages
- Python (Hypothesis)
- JavaScript/TypeScript (fast-check)
- Rust (proptest, quickcheck)
- Go (rapid, gopter)
- Java (jqwik)
- Scala (ScalaCheck)
- Solidity/Vyper (Echidna, Medusa)
- And many more...
See skills/property-based-testing/references/libraries.md for the complete list.
Evals
The skill ships three evals, because "the skill fires" and "the skill helps" are different claims and only the second one matters to a user.
./evals-extra/run.sh # trigger rate against labelled queries
EFFORTS=low ./evals-extra/effectiveness.sh # does the generated suite catch a real bug?Both spend real API budget — run.sh runs one session per query per run, 45 at its defaults (measured at 51.9 min and $36.50 at JOBS=4), and effectiveness.sh is 3. Neither runs in CI for that reason; they are what you run when you change the description or the guidance. RUNS=1 ./evals-extra/run.sh is the cheap smoke test.
Both harnesses ship a --self-test that costs nothing and runs in make check, so a harness that has stopped discriminating fails the build instead of reporting a green skill forever. run.sh --self-test drives the classifier with a stub binary and asserts, among other things, that a crashed session invalidates the sweep rather than being absorbed by the pass threshold.
effectiveness.sh grades by running the generated tests against a fixture with a known defect, not by reading what the model said about its own work. Both scripts exit non-zero when they inspect nothing, so a broken harness fails loudly instead of reporting a clean pass.
See skills/property-based-testing/README.md for what the queries cover and how to run the no-skill baseline.
evals/ — the ablation suite
The third eval is not a shell harness. evals/ holds claude plugin eval cases, run from the plugin root, and every case runs twice — once with the plugin loaded and once without — so the number it reports is Δ against the unaided model rather than a raw score. A skill that scores full marks in both arms is spending context and buying nothing, and that is the failure this suite exists to catch.
claude plugin eval . --ablation with-without --judge-model sonnet --allow-tools WriteTwo operational notes, both learned the hard way:
- It needs ANTHROPIC_API_KEY. Each case runs in a sandboxed config dir, so an interactive login is not visible to it and a subscription OAuth token is ignored. Without the key every session dies instantly as Not logged in, costs a cent, and the judge then grades that string — which scores zero and looks like a real result. Check apiKeySource in a trace before believing any number.
Before you install
- Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
- Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
- Try it in a test project or a copy of your files before pointing it at real work.
- Pin the version you tested, and review changes before updating.
- Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.
FAQ
What is Property Based Testing?
Property Based Testing is a plugin for Claude Code and Claude Cowork from the trailofbits/skills repository on GitHub. Write, review, and triage property-based tests — Hypothesis, fast-check, proptest, and Echidna or Medusa for Solidity invariants
How do I install Property Based Testing in Claude Code?
Add the repository as a plugin marketplace: claude plugin marketplace add trailofbits/skills Install the plugin: claude plugin install property-based-testing@<marketplace-name>, using the marketplace name from the repository's .claude-plugin/marketplace.json. Restart the session if the new skills or commands don't appear straight away.
Can I use Property Based Testing in Claude Cowork?
Open Customize → Plugins and choose Add marketplace. Enter trailofbits/skills (the owner/repo shorthand works for GitHub). Find Property Based Testing in the list, click Install, then connect any connectors it needs from its Connectors tab.
Is Property Based Testing safe to install?
It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.
Similar resources
- Workflow Skill Reviewer Reviews workflow-based Claude Code skills for structural quality, pattern adherence, tool assignment correctness, and anti-pattern… Subagent · trailofbits/skills
- Trailofbits:Semgrep Rule Creates Semgrep rules with test-first methodology Slash Command · trailofbits/skills
- Trailofbits:Variants Finds similar vulnerabilities using pattern-based analysis Slash Command · trailofbits/skills
- Variant Analysis Find similar vulnerabilities and bugs across codebases using pattern-based analysis Plugin · trailofbits/skills
- Rust Review Comprehensive Rust security code review with specialized bug-finding agents covering the safe/unsafe boundary, memory safety in unsafe blocks, concurrency, panic-induced DoS, recursion-induced stack overflow, FFI, and async runtime hazards Plugin · trailofbits/skills
- Mutation Testing Configures mewt or muton campaigns, analyzes surviving mutants, and investigates bugs exposed by testing gaps. Use when setting up mutation testing, reviewing campaign results, identifying equivalent mutants, or finding bugs from surviving mutations. Plugin · trailofbits/skills
- Seatbelt Sandboxer Generate minimal macOS Seatbelt sandbox configurations for applications Plugin · trailofbits/skills
- Modern Python Modern Python best practices. Use when creating new Python projects, and writing Python scripts, or migrating existing projects from legacy tools. Plugin · trailofbits/skills