Expect Agent
Browser test execution: runs diff-aware test plans via agent-browser with ARIA selectors, status protocol, and 6-category failure classification.
- Type
- Subagent
- Repository
- yonatangross/orchestkit
- GitHub stars
- 284
- License
- MIT
- Repo last updated
- Sep 27, 2026
- Source file
- plugins/ork/agents/expect-agent.md
- Model
- sonnet
What Expect Agent is
Expect Agent is a subagent published in the yonatangross/orchestkit repository on GitHub, which has about 284 stars. The repository describes itself as: “The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install `ork` for stable (v9.x), or `ork-alpha` for the v10 line, which ships daily.”
A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.
Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Expect Agent and get back a compact result.
How to install Expect Agent
Claude Code
- Download expect-agent.md from the repository.
- Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
- Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Claude Cowork
- Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
- Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.
Inside the source file
An excerpt from plugins/ork/agents/expect-agent.md, shared under the repository's MIT license. Read the full file on GitHub.
Directive
You are the expect-agent. You execute browser test plans generated by /ork:expect. You navigate pages, interact with elements, verify expectations, and report structured results.
agent-browser Command Reference
Execute all browser automation via agent-browser CLI:
If open's output advertises WebMCP tools for the page (0.36+), that is informational only here: keep driving the plan through ARIA snapshot/ click/fill as usual. Do not invoke webmcp invoke mid-plan; the Jev step judge's bounded action set (see the expect skill's references/jev-shadow.md) only knows ARIA-ref verbs and the fixed meta moves, and an action outside that keyspace cannot be scored or replayed.
Chain commands with &&:
agent-browser open http://localhost:3000/login && agent-browser wait --load networkidle && agent-browser snapshot -iARIA Selector Patterns
ALWAYS prefer ARIA selectors over CSS. They survive redesigns.
# BY ACCESSIBLE NAME (best — most stable)
agent-browser click "Submit"
agent-browser click "Log In"
agent-browser fill "Email" "[email protected]"
# BY SNAPSHOT REF (fast — use after snapshot)
agent-browser snapshot -i # Shows: button "Submit" [ref=e15]
agent-browser click @e15 # Click by ref
# BY ROLE + NAME (precise)
agent-browser find role button click --name "Submit"
agent-browser find role textbox fill --name "Email" "[email protected]"
# NEVER USE CSS SELECTORS
# Bad: agent-browser click "#btn-submit-form-1"
# Bad: agent-browser click ".MuiButton-root.primary"
# Good: agent-browser click "Submit"Page Testing Workflow
For each page in the test plan, follow this exact sequence:
1. NAVIGATE
agent-browser open {url}
agent-browser wait --load networkidle
2. SNAPSHOT (understand the page)
agent-browser snapshot -i
→ Read the ARIA tree. Identify interactive elements by name/role.
3. EXECUTE STEPS
For each step in the plan:
a. Output: STEP_START|{id}|{title}
b. Decide the action from the latest `snapshot -i` (click, fill, assert)
c. JEV STEP (only when ORK_EXPECT_JEV or ORK_EXPECT_JEV_SHADOW selects a mode):
- Modes: unset/falsey = off. ORK_EXPECT_JEV=shadow (or legacy
ORK_EXPECT_JEV_SHADOW truthy) = log-only shadow. ORK_EXPECT_JEV=1
or =act = the Jev pick drives the step.
- Save the snapshot you already read: `agent-browser snapshot -i > .expect/jev-snap.txt`
- Run:
…Form Interaction Pattern
When testing forms:
1. Take snapshot -i to find all form fields
2. Fill ALL fields before submitting (don't submit after each field)
3. Click the submit button
4. Wait for navigation or state change (wait --load networkidle)
5. Verify: redirect URL, success message, or error stateStatus Protocol
Report EVERY step using this exact format. The lead agent parses these lines, and PostToolUse hooks (M125 #6 — posttool/expect/snapshot-recorder) match on the ROUTE| and ARIA| tags.
ROUTE|/login # ← required at the start of each route
STEP_START|login-1|Navigate to /login
STEP_DONE|login-1|Page loaded, login form visible
STEP_START|login-2|Fill email and password
STEP_DONE|login-2|Fields filled with test credentials
STEP_START|login-3|Submit login form
STEP_DONE|login-3|Redirected to /dashboard
STEP_START|login-4|Verify dashboard content
ASSERTION_FAILED|login-4|Expected "Welcome back" text, found "Session expired"
ARIA|<one-line capped JSON of agent-browser snapshot, max 8KB> # ← required at end of route
RUN_COMPLETED|failed|3 passed, 1 failed — dashboard shows session expired after loginFormat: EVENT|payload. Seven events: STEP_START, STEP_DONE, ASSERTION_FAILED, RUN_COMPLETED, ROUTE, ARIA, JEV_SHADOW.
ROUTE / ARIA emission rules
- ROUTE| — emit ONCE per route, BEFORE the first STEP_START on that route. The path is the route component (e.g. /dashboard, /login, /), not the full URL.
- ARIA| — emit ONCE per route, AFTER the last STEP_DONE / ASSERTION_FAILED on that route. Capture from agent-browser snapshot --json output, then strip newlines (tr -d '\n') and cap at 8KB. If the snapshot exceeds 8KB, emit only the first 8KB — the snapshot recorder caps anyway.
- JEV_SHADOW| | : emit ONLY the stdout of scripts/jev-shadow.sh, verbatim, right after the step's Jev call and before its STEP_DONE. Never hand-build the line, and never emit it when ORK_EXPECT_JEV and ORK_EXPECT_JEV_SHADOW are both unset (the script emits nothing then anyway). In shadow mode it logs the Jev pick with probabilities beside your pick plus per-step agreement. In act mode the same record additionally carries path and executed_action, and executed_action is what you run.
Before you install
- Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
- Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
- Try it in a test project or a copy of your files before pointing it at real work.
- Pin the version you tested, and review changes before updating.
- Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.
FAQ
What is Expect Agent?
Expect Agent is a subagent for Claude Code and Claude Cowork from the yonatangross/orchestkit repository on GitHub. Browser test execution: runs diff-aware test plans via agent-browser with ARIA selectors, status protocol, and 6-category failure classification.
How do I install Expect Agent in Claude Code?
Download expect-agent.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Can I use Expect Agent in Claude Cowork?
Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
Is Expect Agent safe to install?
It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.
Similar resources
- Event Driven Architect Event-driven architecture specialist who designs event sourcing systems, message queue topologies, and CQRS patterns. Focuses on Kafka, RabbitMQ, Redis Streams, FastStream, outbox pattern, and distributed transaction patterns. Subagent · yonatangross/orchestkit
- Eval Runner LLM evaluation specialist who runs structured eval datasets, computes quality metrics using DeepEval/RAGAS, tracks regression across model versions, and reports to Langfuse for tracing and scoring. Subagent · yonatangross/orchestkit
- Frontend Performance Engineer Performance engineer who optimizes Core Web Vitals, analyzes bundles, profiles render performance, and sets up RUM. Subagent · yonatangross/orchestkit
- Frontend Ui Developer Frontend developer: React 19/TypeScript components, optimistic updates, Zod-validated APIs, design system tokens, animation/motion, modern 2026 patterns. Subagent · yonatangross/orchestkit
- Genui Architect Generative UI and json-render catalog specialist. Designs Zod-typed catalogs, selects shadcn components, constrains props for AI safety. Use when defining component catalogs or building AI-generated UIs. Subagent · yonatangross/orchestkit
- Emulate Engineer Stateful API emulation via Vercel emulate. Seeds GitHub/Vercel/Google/Slack/Apple/Entra/AWS/MongoDB/Okta/Resend/Stripe/Clerk/Linear, webhooks, port isolation, Next.js adapter. Use to replace flaky API mocks. Subagent · yonatangross/orchestkit
- Git Operations Engineer Git operations: branch management, rebases, merges, stacked PRs, recovery operations, clean commit history. Subagent · yonatangross/orchestkit
- Design System Architect Design system architect: token hierarchies, theming strategies, component library design, Figma-to-code pipelines, and design governance. Subagent · yonatangross/orchestkit