do-and-judge

Execute a task with sub-agent implementation and LLM-as-a-judge verification with automatic retry loop

By neolabhq · 1,098 installs

npx skills add neolabhq/context-engineering-kit --skill do-and-judge

Source repository · Upstream listing

do and judge Task Execute a single task by dispatching an implementation sub agent, verifying with an independent judge, and iterating with feedback until passing or max retries exceeded. Arguments Argument Format Default Description task Free form text Required Task description to execute model haiku\ sonnet\ opus auto selected Explicit user override for all sub agents: implementation, meta judge, and judge. When omitted, you MUST select the model per the [Model Selection Policy]( model selection policy) — there is no fixed fallback tier. When provided, the user's choice wins over the policy for every sub agent — see the [Escalation Rule]( escalation rule) for how escalation interacts with an explicit override. strict strict false Disable the [Iteration Discretion Rule]( iteration discretion rule) the task passes ONLY when score = 4.0 , otherwise retry until max retries is reached. Example: /do and judge Refactor the UserService class to use dependency injection strict Context This command implements a single task execution pattern with meta judge → LLM as a judge verification . You (the orchestrator) dispatch a meta judge (to generate evaluation criteria) and an implementation agent in parallel , then dispatch a judge with the meta judge's evaluation specification to verify quality. If verification fails, you launch new implementation agent with judge feedback and iterate until passing (score ≥4, or accepted per the [Iteration Discretion Rule]( iteration discretion rule)) or max retries (3) exceeded. Key benefits: Fresh context Implementation agent works with clean context window Structured evaluation Meta judge produces tailored rubrics and checklists before judging External verification Judge applies meta judge specification mechanically — catches blind spots self critique misses Parallel speed Meta judge and implementation run simultaneously Feedback loop Retry with specific issues identified by judge Quality gate Work doesn't ship until it meets threshold CRITICAL: You are the orchestrator only you MUST NOT perform the task yourself. IF you read, write or run bash tools you failed task imidiatly. It is single most critical criteria for you. If you used anyting except sub agents you will be killed immediatly!!!! Your role is to: 1. Analyze the task and select the model per the [Model Selection Policy]( model selection policy) — sonnet / haiku by default, opus only when earned 2. Dispatch meta judge AND implementation agent in parallel as foreground agents (meta judge first in dispatch order) 3. Dispatch judge agent with meta judge's evaluation specification 4. Parse verdict and iterate if needed (max 3 retries) 5. Report final results or escalate RED FLAGS Never Do These NEVER: Read implementation files to understand code details (let sub agents do this) Write code or make changes to source files directly Skip judge verification to "save time" Read judge reports in full (only parse structured headers) Proceed after max retries without user decision ALWAYS: Use Task tool to dispatch sub agents for ALL implementation work Dispatch meta judge and implementation agent in parallel (meta judge FIRST in dispatch order) Wait for BOTH meta judge and implementation to complete before dispatching judge Pass meta judge evaluation specification to the judge agent Include CLAUDE PLUGIN ROOT= ${CLAUDE PLUGIN ROOT} in prompts to meta judge and judge agents Parse only VERDICT/SCORE/ISSUES from judge output Iterate with feedback if verification fails Model Selection Policy Picking the model is the single highest leverage decision you make — more than any prompt wording, it decides whether the task comes back correct and how long it takes. You MUST NOT treat it as a formality: name the tier and give a one line justification before dispatching. Reaching for the strongest model because you did not want to think is a failure, not caution. Tier default: sonnet and haiku are the default. opus is reserved and opt in — it MUST be earned by a trigger in the table below, never picked because you are unsure. Selection Rules Task shape Tier Examples Single documentation/text file correction — no code, no cross file reasoning haiku Fix a typo, update a link, correct a stale command in a README Small, few line (~10 lines or fewer), mechanical code change confined to one file haiku Bump a constant, add a guard clause, rename a local, edit a config value Code writing — new functions, components or tests, single module changes, established patterns sonnet Add an endpoint, write a service method plus tests, refactor one module Multi file refactoring (~3+ files, or any file count when a shared contract changes) OR critical (auth, payments/billing, data integrity, irreversible migration, public API break) OR complex logic (concurrency, non trivial algorithms, architectural decisions) opus Cross cutting refactor, auth or payment logic, schema migration, novel algorithm design Precedence (MANDATORY): evaluate EVERY row, not just the first that matches. When more than one row matches, the HIGHEST matching tier wins — criticality and complexity always override size. A four line null check inside a security critical auth handler matches both the haiku row and the opus row, and is therefore opus . The critical list is exhaustive, not illustrative: shipping to production, touching real users, or adding to a public API are NOT triggers, so a new endpoint with validation in one service file stays sonnet . Mechanical breadth carve out: breadth alone is not complexity — a purely mechanical change (e.g., renaming a symbol across many files, with no logic or contract change) stays at the tier its content earns no matter how many files it touches, so mechanically renaming a symbol across 40 files is haiku or sonnet work, not opus ; this carve out does NOT cover a shared contract change (already an opus trigger above), so extracting a shared interface across files remains opus . Tie breaker: ONLY when no row matches cleanly — the task sits genuinely between two tiers — pick the cheaper tier. You MUST NOT bias up to opus to hedge; the [Escalation Rule]( escalation rule) makes a cheap first guess recoverable, and one recovered run costs far less than over provisioning every run. Role Pairing Any model assigned pipeline has up to three roles — producer (does the work), criteria setter (defines what "correct" means), evaluator (checks the work against those criteria); in this skill they instantiate as implementation / meta judge / judge. Default: the SAME tier for all roles — and where a pipeline has no separate criteria setter (e.g. a plan step or stage simply assigned a model), this default is the whole rule. Only for a non obvious task you MAY raise the criteria setter alone by one tier, so the criteria are sharper than the work being evaluated. Non obvious is testable: the tier was decided by the Tie breaker (no Selection Rules row matched cleanly), OR the task states no checkable acceptance condition. Pattern Criteria setter (meta judge) Producer + evaluator (implementation + judge) Use when Sharpened haiku sonnet haiku The work is trivial, but what counts as "correct" is not obvious Sharpened sonnet opus sonnet Code work with ambiguous or high consequence acceptance criteria that does not itself hit an opus trigger Producer and evaluator MAY be a differnt tier. You MAY decide to raise the evaluator alone if criteria list produced by criteria setter looks too complex, but you MUST NOT set the criteria setter below the producer tier. Escalation Rule Bump BOTH producer and evaluator (implementation and judge) one tier for the next iteration when either trigger fires: 1. Low first iteration quality — a low score, or issues showing the model misunderstood the task rather than merely missing details. 2. The user complains that quality is too low or the results are wrong — at any point, including after a reported PASS. Ladder: haiku → sonnet → opus . opus is the ceiling — there is no further tier. If opus tier work still fails, escalate to the user , never loop. Explicit model carve out (the ONLY statement of this rule): an explicit model is a user override, so trigger (1) MUST NOT silently overrule it — continue iterate with override model till you reach max retry limit. If target still not meet at the end, highlight the found issues and propose to the bump to user. Trigger (2) IS that approval, so it bumps immediately. Escalation moves producer and evaluator only. A criteria setter that already produced the evaluation specification is NOT re run and NOT re tiered — changing the criteria mid task invalidates the comparison across attempts. Escalation is a complement to, never a substitute for, a genuine root cause fix. You MUST still pass the judge's specific feedback into the retry; re dispatching the same prompt at a higher tier and hoping is prohibited. Escalation is orthogonal to the score thresholds and the [Iteration Discretion Rule]( iteration discretion rule) — it changes which model runs the next iteration, never whether an iteration is warranted. Cross Provider Equivalence When this skill runs outside the Anthropic model context, map the tier to the nearest model of the same class: Tier Role Comparable models from other providers haiku Fast and cheap; mechanical work gemini flash lite , gemma class, gpt oss class, small open weight models sonnet Balanced workhorse; most code writing gemini pro class and full gemini flash ( not the lite variant, which is haiku tier), GPT 5 mini class, large Qwen / DeepSeek class opus Frontier reasoning; critical or complex work whatever the provider sells as its extended / deliberate reasoning tier — currently GPT 5.5 , deep think modes, Kimi K3 class, any model whose advantage is longer deliberation rather than throughput The mapping is by capability tier, not by name — exact names drift as vendors ship new models. Every rule above is expressed in tiers, so on another provider: map tier → your model of that class, then apply the selection, pairing and escalation rules unchanged. Process Phase 1: Task Analysis and Model Selection Resolve configuration first: STRICT MODE = strict present false . Strip all flags from the task text — never pass them into sub agent prompts. Unless the user passed model , assess the task on three axes, then read the tier straight off the [Selection Rules]( selection rules) table: Scope — one file, one module, or multiple files? Complexity — mechanical edit, established pattern, or novel/intricate logic? Risk — isolated and reversible, internal, or critical per the exhaustive list in the [Selection Rules]( selection rules) opus row? State the three findings, the chosen tier, and a one line justification before dispatching. Then apply [Role Pairing]( role pairing) to decide the meta judge tier — same tier as implementation unless the task is genuinely non obvious. If the user passed model , neither step runs: that one tier is used for implementation, meta judge and judge alike, and Role Pairing MUST NOT raise the meta judge above it. Specialized Agents: Common agents from the sdd plugin include: sdd:developer , sdd:researcher , sdd:software architect , sdd:tech lead , sdd:business analyst . If the appropriate specialized agent is not available, fallback to a general agent without specialization. You MUST use general purpose every time, when there no direct coralation between task and specialized agent, or agent is not available! Phase 2: Dispatch Meta Jud