hunt
Finds root cause before applying fixes for errors, crashes, regressions, failing tests, broken behavior, and screenshot-reported defects. Use when users report in any language errors, crashes, broken behavior, regressions, failing tests, screenshot evidence, or something that used to work and now fa
By tw93 · 13,766 installs
npx skills add tw93/waza --skill hunt
Source repository · Upstream listing
Hunt: Diagnose Before You Fix
Prefix your first line with 🥷 inline, not as its own paragraph.
A patch applied to a symptom creates a new bug somewhere else.
Outcome Contract
Outcome: the root cause is identified before any fix is applied.
Done when: one sentence explains the cause, every observed symptom fits it, and the fix or handoff is verified against a reproducible check.
Evidence: source trace, repro command or UI path, logs or state, targeted test/build output, and runtime evidence for UI or native defects.
Output: root cause, fix or handoff, verification result, and any unswept sibling risks.
Authorization: "diagnose", "investigate", "why", "look into", "排查", "看看", or equivalent is report only. Apply a fix only when the current turn explicitly asks to fix, change, implement, or optimize; root cause proof is still required first.
Do not touch code until you can state the root cause in one sentence:
"I believe the root cause is [X] because [evidence]."
Name a specific file, function, line, or condition. "A state management issue" is not testable. "Stale cache in useUser at src/hooks/user.ts:42 because the dependency array is missing userId " is testable. If you cannot be that specific, you do not have a hypothesis yet.
Diagnosis Signals
Hypothesis quality gate: the hypothesis must explain every observable symptom, not just the one reported first; partial coverage is a symptom level guess, not a root cause. A symptom the reporter waves off as unrelated is still a symptom the hypothesis has to cover. For timing dependent issues (flicker, intermittent failure, race), reproduce reliably before diagnosing.
Rationalization smells: "I'll just try this" = no hypothesis, write it first. "I'm confident" = run the instrument that proves it. "Probably the same issue" = re read the execution path from scratch. "It works on my machine" = enumerate env differences before dismissing. "One more restart" = read the last error verbatim; never restart more than twice without new evidence.
Durable Context Preflight
See [references/durable context.md](references/durable context.md) for when durable context is in scope and the redaction gate that applies before any of it becomes a durable rule.
For /hunt : durable context is hypothesis fuel only, and current code, logs, and repro evidence override memory. It never replaces a fresh root cause sentence or a reproducible symptom list.
Fix Scope Discipline
If the bug needs a prerequisite refactor (e.g. a shared interface must change), state why it is necessary and check the authorized scope. Continue if that work is covered; ask before expanding the scope or choosing an unresolved behavior tradeoff. Keep unrelated refactors separate.
Bisect Mode
Activate when: "以前是好的", "之前是好的", "used to work", "上一次提交还是对的", "broke after update", or the user remembers a specific good commit or version.
Protect the user's worktree first: git status short branch uall . Any modified, staged, or untracked files mean no bisect in the current checkout: run it in a temporary detached worktree and remove that worktree when done. If a temporary worktree is impossible, stop and ask for explicit cleanup/stash approval.
If the last good version is only a few releases back, git diff <last good ..HEAD <suspect path and read the delta first. The regression is usually visible there at a fraction of a bisect's cost; fall through to bisect only when the diff is too large or the culprit is not obvious.
Bisect only with a non interactive pass/fail command defined up front, and keep the bookkeeping in git ( git bisect good/bad ), including when you test a suspect commit directly. When it names the culprit, read only that diff down to the specific line, then run git bisect reset before removing the temporary worktree.
Repeated Regression / Screenshot Reference Mode
Activate when the user says the same issue is still wrong after a fix, provides a "good" screenshot/version/file, or describes a visual result as previously correct.
Treat the reference as evidence, not decoration: list every reported and visible symptom in the user's concrete words; identify the reference oracle (last good commit, old build, fixture, screenshot, described expected state); define the pass/fail check before editing; then name the exact current vs reference delta. Do not generalize a visual defect into "style polish" when the evidence points to a broken render, race, font pipeline, or state path.
If the issue is purely subjective UI taste, route to /ui . If it is rendering, state, timing, build output, font generation, or a regression from a known good version, stay in /hunt .
Scope Blast Mode
Activate after fixing a root cause pattern, before declaring the bug done; also when the user says "举一反三", "举一反三深入看看", or "其他地方有没有同样问题". The same shape often hides in N other places; one local fix that ignores the blast leaves N 1 bugs in the tree.
Extract the pattern signature (the specific function, regex, API call, CSS selector, lock acquisition, validation skip, or input boundary that produced the bug) and grep rn it across the repo, excluding generated dirs, build output, and vendored deps; for class of bug patterns ("any handler missing the lock"), grep the surrounding shape, not just the literal text. For every match, answer in writing: same bug / safe to leave (why) / unsure (ask the user). Do not silently skip a match, and do not claim "fixed" until the blast report is in the Output block. Unrelated bugs the sweep surfaces get listed, not fixed in this PR, unless the user agrees.
Confirm or Discard
Run the one probe that would fail if the hypothesis were wrong, then read it. If the evidence contradicts the hypothesis, discard it completely and re orient on what the probe just showed. Do not stack a fix onto a disproven hypothesis, and do not keep one just because the code "looks like" the cause.
Runtime Evidence Ladder
Use this ladder before claiming a bug is fixed:
1. Source trace: name the exact function, state transition, file, line, or condition that can produce the symptom.
2. Deterministic repro: run or write the smallest command, fixture, UI path, or scenario that produces it.
3. Logs/state/cache: inspect the runtime state that proves the path was reached, including queues, DB rows, caches, temp files, generated outputs, or external tool logs.
4. Build/test: run the narrow test or build that exercises the fix.
5. Real runtime check: for UI, native app, browser, rendering, or visual bugs, open the app/page/artifact and verify the visible result with a screenshot or concrete checklist.
Compile only is not enough for UI, native app, visual, rendering, or generated artifact bugs. If the runtime check is impossible in the environment, say why and hand off the exact screen, command, or artifact to verify.
When the reporter's environment is the missing rung and it cannot be reproduced locally, the next artifact is a read only probe they can paste and run, not another hypothesis. Have it print the environment, the disputed measurement, and the state of whatever the hypothesis turns on, and nothing that could carry a secret or a private path. Assume none of your own layout: their install method, directory conventions, locale, shell, and version all differ, so discover rather than hardcode. Ship it as plain copyable text with one command to run and one block to paste back.
For recurring classes of failures, load references/failure patterns.md before adding a second fix.
Native App Freeze Mode
For beachballs, not responding, tab switch freezes, first open lag, idle wake stalls, overlay lockups, or frozen app screenshots, load references/logging techniques.md (Native App Freeze Mode) before changing code.
Targeted Logging
Every log is a yes/no question: "if this prints X before Y, hypothesis A survives; otherwise A is dead." A log that cannot rule a hypothesis in or out is noise. Remove temporary logs before finishing; gate persistent diagnostics behind the project's debug flag. If adding a log changes the behavior, that is itself evidence of a timing, lifecycle, or concurrency problem. Full playbook: references/logging techniques.md .
Rendering Bug Mode
For PDF output, page breaks, font rendering, or print layout defects, load references/rendering debug.md ; it carries the activation triggers and the diagnosis checklist (WeasyPrint quirks, font loading, page overflow, browser print CSS).
IME / Unicode Issues
For input method, character rendering, or text encoding bugs (IME state, cursor drift, emoji splitting, composition events), check references/ime unicode.md first before forming a hypothesis.
Hard Rules
Same symptom after a fix is a hard stop; so is "let me just try this." Both mean the hypothesis is unfinished. Re read the execution path from scratch before touching code again.
After three failed hypotheses, stop. Use the Handoff format below to surface what was checked, what was ruled out, and what is unknown. Ask how to proceed.
External tool failure: diagnose before switching. When an MCP tool or API fails, determine why first (server running? API key valid? Config correct?) before trying an alternative.
System/tooling symptoms need a lower layer baseline. Before blaming the visible app, generated file, or top level feature, measure the raw lower layer first: OS capture versus post processing, runtime service versus UI, compiler/toolchain versus test assertion, network/API versus client handling. Retire hypotheses that the baseline disproves instead of circling them.
Visual/rendering bugs: static analysis first. Trace paint layers, stacking contexts, and layer order in DevTools before adding console.log or visual debug overlays. Logs cannot capture what the compositor does. Only add instrumentation after static analysis fails.
Behavioral / lifecycle / async bugs: instrument while forming the hypothesis. Window lifecycle, event delivery, navigation, focus, timer, state machine, and async ordering bugs almost never yield to static reading alone. The moment the hypothesis involves "this callback fires before/after that one", "this state should be X when Y runs", or "this object should still be alive here", add the log before writing any fix (anti pattern 28); two guesses in a row is the hard stop signal. Compositor behavior needs DevTools, not logs; pure logic bugs (wrong formula, off by one) need only static analysis.
Tuning magic numbers past round three: stop, unify. When a spacing / sizing / threshold value has been adjusted three times and still looks wrong, the bug is structural, not numeric. Replace the N independent values with one named token ( Spacing.s4 , gap content , etc.) and verify the asymmetry was hiding a missing constraint. Asymmetry that survives tuning is structural; more tuning will not converge.
Performance complaints need numbers. For "slow", "laggy", or memory growth reports outside Native App Freeze Mode, measure the baseline first (wall clock time, profile sample, memory footprint), fix, then re measure and report before/after numbers. "Feels faster" is not evidence.
Fix the cause, not the symptom. Continue necessary fixes within the user's authorized scope. Ask only when the fix expands that scope or requires a user decision; file count alone is not an approval boundary.
Gotchas
What happened Rule
Patched the wrong copy of a duplicated surface Trace the execution path backward to the instance that actually renders before touching any file
Orchestrator reported RUNNING while a downstream stage was misconfigured In multi stage pipelines, test each stage in isolation
Race condition diagnosed as a stale state bug For timing sensitive issues, inspect event timestamps and ordering before state
Reproduced locally but failed in CI Align the environment first (runtime version, env vars, timezone