asb-interview-learning

Facilitates the learning step of a proven customer-interview method — the synthesis half: reads a directory of per-interview debrief files against the working HYPOTHESES.md and QUESTIONS.md and proposes evidence-cited updates ONE at a time — double down, tune numbers, mark disproved, park heard-once

By asmartbear · 442 installs

npx skills add asmartbear/asb-skills --skill asb-interview-learning

Source repository · Upstream listing

Learning: Let the Interviews Rewrite Your Hypotheses Written predictions only pay off at reconciliation: the hypothesis list was recorded precisely so that reality could argue with it, and this is the step where the argument happens. The skill reads the accumulated per interview debriefs against the working hypothesis and question files, builds a queue of proposed changes with the evidence for each, and walks the user through it one proposal at a time — the user decides every change, agreed changes land in the files immediately, and the session ends with a real decision about the interviewing itself: keep going, stop and act, or admit the answers are diverging and re aim. The mental model The three tier update rule Not everything heard deserves a reaction, and the tiers are the core of this step: A stray voice — one person said something contradictory. Change nothing: in the real world nothing is universal, and a belief file that flinches at every anecdote never converges. But notice it: the observation goes in a "That's funny" watch section, on the record, so the next synthesis checks whether it's become a pattern. A pattern — the same story, number, or attitude across several conversations; never claimed from a corpus of only one or two, where the first tier governs no matter how unanimous it looks. Update the hypothesis: tune its number, tighten its segment, or mark it disproved. Patterns are the fundamental truth the interviews exist to find. A revelation — even one customer says something that strikes the user as revelatory, reframing how they see the problem. That single voice may legitimately update a hypothesis or spawn a new one. The test is not vote counting but the felt shock of it — surprise is the signal that learning is happening. The watch section takes its name from the old line that the most exciting phrase in science isn't "eureka" but "that's funny." Anything noticed but below pattern depth parks there — heard once, or even twice in a small corpus — with a weight note ("heard in two of five; priority watch") when it's knocking on the pattern door. The promotion path (funny → pattern → hypothesis change) is this skill re run as debriefs accumulate. Expect contradictions — and "no pattern" is still learning People differ: different goals, roles, past experiences, or no discernible reason at all. The data will be noisy, and some hypotheses resolve not to true or false but to "this varies wildly" or "there is no pattern here." Record that as the resolution — knowing a pattern doesn't exist prevents building on a false assumption. Real patterns stand out from noise; that contrast is the finding. Convergence is what truth feels like Across many interviews, a validated idea behaves like a law of nature: the more people asked, the more the answers agree — same pain, same acceptable solution, same money. A weak idea does the opposite: everyone is positive, but each conversation points a different direction — different buyer, different price, different product — a Venn diagram with twenty lobes and no center. Convergence and divergence, not enthusiasm, are the read on whether the interviews are closing in on something. Watch which one is happening; it drives the final verdict. Emergent segmentation Sometimes the "contradiction" is structure: one type of customer answers one way, another type answers another — marketing departments think about security, solo bloggers never do. When answers cluster by customer type, propose making the segment explicit: rewrite the affected hypotheses to name whom they're about ("Freelancers building client sites will pay …" vs. "Solo owners will …" — two hypotheses, two numbers), and make sure an early interview question sorts which segment the interviewee belongs to, so every later answer gets filed under the right lens. Keep it in one hypothesis file with segment scoped claims; the segmentation itself is one of the most valuable findings this step can produce. Conflicting signals get a choice, not "a balance" When the debriefs pull in two directions — some want cheap and simple, some want premium and full service — the lazy synthesis is "it's a balance," which is usually a refusal to decide. Force the real resolution: A balance is correct only when both extremes are genuinely bad and something in between beats either. Rare in interview findings. A choice is correct when both signals are rational but contradictory — which is usually an emergent segment, and the question becomes which segment is the ideal customer. That's the user's decision to make with eyes open, not this skill's. A choice to a limit : maximize one, hold the other above a threshold — a goal and a governor, not a compromise. Why not both — occasionally the conflict points at an invention no one asked for: a new offering that satisfies both signals at once, as when customers who happily paid full price for big sites resented paying it for trivial side sites, and the answer was neither price point but multi site plans — a thing no competitor offered. The test for such a move is the strategist's question: what else would have to be true for this to work? An invention plus the named set of supporting decisions is a strategy; an invention alone is a wish. When a proposal involves conflicting signals, present which of these shapes the conflict has — and if it's a choice, say so plainly instead of splitting the difference. Stop when it's boring When the surprises cease, learning has ceased, and it's time to stop interviewing and start acting. There's no magic number of interviews — three is definitely too few (that's a marketer trying two ad variants and quitting); ten has validated companies; one famous validation took forty; some take over a hundred. The signals, made concrete: Surprise rate. Are recent debriefs still producing surprises (well kept debriefs mark them, e.g. with ❗) and new addenda themes, or is the same stories same numbers same language pattern setting in? Falling surprise rate = approaching done. Convergence. Are the core hypotheses (pain, coping, money) settling to stable resolutions — or still scattering? Would more change anything? If five more identical conversations wouldn't change a single decision, they're not worth having. Honesty checks , for when the numbers don't speak clearly: don't let sunk cost decide (interviews already scheduled are not a reason to keep learning nothing); timebox the remainder rather than drifting; and the deep inside test — the user often already knows the answer and doesn't want to admit it. Three verdicts are possible, and one must be delivered: continue (surprises still coming — and say what the remaining interviews should focus on: which hypotheses are unresolved, which segments are underrepresented); stop and act (converged — the validated facts are ready to be combined with strategy); or re aim (everyone is polite but the answers diverge with no center — the problem may not be the interviews but the idea or the audience; the next move is different hypotheses or different people, not more of the same conversations). "Let's just do a few more and see" is the non answer this step exists to refuse — if more interviews are the call, it comes with a focus and a number. And a question more interviews cannot answer — which segment to serve, what to build, what to charge — is never a reason to continue: when the only open items are strategy choices, the verdict is stop and act. Vocabulary Debrief — the per conversation record this skill consumes: brief answers mapped to Q numbers (which carry H numbers), plus addenda of slotless findings, one file per interview. Double down — a hypothesis the debriefs confirm; recorded as validated, and worth leaning into when acting. Tune — keep the kind of claim, correct the number, threshold, or segment. Disproved — reality said no. The hypothesis stays in the file, marked, because a disproven belief is a finding — and its number is never reused. That's funny — the watch section for observations below pattern depth: noticed, recorded, not yet acted on. Change log — one line per change at the bottom of each working file; the trail that makes the current file trustworthy. The synthesizer's posture Be clear, not clever Write to be understood, not admired. The work here wrestles with hard concepts, and clever metaphors, wordplay, or cute turns of phrase make them harder to grasp, not easier. Say plainly what you mean. If a sentence reads more clearly without a flourish, cut the flourish. State the actual point rather than gesturing wittily at it. Restate references; never cite a bare token When you mention a numbered or lettered item to the user — K4, W2, O17, H3, and the like — add a few plain words on what it actually is ("K4 — the owner whose career rides on the site"). A bare token is unreadable to a human who saw it defined hours or days ago: the tag is for traceability, the gloss is for comprehension. Keep the tag for accuracy; always add the gloss. You propose, the user decides Half the value of this exercise is the user thinking it through — the "aha" comes from wrestling with the contradictions, not from reading a memo about them. So this skill never batch applies anything: every change is proposed singly, argued from evidence, and applied only when the user accepts it (with whatever adjustment they make — it's their belief file). But deciding is not the same as waving through: batch nodding ("sure, apply all nine") gets declined, and when the user keeps a hypothesis against strong contrary evidence, make them defend it once — "five of seven said the opposite; what do you know that they don't?" — then record their call. Their genuine belief goes in the file even when the evidence weighing would go the other way; the log records what the evidence said. Every proposal cites its evidence A proposal without citations is an opinion. Each one names the debrief files behind it and quotes the operative words — "2026 06 30 dana.md: 'I'd switch tomorrow if migration were handled'" — so the user can weigh the evidence, not the summary of it. Market guru flagged material weighs almost nothing as evidence about the market and full weight as evidence about that speaker — and apply the same discount to guru shaped material the debriefs failed to flag; "most people would…" is hearsay whoever recorded it. Never pad: if only two debriefs speak to a hypothesis, say two, and let the tiers do their work. One proposal per exchange The opening move is small: what was read, how many debriefs are new since the last synthesis, the queue's shape in one line per category — then the first proposal, which may ride in that same opening message: one fully formed proposal is a start, not a wall; two is a wall. After that, strictly one proposal per exchange: presented, decided, applied, logged, next. A user who can't react to each is being performed for, not facilitated. If the user asks to speed up, compress the ceremony (shorter evidence displays, quicker confirms), never the structure — a consolidated diff of everything applied, delivered after the walk, is a fine courtesy; batch review is legitimate, batch deciding never is. Numbers are frozen; the log is mandatory These mechanics are non negotiable, however the user pushes, because other artifacts cite these numbers: A hypothesis or question number is never renumbered and never reused , even for a disproved or retired entry. A rewrite that keeps the claim's subject — tuning a number, sharpening wording, scoping a condition — edits in place under its own number, with the log preserving what changed. A rewrite that changes whom or what the claim is about (a segmentation split, a different actor) retires the o