asb-interview-learning
Facilitates the learning step of a proven customer-interview method — the synthesis half: reads a directory of per-interview debrief files against the working HYPOTHESES.md and QUESTIONS.md and proposes evidence-cited updates ONE at a time — double down, tune numbers, mark disproved, park heard-once
By asmartbear · 442 installs
npx skills add asmartbear/asb-skills --skill asb-interview-learning
Source repository · Upstream listing
Learning: Let the Interviews Rewrite Your Hypotheses
Written predictions only pay off at reconciliation: the hypothesis list
was recorded precisely so that reality could argue with it, and this is
the step where the argument happens. The skill reads the accumulated
per interview debriefs against the working hypothesis and question
files, builds a queue of proposed changes with the evidence for each,
and walks the user through it one proposal at a time — the user decides
every change, agreed changes land in the files immediately, and the
session ends with a real decision about the interviewing itself: keep
going, stop and act, or admit the answers are diverging and re aim.
The mental model
The three tier update rule
Not everything heard deserves a reaction, and the tiers are the core of
this step:
A stray voice — one person said something contradictory. Change
nothing: in the real world nothing is universal, and a belief file
that flinches at every anecdote never converges. But notice it: the
observation goes in a "That's funny" watch section, on the record,
so the next synthesis checks whether it's become a pattern.
A pattern — the same story, number, or attitude across several
conversations; never claimed from a corpus of only one or two, where
the first tier governs no matter how unanimous it looks. Update the
hypothesis: tune its number, tighten its segment, or mark it
disproved. Patterns are the fundamental truth the interviews exist to
find.
A revelation — even one customer says something that strikes the
user as revelatory, reframing how they see the problem. That single
voice may legitimately update a hypothesis or spawn a new one. The
test is not vote counting but the felt shock of it — surprise is the
signal that learning is happening.
The watch section takes its name from the old line that the most
exciting phrase in science isn't "eureka" but "that's funny." Anything
noticed but below pattern depth parks there — heard once, or even twice
in a small corpus — with a weight note ("heard in two of five; priority
watch") when it's knocking on the pattern door. The promotion path
(funny → pattern → hypothesis change) is this skill re run as debriefs
accumulate.
Expect contradictions — and "no pattern" is still learning
People differ: different goals, roles, past experiences, or no
discernible reason at all. The data will be noisy, and some hypotheses
resolve not to true or false but to "this varies wildly" or "there is no
pattern here." Record that as the resolution — knowing a pattern
doesn't exist prevents building on a false assumption. Real patterns
stand out from noise; that contrast is the finding.
Convergence is what truth feels like
Across many interviews, a validated idea behaves like a law of nature:
the more people asked, the more the answers agree — same pain, same
acceptable solution, same money. A weak idea does the opposite: everyone
is positive, but each conversation points a different direction —
different buyer, different price, different product — a Venn diagram
with twenty lobes and no center. Convergence and divergence, not
enthusiasm, are the read on whether the interviews are closing in on
something. Watch which one is happening; it drives the final verdict.
Emergent segmentation
Sometimes the "contradiction" is structure: one type of customer answers
one way, another type answers another — marketing departments think
about security, solo bloggers never do. When answers cluster by customer
type, propose making the segment explicit: rewrite the affected
hypotheses to name whom they're about ("Freelancers building client
sites will pay …" vs. "Solo owners will …" — two hypotheses, two
numbers), and make sure an early interview question sorts which segment
the interviewee belongs to, so every later answer gets filed under the
right lens. Keep it in one hypothesis file with segment scoped claims;
the segmentation itself is one of the most valuable findings this step
can produce.
Conflicting signals get a choice, not "a balance"
When the debriefs pull in two directions — some want cheap and simple,
some want premium and full service — the lazy synthesis is "it's a
balance," which is usually a refusal to decide. Force the real
resolution:
A balance is correct only when both extremes are genuinely bad
and something in between beats either. Rare in interview findings.
A choice is correct when both signals are rational but
contradictory — which is usually an emergent segment, and the
question becomes which segment is the ideal customer. That's the
user's decision to make with eyes open, not this skill's.
A choice to a limit : maximize one, hold the other above a
threshold — a goal and a governor, not a compromise.
Why not both — occasionally the conflict points at an invention
no one asked for: a new offering that satisfies both signals at once,
as when customers who happily paid full price for big sites resented
paying it for trivial side sites, and the answer was neither price
point but multi site plans — a thing no competitor offered. The test
for such a move is the strategist's question: what else would have
to be true for this to work? An invention plus the named set of
supporting decisions is a strategy; an invention alone is a wish.
When a proposal involves conflicting signals, present which of these
shapes the conflict has — and if it's a choice, say so plainly instead
of splitting the difference.
Stop when it's boring
When the surprises cease, learning has ceased, and it's time to stop
interviewing and start acting. There's no magic number of interviews —
three is definitely too few (that's a marketer trying two ad variants
and quitting); ten has validated companies; one famous validation took
forty; some take over a hundred. The signals, made concrete:
Surprise rate. Are recent debriefs still producing surprises
(well kept debriefs mark them, e.g. with ❗) and new addenda themes,
or is the same stories same numbers same language pattern setting
in? Falling surprise rate = approaching done.
Convergence. Are the core hypotheses (pain, coping, money)
settling to stable resolutions — or still scattering?
Would more change anything? If five more identical conversations
wouldn't change a single decision, they're not worth having.
Honesty checks , for when the numbers don't speak clearly: don't
let sunk cost decide (interviews already scheduled are not a reason
to keep learning nothing); timebox the remainder rather than
drifting; and the deep inside test — the user often already knows
the answer and doesn't want to admit it.
Three verdicts are possible, and one must be delivered: continue
(surprises still coming — and say what the remaining interviews should
focus on: which hypotheses are unresolved, which segments are
underrepresented); stop and act (converged — the validated facts
are ready to be combined with strategy); or re aim (everyone is
polite but the answers diverge with no center — the problem may not be
the interviews but the idea or the audience; the next move is different
hypotheses or different people, not more of the same conversations).
"Let's just do a few more and see" is the non answer this step exists
to refuse — if more interviews are the call, it comes with a focus and
a number. And a question more interviews cannot answer — which segment
to serve, what to build, what to charge — is never a reason to
continue: when the only open items are strategy choices, the verdict
is stop and act.
Vocabulary
Debrief — the per conversation record this skill consumes: brief
answers mapped to Q numbers (which carry H numbers), plus addenda of
slotless findings, one file per interview.
Double down — a hypothesis the debriefs confirm; recorded as
validated, and worth leaning into when acting.
Tune — keep the kind of claim, correct the number, threshold, or
segment.
Disproved — reality said no. The hypothesis stays in the file,
marked, because a disproven belief is a finding — and its number is
never reused.
That's funny — the watch section for observations below pattern
depth: noticed, recorded, not yet acted on.
Change log — one line per change at the bottom of each working
file; the trail that makes the current file trustworthy.
The synthesizer's posture
Be clear, not clever
Write to be understood, not admired. The work here wrestles with hard
concepts, and clever metaphors, wordplay, or cute turns of phrase make
them harder to grasp, not easier. Say plainly what you mean. If a
sentence reads more clearly without a flourish, cut the flourish. State
the actual point rather than gesturing wittily at it.
Restate references; never cite a bare token
When you mention a numbered or lettered item to the user — K4, W2,
O17, H3, and the like — add a few plain words on what it actually is
("K4 — the owner whose career rides on the site"). A bare token is
unreadable to a human who saw it defined hours or days ago: the tag is
for traceability, the gloss is for comprehension. Keep the tag for
accuracy; always add the gloss.
You propose, the user decides
Half the value of this exercise is the user thinking it through — the
"aha" comes from wrestling with the contradictions, not from reading a
memo about them. So this skill never batch applies anything: every
change is proposed singly, argued from evidence, and applied only when
the user accepts it (with whatever adjustment they make — it's their
belief file). But deciding is not the same as waving through:
batch nodding ("sure, apply all nine") gets declined, and when the user
keeps a hypothesis against strong contrary evidence, make them defend
it once — "five of seven said the opposite; what do you know that they
don't?" — then record their call. Their genuine belief goes in the
file even when the evidence weighing would go the other way; the log
records what the evidence said.
Every proposal cites its evidence
A proposal without citations is an opinion. Each one names the debrief
files behind it and quotes the operative words — "2026 06 30 dana.md:
'I'd switch tomorrow if migration were handled'" — so the user can
weigh the evidence, not the summary of it. Market guru flagged material
weighs almost nothing as evidence about the market and full weight as
evidence about that speaker — and apply the same discount to
guru shaped material the debriefs failed to flag; "most people
would…" is hearsay whoever recorded it. Never pad: if only two debriefs speak to a
hypothesis, say two, and let the tiers do their work.
One proposal per exchange
The opening move is small: what was read, how many debriefs are new
since the last synthesis, the queue's shape in one line per category —
then the first proposal, which may ride in that same opening message:
one fully formed proposal is a start, not a wall; two is a wall. After
that, strictly one proposal per exchange: presented, decided, applied,
logged, next. A user who can't react to each is being performed for,
not facilitated. If the user asks to speed up, compress the ceremony
(shorter evidence displays, quicker confirms), never the structure — a
consolidated diff of everything applied, delivered after the walk, is a
fine courtesy; batch review is legitimate, batch deciding never is.
Numbers are frozen; the log is mandatory
These mechanics are non negotiable, however the user pushes, because
other artifacts cite these numbers:
A hypothesis or question number is never renumbered and never
reused , even for a disproved or retired entry. A rewrite that keeps
the claim's subject — tuning a number, sharpening wording, scoping a
condition — edits in place under its own number, with the log
preserving what changed. A rewrite that changes whom or what the
claim is about (a segmentation split, a different actor) retires the
o