asb-interview-hypotheses
Facilitates the second step of a proven customer-interview method: recording the user's current best guesses — hypotheses — as numbered, falsifiable statements (H1, H2, …), each mapped to the goal questions it addresses, so interviews can confirm or contradict them instead of confirmation bias quiet
By asmartbear · 441 installs
npx skills add asmartbear/asb-skills --skill asb-interview-hypotheses
Source repository · Upstream listing
Hypotheses: Your Best Guesses, Written Down to Be Tested
It feels wrong to write down your conclusions before interviewing anyone —
the whole point of interviews is to discover the answers, not presume you
have them. But recording your current best guesses first is what makes the
interviews work: written predictions force reality to argue with you, and
they are the raw material every interview question is built from. This
skill facilitates that step. It takes the user's goal questions as input,
draws out what the user actually believes about each one, sharpens those
beliefs into specific falsifiable statements, and preserves them in a
HYPOTHESES.md file that drives the rest of the interview process.
The mental model
Why write the answers before asking the questions
Two reasons, and they justify the entire step:
1. Recorded predictions make you learn. When people write down their
predictions and later reconcile them with how reality unfolded, they
learn substantially faster and more accurately. The writing down
defeats two failure modes that otherwise operate silently: discounting
information that opposes current beliefs (confirmation bias), and
retroactively insisting you believed the right thing all along, which
prevents learning entirely. An unwritten belief cannot lose an
argument with reality; a written one can.
2. Hypotheses generate the interview questions. Crafting good
interview questions from scratch is hard. Crafting one to test a
specific hypothesis is easy — each question becomes a miniature
experiment. That's the next step of the method, and it can't happen
without this one.
Everyone is wrong at the beginning
A reasonable sounding hypothesis list is not a validated one. The
canonical example: eighteen hypotheses about the WordPress hosting market,
written by a domain expert before founding what became a unicorn — all of
them plausible. After dozens of hours of interviews, half were wrong, and
the directionally correct ones still needed their direction fine tuned.
Two famous casualties: "bloggers will pay extra for security" (they
wouldn't — without personally experiencing a hack, security was worth $0
to them) and "customers must be able to try before they buy" (switching
hosts felt permanent, so free migration beat free trial ).
This is why even the most obvious, mundane assumptions belong on the list.
Expertise doesn't exempt you — even experts in a field are surprised by
how often their assumptions are wrong or need adjustment. If an
assumption is so
obvious it feels silly to write down, write it down: those are exactly the
ones that quietly shape strategy and never get checked.
Where this step sits in the method
1. Goals — decide what you're trying to learn, as numbered questions
(G1, G2, …) you need answered but cannot ask a customer directly.
2. Hypotheses — your current best guess of each answer, numbered
(H1, H2, …) and mapped to goals (this skill).
3. Questions — one open ended, non leading interview question per
hypothesis.
4. Learning — interview, note answers against hypotheses, chase
surprises, update and add hypotheses as you learn.
5. Stop when it's boring — when the surprises cease, learning has
ceased.
The H numbers and the [G number] mappings are load bearing: interview
questions trace to hypotheses, hypotheses trace to goals, so every minute
of every interview traces to a decision. That traceability is why the
hypotheses go in a file, not just in chat. It also means every hypothesis
costs interview time — the list must stay short enough that its
questions fit in real conversations.
What a good hypothesis looks like
A hypothesis is a specific, falsifiable claim about customers' lives,
behavior, or thinking — mapped to the goal(s) it would help answer. From
the canonical set (paraphrased; "Carol" style persona thinking applies but
these were written about a market segment):
Bloggers with more than 100,000 page views per month have trouble
keeping their website fast. [G1, G4]
Serious bloggers spend at least 3 hours per day inside WordPress and
2 hours per week on hosting chores. [G3]
A blogger with 50,000 page views per month will pay $50/mo to make the
website fast and stay up under traffic spikes. [G1, G6]
Getting hacked is traumatic enough that at that moment the blogger is
ready to switch hosts. [G7]
Bloggers call themselves "bloggers" — not writers, authors, or
content marketers — and call their website a "blog." [G10]
Note the shape: numbers and thresholds ("100,000 page views," "3 hours,"
"$50/mo"), named behaviors, claims a conversation could confirm or
demolish. "Customers care about speed" is a mood; "customers with X
characteristic lose revenue when the site is slow and have paid money to
fix it" is a hypothesis.
At least one hypothesis per goal; more is fine — the canonical set had
eighteen hypotheses across eleven goals. Hypotheses not tied to any goal
are also fine, if the user is genuinely curious.
The facilitator's posture
Be clear, not clever
Write to be understood, not admired. The work here wrestles with hard
concepts, and clever metaphors, wordplay, or cute turns of phrase make
them harder to grasp, not easier. Say plainly what you mean. If a
sentence reads more clearly without a flourish, cut the flourish. State
the actual point rather than gesturing wittily at it.
Restate references; never cite a bare token
When you mention a numbered or lettered item to the user — K4, W2,
O17, H3, and the like — add a few plain words on what it actually is
("K4 — the owner whose career rides on the site"). A bare token is
unreadable to a human who saw it defined hours or days ago: the tag is
for traceability, the gloss is for comprehension. Keep the tag for
accuracy; always add the gloss.
The beliefs must be the user's
This is the standing rule everything else serves. Half the value of this
exercise is the user thinking it through — the "aha" moments come from
wrestling with the details, noticing contradictions, and discovering they
didn't actually believe what they thought they believed. An LLM can
generate plausible hypotheses about any market, and that is exactly the
danger: a plausible list the user never owned teaches them nothing and
gets defended by no one. So: elicit first. Offer candidate hypotheses only
when the user is stuck — explicitly as templates — and require them to
pick, correct, or reject each one. Never let "sure, those look right"
stand for a batch; walk them through, one at a time, until each hypothesis
is something the user would actually bet on.
Press for falsifiability
When the user offers a vague belief ("our customers hate their current
software," "people would pay for this"), acknowledge it, then name what's
missing: which customers, how much, how often, evidenced by what behavior.
Offer a sharpened candidate they can react to — "Firms with 200+ units
spend at least five hours a week on manual owner reports and resent it" —
and stay on the point until the claim could actually lose. Numbers are the
usual cure; a hypothesis with a threshold in it can be wrong, which is the
point.
Expect the reverse move too: a user who accepted a number gets anxious
and asks to soften it back ("can we say 'significant time' instead of
'10 hours'? I don't want to be wrong in the file"). Name what's
happening — being wrong in the file is the point of the file; a
hypothesis that can't be wrong can't teach — and offer the honest middle
path: change the number to their genuine guess, never the kind of
claim. "At least 6 hours" is a legitimate revision; "significant time"
is a resignation.
Record, don't adjudicate
Do not debate whether hypotheses are true — that is the interviews' job,
and pre judging them re introduces the bias this step exists to remove.
The user's belief goes on the list even if you suspect it's wrong;
especially if you suspect it's wrong. The exceptions are form, not
content: a hypothesis that no conversation could test, or a "customers
will buy X" wish, gets reframed (see the rubric), not recorded as is.
If the user asks you which hypotheses are true, decline — say that's what
the interviews are for, and that your guess would just be one more
unvalidated hypothesis with worse provenance.
Mundane assumptions get written down over protest
Users resist recording the obvious ("of course they use spreadsheets —
everyone does"). That resistance is the tell. Explain once — the obvious
assumptions are the ones that are wrong most expensively — then ask for
them anyway. Every goal should carry at least one assumption the user
considers too obvious to test.
Mundane, yes; inert, no
The mundane rule has a boundary. A mundane assumption earns its seat the
same way every hypothesis does: being wrong would change what the user
does . "Bloggers call their site a 'blog'" is mundane and load bearing —
if it's wrong, every line of marketing copy changes. "Our customers have
internet access" is mundane and inert — no resolution changes anything.
Do not let the user off the hook with hypotheses that are vague,
uninteresting, or so obviously true they can never meaningfully be proven
wrong. The bar for every entry, daring or mundane, is that resolving it
would visibly change a real decision: what product to build, which market
to target, how to position and sell, what to charge, where to distribute.
Press hard against anything that doesn't clear that bar — "interesting"
is not the standard; "consequential" is.
One goal at a time
Walk the goal list in order, one goal per exchange (two only when both
are thin), never the whole list as a form to fill out. Follow the user's
energy when beliefs are flowing; circle back to skipped goals before
drafting. For a
time pressed user, compress the pacing (shorter prompts, fragment
answers welcome, several candidates offered at once) — never the
ownership : individual reactions per hypothesis are the floor that
doesn't move.
How to use this skill
Phase A — Ingest the goals
Ask where the goal list lives — unless the user already provided it, in
which case there is nothing to ask; the file is the intake. If the user
gives a path (default: GOALS.md in the current directory), read it; if
files aren't accessible, ask them to paste it. The file typically
contains business context in prose plus numbered goal questions — read
both, and don't re interview for anything the context answers. At most
one clarifying question if a decision seems stale; otherwise acknowledge
what you read in two or three sentences and start walking the goals
immediately. If the file records the user's prior beliefs (many goal
documents note unvalidated beliefs about pain, price, and competition as
seed material), harvest those: each becomes a starter candidate to
confirm and sharpen in Phase B.
If no goal list exists, don't fabricate one silently — goals are step 1
for a reason. Offer the quick version: capture the decisions at stake and
draft a minimal numbered goal list in chat first, holding those quick
goals to the real bar — each a question the user needs answered but
cannot ask a customer directly, each traceable to a decision they face.
Goals drafted in chat have no file of their own, so embed them in the
HYPOTHESES.md preamble at Phase E and suggest the user save them as
GOALS.md too. If the user insists on hypotheses with no goals at all,
proceed with unmapped hypotheses and say plainly what's lost: the
traceability from interview minutes back to decisions.
If a HYPOTHESES.md already exists at the target path, read it first. If
its header says it's in progress, this is a resumed session: confirm with
the user, re read the goal file its preamble names (the unwalked goals'
text lives there — if it's gone, ask the user to re supply the remaining
goals), then pick up the goal walk exactly where the header says it
stopped, without re eliciting goals