design-testing-strategy

Use before writing any type of tests. Distills 14 industry sources into deterministic decision gates, schemas, and worked test examples.

By neolabhq · 516 installs

npx skills add neolabhq/context-engineering-kit --skill design-testing-strategy

Source repository · Upstream listing

Design Testing Strategy A reference manual for designing a fit for purpose, fit for criticality testing strategy. This skill is decision oriented , not philosophical: every gate is deterministic (ON when X / OFF when Y), every schema is enforced (field ordering matters), every example is worked end to end. How To Use This Skill 1. Read Decision Gates in order (Gate 0 Gate 6). Each gate is independent — you may finish with any subset of test types ON. 2. Apply Strategic Skip Heuristics to remove ON gates that would yield low ROI for this artifact. 3. For each ON gate, fill the Test Matrix Schema ( selected types entry) — the field order is load bearing. 4. List rejected types in rejected types and deliberate skips in deliberately skipped . 5. Produce a Test Cases to Cover markdown bullet list using ISTQB techniques from Case Design Techniques . 6. Cross check against the matching Worked Example (A pure function / B HTTP+DB endpoint / C UI component). Decision Gates Apply gates in numeric order. Each gate produces an independent boolean ( applies: true false ). Gates do NOT veto each other — a single artifact may have unit + integration + contract + property based all ON. Type ON when OFF when Source 0 Skip All Criticality is NONE (docs only, comments, formatting, generated code, config without logic, throwaway prototypes) Anything with branching, computed output, side effects, or user visible behavior [Pragmatic Programmer](https://pragprog.com/titles/tpp20/the pragmatic programmer 20th anniversary edition/) — "Test ruthlessly and effectively" implies effective skipping when ROI is zero 1 Unit Code contains any logic: branches, loops, conditionals, computation, transformation, parsing, validation, formatting Pure declarative wiring (DI registration, route table) with no behavior [Test Pyramid (Vocke)](https://martinfowler.com/articles/practical test pyramid.html) base layer + [Beck TDD](https://www.oreilly.com/library/view/test driven development/0321146530/) Red Green Refactor unit 2 Integration Boundary crossing: HTTP call, DB query, external SDK, message queue, filesystem I/O, OR collaboration with =2 distinct collaborators where unit doubles distort behavior Pure function with no I/O and 0 1 stable collaborators [Testing Trophy (Dodds)](https://kentcdodds.com/blog/the testing trophy and testing classifications) — integration is the highest ROI layer; [Google "Follow the User"](https://testing.googleblog.com/2020/10/testing on toilet testing ui logic.html) 3 Component or E2E UI surface AND criticality = MEDIUM HIGH AND user facing critical path (signup, checkout, auth, payment, primary CTA) Internal admin only screens, dev tooling, or non critical UI [Test Pyramid top](https://martinfowler.com/articles/practical test pyramid.html) + [ISO/IEC/IEEE 29119](https://en.wikipedia.org/wiki/ISO/IEC 29119) risk ranking + [Google e2e principles](https://testing.googleblog.com/2016/09/testing on toilet what makes good end.html) 4 Contract Public API consumed by =1 distinct clients (mobile + web, multiple internal services, external partners) AND independent deploy cadence API where consumer and provider deploy together [Pact / CDC](https://docs.pact.io/) + [Pactflow CDC explainer](https://pactflow.io/what is consumer driven contract testing/) 5 Smoke Deployable surface (web app, API, service) AND a deploy/CI pipeline exists where post deploy validation is meaningful Library, internal helper, or no deploy pipeline [Google "What Makes a Good End to End Test"](https://testing.googleblog.com/2016/09/testing on toilet what makes good end.html) — smoke = minimal e2e for deploy gate 6 Property Based Input domain is large or unbounded (numeric ranges, strings, lists, parsers, serializers, encoders, math) AND invariants are stable (round trip, idempotency, monotonicity, commutativity) AND criticality = MEDIUM HIGH Small finite input domain, unstable invariants, or LOW criticality [Hypothesis / QuickCheck](https://hypothesis.works/articles/what is property based testing/) Gate Application Algorithm Criticality Scale (used by Gates 3 and 6): Level Definition NONE Docs, formatting, generated code, throwaway code, configs without logic LOW Internal dev tooling, admin only screens, logging formatters MEDIUM Standard CRUD, internal APIs with a single team consumer, non critical UI, helpers and utilities MEDIUM HIGH User facing UI on critical paths, public APIs with multiple consumers, business workflows HIGH Money movement, auth/authz decisions, security critical validation, data integrity, regulated domains Test Type Reference Type Use when Do NOT use when Frameworks Typical dependencies Google Size unit Pure logic, single function/method/class, deterministic inputs Code is just I/O orchestration with no logic vitest, jest, pytest, go test, JUnit, xUnit, RSpec None (or in memory fakes) [Small](https://testing.googleblog.com/2010/12/test sizes.html) integration Boundary crossing (DB, HTTP, queue, FS); multiple collaborators where mocking distorts behavior Pure function with no boundary vitest, jest, pytest, go test, JUnit + [Testcontainers](https://testcontainers.com/), supertest, TestRestTemplate Real Postgres/Redis/Kafka via Testcontainers, in process HTTP server, real FS in tmpdir [Medium](https://testing.googleblog.com/2010/12/test sizes.html) (single machine, localhost OK) component UI rendering + interaction within a single component, no full app context Backend only logic; multi page user flow React Testing Library, Vue Test Utils, Angular TestBed, Storybook interaction tests jsdom or happy dom, mocked network at fetch/axios level Small to Medium e2e Full user path through running app: real browser, real backend, real DB Internal helper, single component, non critical UI [Playwright](https://playwright.dev/), [Cypress](https://www.cypress.io/), Selenium Real running app + Testcontainers backed DB or seeded staging [Large](https://abseil.io/resources/swe book/html/ch11.html) (multi process, possibly multi machine) smoke Post deploy go/no go: hit / health, key endpoints respond, login works Detailed correctness; smoke is shallow by design Playwright (1 3 critical paths), HTTP probe scripts, k6 minimal scenarios Real deployed environment Large contract Public API consumed by 2+ distinct clients with independent deploy cadence Single consumer internal API; provider and consumer deploy together [Pact](https://docs.pact.io/), Spring Cloud Contract, OpenAPI schema validators Pact broker or contract files in repo Medium property based Large/unbounded input domain with stable invariants (parser, serializer, encoder, math) Small finite input space; unstable invariants [Hypothesis](https://hypothesis.works/) (Python), fast check (TS), QuickCheck (Haskell), jqwik (Java), proptest (Rust) Same as unit Small Google Test Size Mapping [Google Test Sizes (Bland)](https://mike bland.com/2011/11/01/small medium large.html) and [SWE at Google Ch.11](https://abseil.io/resources/swe book/html/ch11.html) classify tests by resources (size), independent of scope (paths covered): Size Process model Network Filesystem Time budget Notes small Single process, single thread None None (in memory only) < 100ms Fast, hermetic, parallelizable medium Single machine, multiple processes allowed localhost only tmpdir allowed < 1s Testcontainers fits here large Multi machine External network allowed Persistent FS allowed < 15min Full e2e enormous Distributed Wide network Anywhere longer Cluster / chaos A test's type (unit/integration/e2e) and size (small/medium/large) are orthogonal: a small integration test (Testcontainers Postgres in same process via JDBC) is legitimate. Playwright vs Cypress (UI e2e) Dimension [Playwright](https://playwright.dev/) [Cypress](https://www.cypress.io/) Browsers Chromium, Firefox, WebKit Chromium, Firefox, WebKit (limited) Multi tab / multi origin Yes Limited Parallelism Built in shards Paid dashboard or external Network interception Robust route level cy.intercept Default Choose Playwright for new projects unless team already standardized on Cypress Choose Cypress when team has heavy investment Case Design Techniques Use ISTQB Foundation Level black box techniques to derive what to test inside each chosen test type. References: [ISTQB BVA white paper](https://istqb.org/wp content/uploads/2025/10/Boundary Value Analysis white paper.pdf), [ASTQB black box techniques](https://astqb.org/4 2 black box test techniques/). 1. Equivalence Partitioning (EP) Divide input domain into partitions where the system is expected to behave the same way; ONE test per partition is sufficient. Worked example — discount(orderTotal: number) number : Partition Range Representative test input Expected Below threshold 0 <= total < 100 50 0% discount Mid tier 100 <= total < 500 250 5% discount Top tier total = 500 1000 10% discount Invalid (negative) total < 0 1 throw / error Four tests cover all partitions. EP alone misses boundaries — combine with BVA. 2. Boundary Value Analysis (BVA) Bugs cluster at boundaries. For every boundary value B , test B 1 , B , B+1 (or for floats, the smallest representable step). Worked example — same discount function, boundary at 100 : Test input Why Expected 99 (= B 1) Last value of "below threshold" partition 0% discount 100 (= B) First value of "mid tier" partition 5% discount 101 (= B+1) Confirms not off by two 5% discount Repeat for boundary at 500 : test 499 , 500 , 501 . Total: 6 boundary tests + 4 EP tests = 10 cases. The B 1 / B / B+1 triplet has the same shape across boundaries (vary input, vary expected output, identical assertion); this is a natural fit for a table driven test (see sub section 5 below). 3. Decision Tables When output depends on combinations of conditions. Each column is a rule. Worked example — canCheckout(cartHasItems, paymentValid, addressOnFile) : Condition / Rule R1 R2 R3 R4 cartHasItems T T T F paymentValid T T F addressOnFile T F Result allow block:address block:payment block:cart Four tests, one per rule ( = don't care, dropped via merging). 4. State Transition When behavior depends on history. Identify states, events, and forbidden transitions. Worked example — Order state machine with states {draft, submitted, paid, shipped, cancelled} : From Event To Test draft submit submitted happy path submitted pay paid happy path paid ship shipped happy path draft cancel cancelled early cancel paid cancel reject forbidden — refund flow required, NOT direct cancel shipped submit reject forbidden Cover one test per legal transition + one per forbidden transition (negative path). 5. Table Driven Tests When EP, BVA, or decision table analysis yields 3+ cases with the same shape (same setup, same assertion, only inputs and expected outputs differ — e.g., parsing valid/invalid date formats; computing tax across brackets; routing rules) collapse them into a single table driven test . The cases become rows in a data table; the test body iterates the rows and runs one assertion per row. References: Dave Cheney, [Prefer table driven tests](https://dave.cheney.net/2019/05/07/prefer table driven tests); [Go wiki: TableDrivenTests](https://go.dev/wiki/T