Practice 01

QA for AI

Testing an AI feature is not testing. There is no deterministic pass, so confidence has to be expressed as a distribution across repeated runs. Someone has to produce it. We do, and we hand you the evidence pack rather than an opinion.

  • Readiness
    Where the system stands against the behavior you say it must hold to.
  • Validation
    Testing against those bounds at scale, with quantified coverage and named gaps.
  • Assurance ops
    Continuous re-grading after every model, prompt, or data change.

What a confidence distribution looks like

agreed boundbehavior score across 2,000 repeated runs →94.5% within bounds · the tail is where incidents live

Illustrative. Your distribution comes from your own runs. A single pass/fail number would report this system as working. The shape is the finding.

Failure taxonomy

What Actually Goes Wrong in Production.

Five patterns account for most of the AI incidents we are called in after. A generic evaluation suite catches roughly one of them.

  • Silent Drift A model or prompt change shifts behavior on a subset of cases nobody re-tested. Suites still pass. Different system.
  • Confident Wrong Answers The system answers in-scope questions correctly and out-of-scope questions with exactly the same confidence.
  • Phantom Actions The agent reports completing an action it never completed. Common in scheduling, booking, ticketing, and claims flows.
  • Untested Conditions Accented speech, background noise, adversarial phrasing, and very long context. Fine in the demo, absent from the test set.
  • Boundary Erosion Multi-turn pressure walks the agent outside the limits it held on turn one, one small concession at a time.
Boundaries

What We Will Not Claim.

This matters more here than anywhere else, and vagueness is how vendors in this category get people hurt.

  • We do not certify anything, and we are not a certification body.
  • We do not assert compliance with any regulation on your behalf.
  • We do not guarantee safety, correctness or production outcomes.
  • We do not assume your production risk.
  • We produce evidence under tested conditions with known gaps stated. You make the release decision.
Tier 1 · No Cost

Verification Gap Diagnostic

Seven questions, two minutes. A scored read on where your gap is widest, an estimate of your unverified merge ratio, and what to fix first, on screen before we ask who you are.

Free · Instant · No Call
Tier 2 · Fixed Fee

Three-week Readiness Assessment

We measure change volume against verified change, find where coverage claims and production reality diverge, and hand you a costed 60–90 day plan you can take to a board.

Fixed Scope · Credited Against the First Three Months
Tier 3 · Talk First

Thirty Minutes With a QA Lead

An engineering lead is in the room, not just a salesperson. Bring the number you have to hit and we will tell you whether it is reachable, including when it is not.

No Pitch · We Will Say if It Is Not a Fit