Bring your AI's homework.
Watch it get graded.
"Your AI reviews your architecture. Has it ever been graded against an answer key it didn't write?" Bring one architecture, your AI's review of it, and one deadline. Live on screen: where the independent read disagrees, and why. You leave with three written decisions by 5 pm, or you've lost half a day and I've lost a client. Fair trade.
The 60-second read of one session
"Our AI's review agreed with itself everywhere, and with the evidence in five fewer places than it reported. The phantom requirement, the assertion-free test, and the hardware it praised that we deleted last year are now three signed decisions."
The sample, condensed: five claims that did not survive
Andare's assistant reviewed their integrated-cockpit program and reported green. Ten of its claims were probed against the artifact set it was describing: the architecture, the 15-row requirement table, the 12-case test summary. Five confirmed. These five did not:
| # | The self-review says | The artifact says | Class |
|---|---|---|---|
| D1 | "…satisfies REQ-ICP-118" | The requirement table ends at REQ-ICP-115. The citation does not exist. | PHANTOM TRACEABILITY |
| D2 | "Verified by TC-0906, which passes" | TC-0906 contains zero assertions. A test that cannot fail verifies nothing. | ASSERTION-FREE TEST |
| D3 | "92% of safety requirements covered" | Joining requirements to tests with real assertions: 9 of 14, or 64%. | COVERAGE VS CORRESPONDENCE |
| D4 | "The external watchdog MCU provides independent supervision" | That part was deleted in a cost-down a year ago. The praised interface connects to nothing. | NONEXISTENT INTERFACE |
| D5 | "Risk R-04: MITIGATED via CRC monitor" | No linked test. The monitor covers one framebuffer region; the regulatory telltales sit in another. | MITIGATED WITHOUT EVIDENCE |
Consistency is not correspondence. A model reviewing artifacts inside your own context inherits your context's blind spots: it grades with the answer key it wrote. The disagreements above are structural, not a model-quality problem, and they are the finding your AI cannot make about itself.
It ends in decisions, not observations
| # | Decision signed by 5 pm | Owner | Date |
|---|---|---|---|
| DEC-1 | Correspondence-checked coverage (64%) becomes the program number; the five uncovered requirements get real tests before re-baseline. | VP | 2026-09-30 |
| DEC-2 | Independent supervision returns to the C-sample agenda before the hardware lock; the review assumed hardware that is not fitted. | Chief Engineer | 2026-11-15 |
| DEC-3 | "MITIGATED" requires a linked test ID within 30 days or reverts to OPEN. | FuSa Lead | 2026-09-20 |
What was tested
Ten claims from your internal assistant's readiness review, probed live against your own artifact set. Five confirmed. Five did not correspond to the artifacts: a cited requirement that does not exist, a test with no assertions, a coverage figure off by 28 points, an interface to deleted hardware, and a risk marked mitigated with no linked evidence…
The method, inspectable
No assessor certification is held or claimed; standards language is engineering interpretation. The referee grades you. This session makes sure nothing the referee will find is news to you.
The session instrument is one readable HTML file. No network calls, no telemetry, nothing retained. Your security team can read every line before the session runs inside your perimeter.
Your AI grades with the answer key it wrote; your QA reports to your VP. The outside key has no stake in the review being right. That property cannot be in-sourced.
Your team picks which claims to probe, live, in any order. Nothing is staged to be found; the disagreements sit wherever your finger lands.
Half a day, three beats
PREP · 30 MIN
One call. You choose the architecture, export your AI's review of it, and name the deadline that makes it matter.
SESSION · LIVE
The Evidence Desk on screen: your artifacts, your AI's claims, the outside key probing them one by one. The disagreement is produced in front of the room; nothing to believe in advance.
48 H · IN WRITING
Three decisions with owners and dates, plus the memo written to be forwarded upward as-is.
Credits 100% against the AI Leverage Audit (from US$7,500) or the SDV Audit-Readiness Sprint (assess from US$6,500). Replies within one business day.