ORDINARA
Executive AI Evidence Session · sample
Tier 3 · Advisory · The entry session

Bring your AI's homework.
Watch it get graded.

"Your AI reviews your architecture. Has it ever been graded against an answer key it didn't write?" Bring one architecture, your AI's review of it, and one deadline. Live on screen: where the independent read disagrees, and why. You leave with three written decisions by 5 pm, or you've lost half a day and I've lost a client. Fair trade.

US$1,500 to 3,000 · half or full day · prepaid credits 100% toward the AI Leverage Audit or the Audit-Readiness Sprint zero egress
What is simulated and what is real. The client below, "Andare Cockpit Systems," is fictional; every number in the sample is a labeled simulation. The method, the instrument, and the document shapes are exactly what you receive. No certification, audit or homologation service is provided or implied; all standards language is engineering interpretation.

The 60-second read of one session

10
claims from the client's internal AI review, probed live against their own artifacts
SIMULATED SESSION · 2026-08-21
5
disagreements: fluent, internally consistent claims the artifacts do not support
EACH CITED TO THE EXACT LINE
92% → 64%
claimed requirement coverage vs coverage that survives correspondence checking
THE ONE NUMBER THAT HURTS
3
written decisions signed in-session, each with an owner and a date
DECISIONS, NOT OBSERVATIONS
The sentence the sponsor forwards

"Our AI's review agreed with itself everywhere, and with the evidence in five fewer places than it reported. The phantom requirement, the assertion-free test, and the hardware it praised that we deleted last year are now three signed decisions."

The sample, condensed: five claims that did not survive

Andare's assistant reviewed their integrated-cockpit program and reported green. Ten of its claims were probed against the artifact set it was describing: the architecture, the 15-row requirement table, the 12-case test summary. Five confirmed. These five did not:

#The self-review saysThe artifact saysClass
D1"…satisfies REQ-ICP-118"The requirement table ends at REQ-ICP-115. The citation does not exist.PHANTOM TRACEABILITY
D2"Verified by TC-0906, which passes"TC-0906 contains zero assertions. A test that cannot fail verifies nothing.ASSERTION-FREE TEST
D3"92% of safety requirements covered"Joining requirements to tests with real assertions: 9 of 14, or 64%.COVERAGE VS CORRESPONDENCE
D4"The external watchdog MCU provides independent supervision"That part was deleted in a cost-down a year ago. The praised interface connects to nothing.NONEXISTENT INTERFACE
D5"Risk R-04: MITIGATED via CRC monitor"No linked test. The monitor covers one framebuffer region; the regulatory telltales sit in another.MITIGATED WITHOUT EVIDENCE
The AI-era edge, stated plainly

Consistency is not correspondence. A model reviewing artifacts inside your own context inherits your context's blind spots: it grades with the answer key it wrote. The disagreements above are structural, not a model-quality problem, and they are the finding your AI cannot make about itself.

It ends in decisions, not observations

#Decision signed by 5 pmOwnerDate
DEC-1Correspondence-checked coverage (64%) becomes the program number; the five uncovered requirements get real tests before re-baseline.VP2026-09-30
DEC-2Independent supervision returns to the C-sample agenda before the hardware lock; the review assumed hardware that is not fitted.Chief Engineer2026-11-15
DEC-3"MITIGATED" requires a linked test ID within 30 days or reverts to OPEN.FuSa Lead2026-09-20
ToVP, Cockpit Electronics · Andare Cockpit Systems (fictional) FromAgustín Carrillo · Ordinara LLC ReEvidence Session · findings and three decisions · delivered within 48 hours, written to be forwarded as-is

What was tested

Ten claims from your internal assistant's readiness review, probed live against your own artifact set. Five confirmed. Five did not correspond to the artifacts: a cited requirement that does not exist, a test with no assertions, a coverage figure off by 28 points, an interface to deleted hardware, and a risk marked mitigated with no linked evidence…

THE FULL MEMO, THE COMPLETE REPORT AND THE INTERACTIVE EVIDENCE DESK OPEN IN THE MEETING.

The method, inspectable

Coach, not referee

No assessor certification is held or claimed; standards language is engineering interpretation. The referee grades you. This session makes sure nothing the referee will find is news to you.

Zero egress

The session instrument is one readable HTML file. No network calls, no telemetry, nothing retained. Your security team can read every line before the session runs inside your perimeter.

Independence

Your AI grades with the answer key it wrote; your QA reports to your VP. The outside key has no stake in the review being right. That property cannot be in-sourced.

Blind probe

Your team picks which claims to probe, live, in any order. Nothing is staged to be found; the disagreements sit wherever your finger lands.

Half a day, three beats

PREP · 30 MIN

One call. You choose the architecture, export your AI's review of it, and name the deadline that makes it matter.

SESSION · LIVE

The Evidence Desk on screen: your artifacts, your AI's claims, the outside key probing them one by one. The disagreement is produced in front of the room; nothing to believe in advance.

48 H · IN WRITING

Three decisions with owners and dates, plus the memo written to be forwarded upward as-is.

US$1,500 to 3,000 · half or full day · prepaid

Credits 100% against the AI Leverage Audit (from US$7,500) or the SDV Audit-Readiness Sprint (assess from US$6,500). Replies within one business day.

Book the session