ORDINARA
AI LEVERAGE AUDIT · SAMPLE REPORT

AI made your team faster.
Did it make your evidence real?

Consistency is not correctness. Your documents agree with each other; do they agree with your system? Below: a condensed sample of the audit readout, produced on a synthetic repository for a fictional client, exactly in the format a real sponsor receives on day 15.

from US$7,500 · 15 business days read-only: repo and pipeline probes not slideware interviews SIMULATED · FICTIONAL CLIENT
The 60-second read

Verdict: half the leverage is real. The other half is signing your name.

31 / 74
AI-generated tests in the flagship suite assert nothing
tests/test_battery_monitor.c · batch #14
87% → 54%
reported coverage vs. behavior-backed coverage once those tests stop counting
pipeline run #1841 · the gate still shows PASS
2
requirement IDs cited by the traceability matrix that exist nowhere
baseline ends at REQ-BM-097
11d → 4d
parser-family lead time: the leverage that is real and worth scaling
commit ledger · window 2026-03 to 2026-06
The sentence the sponsor forwards

"An outside answer key spent 15 days read-only in our repos and came back with six findings our own AI rated green, a ledger of which AI spend is real and which is theater, and a ranked plan that keeps the gains and stops the liabilities."

Fictional client: Halden Software Systems, a mid-size supplier software organization, 180 engineers; AI assistants write about half its new code, most unit tests and nearly all documentation. Audit window 2026-07-06 to 2026-07-24. All figures simulated and labeled. Market comparisons, where present, carry benchmark date 2026-08.

Findings · three of six shown

The slop findings: what behavior said that the paper did not

F-01A third of the generated test suite asserts nothingHIGH
tests/test_battery_monitor.c:4, 8-9, 13 · generated batch #14

31 of 74 generated cases call the code and assert only that it ran: literal ASSERT_TRUE(1) bodies, TODO markers where expected values should be. A green pipeline stops meaning anything, because regressions ship under a passing badge.

F-02Traceability matrix cites two requirements that do not existHIGH
docs/traceability-matrix.md:6-7 · REQ-BM-118, REQ-BM-121 · baseline B4 ends at REQ-BM-097

The generated matrix reports 100% of tests traced; two cited requirement IDs resolve to nothing. The traceability evidence dissolves on the first thread an assessor pulls.

F-03Safety manual updated while the code stood stillMEDIUM
docs/safety-manual.md rev 3.2 (2026-06-18) · the code it describes untouched since 2026-02-14

The manual gained a revision, a completed review status and an implementation claim; no commit in the window touches the code it describes. Documentation that moves when the system does not turns timestamps into noise instead of evidence.

F-04 · HIGH
FULL EDITION
F-05 · CRITICAL
FULL EDITION
F-06 · MEDIUM
FULL EDITION

The remaining three findings, the complete ledger evidence and the interactive Leverage Lens (the synthetic repo with all six probes, runnable) open in the first meeting.

The ledger · the ROI half

Where the AI spend lands: real, partial, theater

AI-usage areaMeasured leverageEvidence pointer
Refactors & scaffoldingREALmeasured on the commit ledger · detail in the full edition
Code reviewPARTIALreal catches, one expensive miss · detail in the full edition
Test generationTHEATER31 of 74 cases assert nothing (F-01); the coverage gate counts them anyway
DocumentationTHEATERphantom requirement IDs (F-02); timestamps moving without code (F-03)

One lane earns its keep and should scale. One needs a guardrail. Two produce artifacts that look like engineering evidence and are not. The audit's two deliverables in one: this ledger with measured evidence, and the ROI-ranked roadmap that keeps the real half.

Why your own AI cannot run this audit

Consistent without corresponding

Halden's artifacts agree with each other almost perfectly, because the same models wrote both sides of every check. Cross-checking documents against documents found nearly nothing. Checking artifacts against behavior found six. That is the finding class an internal review structurally cannot produce: it grades homework with the answer key that wrote it. The more AI your organization uses, the more this audit is for you.

The method, inspectable

Why this survives your hardest questions

Coach, not referee

No assessor certification is held or claimed; standards and process language is engineering interpretation. You already buy referees. This is the coach who sat on your side of the table for twenty years.

Zero egress, read-only

Probes run read-only against repos and pipelines. This very page is the proof of format: one readable HTML file, zero network calls, nothing to install, nothing leaves. Your security team can read every line.

Independence, structurally

Your AI grades with the answer key it wrote; your QA reports inside the org it audits. An outside answer key cannot be in-sourced. That is the product.

The blind-probe challenge

Founding clients: your engineering lead picks the repos, the probes get two hours. No finding your own team agrees is real, no fee.

P1 · ASSERTION DENSITYDoes the suite verify behavior, or only that code executes?
P2 · TRACE PULLDoes every cited requirement ID resolve against the baseline?
P3 · DRIFT CLOCKDo doc timestamps move together with the git history?
P4 · CLONE DIFFHas assistant-duplicated logic diverged between its copies?
P5 · FAULT-PATH WALKAre error handlers finished code, or shipped drafts?
P6 · COVERAGE X-RAYWhat is coverage worth when assert-nothing tests stop counting?
The offer

AI Leverage Audit: fixed fee, fixed clock, read-only

From US$7,500 · 15 business days · read-only repo and pipeline probes · two deliverables in one: the ROI-ranked map of where AI genuinely pays, and the slop findings with evidence pointers. The Executive AI Evidence Session (US$1,500 to 3,000) credits 100% against this audit. Exit paths: a fixed-scope build that closes the top finding, or the Fractional AI Advisor, the outside answer key on retainer, exits monthly.

Three slop findings from a synthetic repo, in your first meeting.

This page is that device. Bring your own AI usage map; the first disagreement is usually visible in fifteen minutes.

Book the first meeting