AI made your team faster.
Did it make your evidence real?
Consistency is not correctness. Your documents agree with each other; do they agree with your system? Below: a condensed sample of the audit readout, produced on a synthetic repository for a fictional client, exactly in the format a real sponsor receives on day 15.
Verdict: half the leverage is real. The other half is signing your name.
"An outside answer key spent 15 days read-only in our repos and came back with six findings our own AI rated green, a ledger of which AI spend is real and which is theater, and a ranked plan that keeps the gains and stops the liabilities."
Fictional client: Halden Software Systems, a mid-size supplier software organization, 180 engineers; AI assistants write about half its new code, most unit tests and nearly all documentation. Audit window 2026-07-06 to 2026-07-24. All figures simulated and labeled. Market comparisons, where present, carry benchmark date 2026-08.
The slop findings: what behavior said that the paper did not
31 of 74 generated cases call the code and assert only that it ran: literal ASSERT_TRUE(1) bodies, TODO markers where expected values should be. A green pipeline stops meaning anything, because regressions ship under a passing badge.
The generated matrix reports 100% of tests traced; two cited requirement IDs resolve to nothing. The traceability evidence dissolves on the first thread an assessor pulls.
The manual gained a revision, a completed review status and an implementation claim; no commit in the window touches the code it describes. Documentation that moves when the system does not turns timestamps into noise instead of evidence.
The remaining three findings, the complete ledger evidence and the interactive Leverage Lens (the synthetic repo with all six probes, runnable) open in the first meeting.
Where the AI spend lands: real, partial, theater
| AI-usage area | Measured leverage | Evidence pointer |
|---|---|---|
| Refactors & scaffolding | REAL | measured on the commit ledger · detail in the full edition |
| Code review | PARTIAL | real catches, one expensive miss · detail in the full edition |
| Test generation | THEATER | 31 of 74 cases assert nothing (F-01); the coverage gate counts them anyway |
| Documentation | THEATER | phantom requirement IDs (F-02); timestamps moving without code (F-03) |
One lane earns its keep and should scale. One needs a guardrail. Two produce artifacts that look like engineering evidence and are not. The audit's two deliverables in one: this ledger with measured evidence, and the ROI-ranked roadmap that keeps the real half.
Consistent without corresponding
Halden's artifacts agree with each other almost perfectly, because the same models wrote both sides of every check. Cross-checking documents against documents found nearly nothing. Checking artifacts against behavior found six. That is the finding class an internal review structurally cannot produce: it grades homework with the answer key that wrote it. The more AI your organization uses, the more this audit is for you.
Why this survives your hardest questions
No assessor certification is held or claimed; standards and process language is engineering interpretation. You already buy referees. This is the coach who sat on your side of the table for twenty years.
Probes run read-only against repos and pipelines. This very page is the proof of format: one readable HTML file, zero network calls, nothing to install, nothing leaves. Your security team can read every line.
Your AI grades with the answer key it wrote; your QA reports inside the org it audits. An outside answer key cannot be in-sourced. That is the product.
Founding clients: your engineering lead picks the repos, the probes get two hours. No finding your own team agrees is real, no fee.
AI Leverage Audit: fixed fee, fixed clock, read-only
From US$7,500 · 15 business days · read-only repo and pipeline probes · two deliverables in one: the ROI-ranked map of where AI genuinely pays, and the slop findings with evidence pointers. The Executive AI Evidence Session (US$1,500 to 3,000) credits 100% against this audit. Exit paths: a fixed-scope build that closes the top finding, or the Fractional AI Advisor, the outside answer key on retainer, exits monthly.
This page is that device. Bring your own AI usage map; the first disagreement is usually visible in fifteen minutes.