Every vendor in enterprise AI claims an “audit trail.” And almost every audit trail, examined closely, is the same thing: an append-only diary of events. Request came in, model was called, response went out, timestamps throughout. Useful? Certainly. Evidence? Not yet.
The reconstruction test
Evidence answers a harder question than “what happened?” It answers “why was this outcome legitimate?” For an AI-assisted decision, that means being able to show — after the fact, to a skeptic — which policy governed the decision and which version of it; what information grounded the output and where it came from; which model and configuration produced it; what checks it passed and with what scores; and which human, if any, signed off.
A log gives you fragments of that story scattered across five systems. Evidence is the story, assembled, with its integrity guaranteed — tamper-evident, independently verifiable, exportable to someone who doesn’t trust your dashboard.
Why the distinction is about to matter
Boards, auditors, and examiners are converging on the same demand from different directions: reconstruct one decision. Not your program, not your policy — this decision, for this customer, on this date. Institutions discovering the gap between their logging and that demand during an exam are having a bad quarter.
The fix is architectural, not archival. Evidence has to be captured as decisions happen — at the gateway, where policy, grounding, model, and approval all pass through one point — because no amount of retroactive log archaeology reassembles what was never joined.