The Questions · 19 August 2026

How do you check an AI's diligence work before it goes in an IC memo?

A verification protocol for AI-generated diligence output, and an honest account of which errors the protocol catches and which it does not.

Key takeaways
  • An associate who cannot find a document says so; a language model will sometimes produce a well-formed, entirely fabricated summary carrying no marker that separates it from an accurate one.
  • The four failure modes, in descending order of frequency, are fabrication, misattribution, silent omission and instability.
  • Every number must carry a citation to a page rather than to a document, or it does not go in the memo.
  • Spot-check at a fixed ratio chosen before you see the output, because checking the answers that look wrong catches only the errors that look wrong.
  • A negative control — asking for something you know is not in the room — takes two minutes and shows you the system's failure mode before it matters.
  • Verification does not catch what was never asked, and a well-cited first pass that never examined supplier concentration is a well-sourced blind spot.

The honest starting position is that you check it the way you would check a first-year associate's work, with one adjustment: the associate knows when she is unsure and the model frequently does not.

That adjustment is the whole problem. An associate who cannot find the supplier agreement says she cannot find the supplier agreement. A language model asked the same question will sometimes produce a well-formed, correctly-formatted, entirely fabricated summary of a supplier agreement, and the output carries no marker distinguishing it from the twelve accurate summaries above it. Confidence is not correlated with accuracy in these systems, and reading for tone will not save you.

So the verification cannot be vibes-based. It has to be a protocol.

The four failure modes, in order of how often we see them

Fabrication. The number, clause, or document does not exist. Rarer than the discourse suggests in well-built systems, common in general-purpose chat pointed at a folder, and catastrophic when it reaches an IC memo.

Misattribution. Every fact is real and one of them came from the wrong document. The 2023 figure is genuine, it is simply from the prior-year presentation. This is far more common than outright fabrication and much harder to spot, because everything checks out until you check where.

Silent omission. The answer is accurate and incomplete. Asked about customer concentration, the system reports the top-five concentration accurately and does not mention that the schedule excludes a segment. Nothing is wrong. Something is missing, and the output does not say so.

Instability. Ask the same question twice, get two answers, both defensible. This is the one prospects raise most often in demos and it is a real property of the technology, not a bug that gets patched. It matters most when the question is evaluative rather than extractive.

The protocol

Every number carries a citation to a page, or it does not go in the memo. Not a document name — a location. If the system cannot say which page of which file the figure came from, the figure is unverified and should be treated as unverified regardless of how right it looks. This one rule eliminates most fabrication risk, because fabricated content generally cannot produce a real citation, and it converts misattribution from invisible to checkable.

Spot-check on a fixed ratio, chosen before you see the output. Pick a number in advance — one in ten cited figures, say — select them at random rather than by which ones look suspicious, and open the source. Checking the ones that look wrong catches only the errors that look wrong, which is a biased sample of the errors.

Check the extractive and evaluative outputs differently. "What is the largest customer's share of revenue" has one right answer and either matches the source or does not. "Is this customer concentration acceptable" has no source to check against; it is judgment wearing the costume of an answer. Extractive output gets verified. Evaluative output gets read as an argument, weighed, and either adopted or discarded by a human whose name goes on the memo.

Test stability deliberately on anything material. Ask the material questions twice, in different sessions. Agreement does not prove correctness, and disagreement is highly informative: it tells you the question is under-determined by the documents, which is itself a finding worth writing down.

Run a negative control. Ask for something you know is not in the room. A system that produces a confident answer to a question about a document that does not exist has told you what its failure mode is, and it will do the same thing on a question where you do not already know the answer. This takes two minutes and is the single most useful test we know of. Run it during vendor evaluation and run it again quarterly.

What the protocol does not catch

Being straight about this matters more than the protocol itself.

It does not catch what was never asked. A verified answer to the wrong question is still a wrong turn, verified. The completeness of the question set is a separate discipline from the accuracy of the answers, and it is the one that actually determines diligence quality. Perfect citation on a first pass that never examined supplier concentration produces a well-sourced blind spot.

It does not catch the wrong judgment. A system can correctly extract every fact about a customer relationship and be wrong about what those facts mean for the durability of the revenue. That is not a verification failure. It is why the investment committee exists.

It does not transfer accountability. The associate whose name is on the memo owns the memo. There is no version of this where a citation makes the number somebody else's problem, and any process design that implies otherwise will eventually produce an IC meeting nobody enjoys.

The uncomfortable comparison

Firms sometimes hold AI output to a standard they have never applied to human output. Worth sitting with for a moment: how often does anyone check the associate's numbers back to source? In most firms the answer is that spot-checking happens when something looks off, which is the same biased-sample problem described above, applied to a process everyone already trusts.

This is not an argument for lowering the bar on AI output. It is an argument that the citation discipline described here is a good discipline generally, and that firms adopting it for machine-generated work often find they have improved the review of human work at the same time. The audit trail is the point. That it arrived attached to a new tool is incidental.

The short version

Require page-level citation. Spot-check at a fixed ratio, chosen blind. Verify extraction, argue with evaluation. Run a negative control before you buy and every quarter after. And keep the question of what was asked separate from the question of whether the answers are right, because the first one is where diligence quality actually lives.

The benchmark

See what the question set surfaced.

A controlled benchmark against a frontier general-purpose model on the same source materials, published in full.

Read the white paper

Work email only. We send it straight over.

Request the white paper See the benchmark results Browse all insights