Category Capture · 19 August 2026

What should you ask an AI diligence vendor before you buy?

Twelve questions that separate purpose-built diligence systems from general-purpose AI with a finance skin, covering security, data portability, and build-versus-buy.

Key takeaways
  • Most software sold as purpose-built diligence tooling is a general-purpose language model with a retrieval layer and a finance-themed interface, and a scripted demo on the vendor's own sample data cannot tell you which is which.
  • Run the system on a real data room from a closed deal where you already know the answers, and ask it, unscripted, for something that is not in the room.
  • 'Cannot assess — the schedule lacks contract dates' is the right answer to an unanswerable question; a confident estimate is disqualifying.
  • The commitment that no uploaded data trains a model belongs in the contract, not the FAQ page.
  • After ingestion your archive exists in a structured, entity-resolved form more valuable than the folder you started with — ask in writing what you receive on termination, in what format, and at what cost.
  • An enterprise ChatGPT licence has no memory of your deals and no completeness standard, so it answers what you ask and cannot tell you what you failed to ask.

The category has a structural problem for buyers. Most of what is sold as purpose-built diligence software is a general-purpose language model with a retrieval layer and a finance-themed interface, and from the outside, in a scripted demo, on the vendor's sample data room, it is close to impossible to tell which is which. Every vendor says the same words. The demos look the same.

Twelve questions do most of the separating. None of them are gotchas and a good vendor will answer all of them without friction.

On the system itself

1. Run it on our data room, not yours. The single most informative thing you can do. Vendor sample rooms are clean, complete, and text-native. Yours has a scanned 2016 lease, a schedule where the header row is in the fourth row, and a supplier agreement in a language nobody at the vendor speaks. Ask for a live run against a real room from a closed deal, where you already know the answers.

2. Ask it something that is not in the room. Described at more length elsewhere on this site, but it belongs on any purchase checklist. A system that confidently answers a question about a document that does not exist has shown you its failure mode. Do this in the demo, unscripted.

3. Where does every number come from? Page-level citation or the output is unverified. Ask to click through from a figure in a generated memo to the exact page. If the answer is a document name rather than a location, the verification burden falls entirely on your associate.

4. What happens when it does not know? Ask to see the output for a question the room cannot support. "Cannot assess — the schedule lacks contract dates" is the right answer. A confident estimate is a disqualifying answer.

5. Can it apply our criteria, or only its own? The difference between a tool and a system. Ask whether you can load your investment principles and have every deal scored against them in your order, or whether you get a generic framework with your logo on it.

On your data

6. Is anything we upload used to train a model, ever? The answer needs to be no, and it needs to be in the contract rather than the FAQ page. A meaningful share of enterprises have restricted or banned generative AI tools over exactly this, and diligence data is the most sensitive corpus a fund holds: MNPI on companies you may not end up owning, and LP terms.

7. Where does it run, and who else is in the tenancy? Shared infrastructure is fine for many purposes. It is worth knowing that it is shared, and knowing whether an in-tenant or single-tenant deployment is available at what price, before rather than after your compliance review.

8. Which underlying models, and what happens when they change? Every vendor in this category sits on top of foundation models they do not control. Model versions get deprecated on the provider's schedule, and output changes when they do. Ask what happens to your saved workflows the week a model is retired, and whether you are notified before or after.

9. What comes out if we leave? The question almost nobody asks and the one with the longest tail. After ingestion, your archive exists in a structured, entity-resolved, extracted form that is more valuable than the folder you started with. Ask, in writing: on termination, do we receive that structured version, in what format, and at what cost? A vendor whose answer is that you get your original files back is telling you the improved asset is theirs.

On the firm

10. How long has this been built and by how many people? Prospects ask us this constantly and they are right to. The answer tells you about roadmap risk, support depth, and whether the thing you saw is production or a well-rehearsed prototype. A small team using frontier tooling can ship remarkable things quickly; you are entitled to know that is what you are buying.

11. Who else in our segment uses it, and will you connect us? Reference calls with a firm of your size and shape are worth more than any feature list. A mid-market buyout fund and a family office running two direct deals a year have almost nothing in common as users.

12. What does year two cost? Including ingestion of new history, seat growth, and any re-processing. First-year pricing in this category is frequently strategic.

The build-versus-buy question underneath all of this

Firms with a strong technical partner reliably ask whether they should just build it. The honest answer is that the first 60% is genuinely easy now and the last 40% is where every internal project stalls.

The easy part is a chat interface over a document store. A capable engineer can build a convincing version in a fortnight, and many have. The part that does not get built is everything in the middle of the ingestion problem: optical character recognition on scanned contracts, table extraction from PDFs where structure is implied by formatting, entity resolution across a decade of naming conventions, and the evaluation harness that tells you when a change to the prompt made the output quietly worse.

That last one is what separates a system from a demo, and it is the thing internal builds almost never have. Without it, nobody can tell whether the tool is getting better or worse, and confidence erodes over about six months until the tool is used for summaries and nothing that matters.

MIT's NANDA work on enterprise AI deployments found that internally built systems succeeded at roughly half the rate of purchased or partnered ones. That ratio matches what we observe: the build succeeds where the firm is prepared to staff it permanently, and fails where it was scoped as a project.

And the one about ChatGPT

Firms with an enterprise license reasonably ask what they are missing. Three things, concretely.

It has no memory of your deals beyond the session or project you put it in, so it cannot tell you that this target resembles one you passed on in 2023. It has no completeness standard, so it answers what you ask and cannot tell you what you failed to ask. And it will read a 200-page CIM with embedded tables and inconsistent formatting less reliably than a system built for that document type, in ways that are not visible in the output.

For market research, drafting, and thinking out loud, it is excellent and you should keep using it. For the specific job of reading a data room against a standard and reporting what is absent, it was not built for that and does not claim to be.

The benchmark

See what the question set surfaced.

A controlled benchmark against a frontier general-purpose model on the same source materials, published in full.

Read the white paper

Work email only. We send it straight over.

Request the white paper See the benchmark results Browse all insights