Seed investors buy a founder and an insight. Series A investors buy evidence. For an AI-native company, that shift is sharper than most founders expect. The demo that raised the seed round shows what the model can do. Technical diligence asks what the system does, every day, at scale, with real data and real customers.
Preparing technical due-diligence packs and data-room answers is part of the CTO work we do for the companies we build. These are the questions we prepare founders for, and the evidence that answers them better than a slide.
1. What is defensible here that isn't the model?
Every serious investor has seen a hundred thin wrappers around the same model APIs. The question underneath is: if a competitor gets the same model tomorrow, what do you still have?
Good evidence: proprietary domain data and the rights to use it; a workflow that encodes expert judgement, such as WaterDoctor's process science or a chartered accountant's close checklist; operating history with customers; and the evaluation sets you built from real cases. A model can be swapped. A library of graded real-world cases cannot be copied.
2. How do you know it works, and how do you know when it stops?
"Our users like it" is not an answer. Investors increasingly ask about evaluation: a fixed set of test cases, a pass bar, and a record showing that every release was measured against it.
There is a quieter version of this question that matters more. Many agent products stall not because the model is wrong but because checking its output costs more than doing the work. If senior staff must read every output before it reaches a customer, the product is not saving what the pitch deck claims. Be ready to show where human review happens, how much of the output needs it, and how that share is falling.
3. What happens when the model provider ships a new version?
This is the risk founders underestimate most. In one of our own internal systems, a routine upgrade to a newer model version, treated as a patch release, silently dropped 14% of inbound leads over three days. Nothing crashed and no alert fired. The new model returned its output in a slightly different structure, and the code downstream quietly discarded what it couldn't parse.
Good evidence: pinned model versions; a replay suite of recorded real inputs that every upgrade must pass; output validation that fails loudly; and a written upgrade procedure. Treat a model upgrade like a data migration, not a dependency bump.
4. What does each unit of value cost to produce?
Gross margin for an AI product depends on inference cost, and inference cost depends on architecture. In 2023, one of our internal marketing agent crews outspent an engineering crew four to one on model usage. It wasn't producing four times the output. A retrieval step was stuffing irrelevant context into every call.
Good evidence: cost per outcome (per report, per case, per filing), not a monthly API bill; hard token budgets per agent or crew; and a clear view of which costs fall as usage grows and which rise with it.
5. When it's wrong, who answers for it?
"The model did it" is not an incident report. Investors in regulated markets such as health, finance, legal and environmental will ask who owns each agent in production, what happens when it makes a mistake, and what the audit trail shows.
A trap here: a model's own "reasoning summary" is not audit evidence. Providers generate those summaries, often hide or encrypt the underlying reasoning, and differ in what they expose. Audit evidence is what your system recorded: the inputs, the tool calls, the checks that ran, the person who approved the output, and when.
Good evidence: a named owner for every production agent; escalation paths; approval gates at the points where judgement matters; and logs you control.
6. Where does personal data go?
In Singapore, the PDPC's Proposed Advisory Guidelines on Use of Personal Data in Generative AI, issued for consultation in June 2026, make an existing PDPA expectation concrete. Organisations should be able to account for personal data across the whole model workflow. That includes the places engineers forget: raw prompts in debug logs, analytics dashboards and observability tools with no retention limit.
Good evidence: a data map that includes prompts and logs; redaction before data reaches the model; retention policies on every log group; and alignment with IMDA's Model AI Governance Framework for Agentic AI if you sell to government or regulated buyers.
Diligence rewards companies that already know their weak spots and have a plan for them. It punishes surprises.
How we use this list
When we build a company with a founder, these questions are part of the build plan from the first month, not a scramble before the data room opens. The evidence is a by-product of building properly: evaluation harnesses, pinned versions, cost tracking, named owners and audit trails. By the time a Series A investor asks, the answer already exists. See how the engineering bench works.
- wGrow field note: Agent adoption is stuck at verification
- wGrow field note: Upgrading an LLM version requires data-migration thinking
- wGrow field note: Token budgets are department budgets in disguise
- wGrow field note: "The model did it" is not an incident report
- wGrow field note: Reasoning summaries are not audit evidence
- wGrow field note: PDPC GenAI guidelines turn training data into a delivery question