Evaluate an AI startup by checking its pitch-deck claims against customer behavior, product performance, operating economics, dependencies, and execution—not by treating the deck itself as proof. This framework is for investors comparing companies at different stages and in different markets; the evidence that matters varies with the business model, deployment context, jurisdiction, and maturity. Diligence can reveal strengths, assumptions, and risks, but it cannot predict success with certainty.
Turn the pitch deck into testable claims
For each major claim, ask what observable evidence would support or weaken it. A statement such as “customers love the product” should lead to questions about who uses it, who pays, how often it is used, and what measurable outcome changed. “Our model is best” calls for task-specific evaluations, a relevant baseline, and examples of failure—not a polished demo. “We have a moat” calls for evidence of durable advantages and an account of dependencies.
Request the underlying records and definitions, not just headline metrics or selected customer stories. Keep verified facts, management assumptions, and unanswered questions distinct. If a metric is unavailable or too immature to interpret, record that limitation and specify what evidence would resolve it rather than estimating a result from the deck.
Check whether customers get durable value
Identify the user, buyer, and costly task
Establish who experiences the problem, who controls the budget, and what important or costly task the product improves. Then trace the product’s effect on the customer’s workflow: what work changed, what still requires human effort, and what outcome can be documented? User enthusiasm alone does not establish willingness to pay, and a buyer’s contract does not prove frequent use.
#1 Best Overall
Look for adoption that lasts
Use customer-level evidence and operating records to examine repeated use, renewal, expansion, contract duration, churn, and customer concentration. Ask whether usage continues after a trial or initial experiment and whether the product becomes important to a workflow. A large signup count or persuasive demonstration, by itself, does not establish recurring value.
CRV’s March 5, 2026 investor guidance highlights use that persists beyond experimentation, usage expansion, and high-value use cases that become indispensable. Renaissance Capital’s AI-company checklist also identifies workflow integration, API usage growth, enterprise adoption, real-world ROI, retention, revenue distribution, contract duration, and recurring revenue as dimensions to investigate. These are diligence prompts, not evidence that a particular startup meets them.
Test product performance and AI-specific risks
Ask for an evaluation you can interpret
Request a defined, task-specific evaluation rather than relying on a demo or a single aggregate score. Understand how test cases were selected, whether the set represents real use, which performance measures were chosen, and what baseline the company used for comparison. Ask to see limitations, uncertainty, and failure examples, along with results from conditions similar to the intended deployment.
Find out who can reproduce or independently review the evaluation. A result is harder to assess when the test set, measurement method, or comparison is unclear. Check whether the system can recognize cases where it should abstain or fail safely, and what happens when its output is wrong.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchReview oversight and ongoing risk management
Ask how performance is monitored after deployment, how incidents are handled, and where human review is required. NIST’s voluntary AI Risk Management Framework (AI RMF) organizes lifecycle work into Govern, Map, Measure, and Manage. Its guidance covers context mapping, documented testing and metrics, deployment-like evaluation, monitoring, and continuing risk management. NIST says the framework is voluntary; using it is not a certification or proof of product quality. NIST’s framework page reports that AI RMF 1.0 is being revised, so check the current version when using it as a reference.
Assess defensibility and dependencies
Separate demonstrable advantages from labels
Ask what gives the company an advantage that could persist: for example, deep workflow integration, properly licensed proprietary data, accumulated feedback, distribution, a specialized model or system, or customer switching costs. The relevant question is not whether the company calls its AI proprietary, but what evidence shows that its advantage is real and difficult to reproduce.
Trace the supply chain and fallback options
Map material reliance on third-party models, data, software, cloud infrastructure, and hardware. For each important dependency, examine access terms, rights and provenance, resilience, and contingency plans. Ask what would change if a foundation-model provider altered prices, access, terms, or capabilities. NIST’s AI RMF calls attention to third-party software and data risks, including possible infringement of third-party rights. Its July 8, 2026 ICT supplier due-diligence guide also discusses ownership and control, provenance, resilience, foundational cybersecurity practices, and supply-chain tiers. That guide is scoped to ICT supplier assessment, so use its dimensions proportionately rather than treating it as a universal startup scorecard.
Reconstruct the economics of growth
Reconcile the metrics to records
Ask how the company defines recurring revenue, gross profit, customer acquisition cost (CAC), customer lifetime value (LTV), payback, burn, and retention. Review the underlying data and cohort assumptions, and reconcile reported figures to financial records. Make sure the definitions, time windows, and customer segments are consistent before comparing metrics across companies.
Include the costs of delivering AI
Where relevant, include inference, hosting, customer-specific training, onboarding, and support in the cost of serving customers. Separate materially different acquisition and delivery motions—such as self-serve, product-led, and enterprise sales—instead of relying on a blended average that conceals their economics.
CRV’s March 5, 2026 AI SaaS guidance notes that inference, hosting, and customer-specific training can scale with usage and pressure gross margins. Its July 23, 2026 Series A article recommends considering CAC, LTV, payback, margin, and retention together, with assumptions and segments visible. These are investor perspectives, not universal cutoffs. A payback figure has meaning only in context: acquisition channel, gross profit, cash timing, and retention all affect how it should be interpreted. Do not turn a rule of thumb into a pass/fail threshold without evidence that it fits the company’s model and stage.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Review team execution, governance, and security
Assess whether the team has relevant technical, product, commercial, and domain expertise—and whether its members can explain tradeoffs candidly. Compare roadmap commitments with what has shipped and what customers can substantiate. A compelling plan is not the same as a record of execution.
Identify who is responsible for model evaluation, privacy, security, incident response, customer complaints, and oversight. Ask whether those roles are documented and resourced. Review data rights, access controls, vulnerability handling, and third-party risk. Determine which legal or regulatory duties apply to the actual product, use case, and jurisdictions; do not infer compliance from a general policy statement or a framework reference. NIST’s AI RMF emphasizes governance, documented roles, context-specific risk mapping, evaluation, feedback, monitoring, third-party risks, and ongoing management. Renaissance Capital’s checklist also flags regulatory readiness, privacy safeguards, security, and governance as areas to assess, but neither source establishes that a particular startup complies with applicable law.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Compare startups using the same evidence
When evaluating alternatives, use consistent definitions and time windows. Compare evidence rather than awarding points for polished presentations. The dimensions below help organize the comparison; they are not a universal weighting system.
Quick Recap
| Dimension | Evidence to compare |
|---|---|
| Customer value | Importance of the use case, documented outcomes, repeat use, renewal, expansion, and customer concentration |
| Product quality | Task-level performance, reliability, failure modes, fit to deployment conditions, and human oversight |
| Economics | Gross and contribution margin, inference and service costs, acquisition channel, payback, cash needs, and retention |
| Defensibility and resilience | Data and intellectual-property rights, workflow integration, vendor dependence, compute access, switching costs, and contingency plans |
| Risk readiness | Relevant privacy, security, fairness, and safety testing; governance; monitoring; incident response; and jurisdiction-specific obligations |
| Execution | Team capability, delivery against milestones, quality of evidence, and connection between milestones and customer or operating outcomes |
Use a practical diligence sequence
- Translate claims into questions. List the important assertions in the deck and request the underlying evidence for each.
- Validate the customer case. Establish the problem, user, buyer, workflow change, and measurable outcome using customer-level evidence.
- Inspect the product in realistic conditions. Use task-specific tests and deployment-like settings; document performance limits, failure cases, and oversight.
- Map material dependencies. Trace model, data, software, compute, and cloud reliance, including rights and fallback plans.
- Rebuild the economics. Reconstruct cohort retention and unit economics using fully loaded costs, including AI-variable delivery costs, and separate distinct sales motions.
- Examine execution and controls. Review team delivery, governance, security, privacy, monitoring, and incident practices.
- Separate conclusion from uncertainty. Record supported evidence, assumptions, unresolved questions, and downside cases separately from the investment thesis.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

