How to Evaluate an AI Startup: What to Actually Verify
Every deck says "AI" now, which means the label tells you almost nothing. Evaluating an AI startup well is about looking past the word to what is actually there: whether there is a real model or a thin wrapper, whether the data is a moat or a liability, whether the traction is usage or demos, and whether any of it is defensible. This guide covers the claims worth verifying on an AI company and the ones that sound impressive but are not.
The AI label has become free to apply, so it has stopped separating companies. A startup that fine-tuned a model and one that pasted an API key into a form both say the same word on the same slide. For anyone screening or funding, that collapses the usual signals: the pitch quality, the demo polish, and the buzzwords are all uncorrelated with whether the company has anything durable. What still separates them is what you can verify underneath the label.
This guide walks through what to actually check on an AI startup: the model claim, the data claim, the traction claim, the moat, and the team, plus the specific ways AI pitches inflate and how to keep verification honest without pretending to detect the undetectable.
First, see past the label
Start by refusing to let "AI" do any work in the evaluation. The word describes a category, not a capability, and the interesting question is always the specific one underneath: what does the model do, what is it trained or grounded on, what would a competitor have to reproduce, and what happens when the underlying provider changes its terms or prices. A company that cannot answer those crisply is telling you the AI is a marketing layer, not a foundation.
This is not skepticism for its own sake. It is that the AI label is the single most inflated claim in a deck right now, and treating it as a neutral input rather than a claim to be checked is how a wrapper gets funded as if it were a model. Evaluate the substance, and let the label earn its place or not.
The model claim: real model or thin wrapper?
The first thing to verify is what the company actually built. There is a spectrum, and where a startup sits on it changes everything about its risk:
- A thin wrapper calls someone else's model through an API and adds a prompt and a UI. This can be a good business, but its "AI" is rented, its margins depend on a provider's pricing, and a competitor can rebuild it in a weekend. That is fine if the company is honest that the value is elsewhere, in distribution, workflow, or data, and a problem if the wrapper is being sold as the moat.
- A real model or a substantial adaptation means the company trained, fine-tuned, or meaningfully engineered something that a competitor cannot trivially copy. This is harder to build and more defensible, and it is also easier to overstate.
The verification is to separate what the company owns from what it rents. A claim of "our proprietary model" against evidence that the product is a straightforward call to a public API is the kind of gap that reorders a ranking. You are not judging that wrappers are bad. You are checking that the company is what it says it is.
The data claim: moat or liability?
AI companies love to claim data as a moat, and data is the claim most worth checking, because it cuts both ways. Real proprietary data that a model genuinely improves on is a durable advantage. But a data claim can also hide two problems: the data may not actually be proprietary, or it may be a liability rather than an asset.
The questions to verify are ownership and standing. Does the company have the rights to the data it trains on, or is it scraping something it does not own and calling it a moat? Licensed data has terms that limit what can be redistributed or built on, which constrains the supposed advantage. Data gathered without clear rights is not a moat, it is exposure. The honest version of a data-moat claim survives the question "and you are allowed to use all of that, how"; the inflated version does not.
The traction claim: usage or demos?
Every category has a traction-inflation problem, and AI has a specific one: the gap between a demo that dazzles and usage that persists. A model that produces an impressive result in a controlled demo is not the same as a product people use twice, and AI demos are unusually good at looking like traction while being neither retained nor paid.
So verify traction the way you would for any company, with extra care for the AI-specific mirage. Real usage, real retention, real revenue are checkable signals. A viral launch, a waitlist, and a demo video are not. This is the same discipline that applies across the Deckwise Method: treat the traction figures as claims to check against evidence, not as facts because they were on a slide, and be careful with metrics that shift over time rather than scoring a stale number as a lie.
The moat and the team
Two more claims deserve verification. The moat: if the AI is rented and the data is not owned, "AI" is not the moat, and the company needs a real one somewhere else, in distribution, a workflow lock-in, or a genuine data or model advantage. Naming the moat honestly is itself a signal; a company that insists the moat is "our AI" while renting the model has not thought it through, or is hoping you will not.
The team: for a company building something technically real, whether the people can actually build it is central, and it is verifiable, through backgrounds, prior work, and public record, more than through the confidence of the pitch. A strong technical claim from a team with no evidence of being able to deliver it is a claim, not a fact.
Keep verification honest
A word of caution that matters especially here. Verifying an AI startup means checking claims against public evidence, corporate standing, funding history, named partnerships, ownership of data, the reality of the model, and labeling honestly what cannot be confirmed. It does not mean claiming to detect whether a deck was written by AI, or scoring a company down for using AI tools. That is not a verifiable signal, and treating it as one would be exactly the kind of false certainty this whole approach exists to avoid. Verification credits what the evidence supports, qualifies what it cannot, and leaves the rest honestly marked as to-confirm.
Frequently asked questions
What should I verify first on an AI startup? The model claim: whether the company built or adapted something defensible, or is calling someone else's model through an API. That distinction changes the whole risk profile, and it is checkable by separating what the company owns from what it rents.
Is a "wrapper" startup automatically a bad investment? No. A wrapper can be a strong business when the value is in distribution, workflow, or data and the company is honest about that. It is a problem only when the rented model is being sold as the moat. Evaluate where the durable advantage actually sits.
How do I check an AI data moat? Verify ownership and rights, not just existence. Ask whether the company owns the data or licensed it, and on what terms, because licensed data limits what can be built on or redistributed, and data gathered without clear rights is a liability rather than a moat.
How is AI traction different to verify? AI demos are unusually good at looking like traction. Separate real, persistent usage and revenue from a viral launch, a waitlist, or a demo video. Verify the usage and retention figures as claims against evidence, with care for metrics that legitimately change over time.
Can you detect whether a startup is "faking" its AI? You can verify concrete claims, whether a model is proprietary, whether data is owned, whether traction is real, against public evidence. You cannot reliably detect AI use itself, and a responsible evaluation does not pretend to. It checks what is checkable and honestly marks what is not.
The bottom line
Evaluating an AI startup is evaluating a company whose most inflated claim happens to be its category label. See past the word, and verify the substance: is the model built or rented, is the data owned or exposed, is the traction usage or a demo, and where does the real moat sit. Check what is checkable, qualify what is not, and never mistake the AI label for the answer to any of those questions.
How this shows up in Deckwise
Deckwise evaluates an AI startup the way it evaluates any candidate, across the full pipeline: it sources against your thesis, scores on one basis, verifies each material claim against public evidence, and goes deep on the finalists. For an AI company that means checking the concrete claims, corporate standing, funding, data ownership signals, named partnerships, and marking honestly what cannot be confirmed, rather than scoring the buzzword. Verification is one of four strengths, not a lie detector: it credits what the evidence supports and qualifies the rest, and the human decides.
Related: The Complete Guide to Startup Due Diligence · The Verification Layer in Deal Flow · The Deckwise Method