Forget About the ARR: The New Playbook for AI Due Diligence

@mardehaym
ENGLISCH20. Aug. 2026
104K
159
28
29
80

TL;DR

Mark Ajzenstadt argues that traditional SaaS diligence fails for AI companies, often hiding 'wrappers' behind strong ARR. He proposes a 4-axis technical scorecard to evaluate true defensibility.

If you're reading this, I trust the algorithm targeted you because you evaluate software companies for a living. Experienced deal teams love saying they run thorough diligence.

I intend to show you why, for AI deals, they don't.

Most of an acquisition's risk surface gets covered by standard financial diligence. Revenue quality, customer concentration, churn cohorts, unit economics, TAM sizing. These playbooks exist because they work. They've killed bad deals and validated good ones for decades.

Technical diligence for traditional SaaS is solid too. Code quality reviews, infrastructure scalability, security audits. Proven. Reliable.

But AI companies are not traditional SaaS companies.

The playbooks that evaluate a CRM platform or a project management tool were never designed for a world where the "proprietary technology" is an API call to someone else's model with a frontend on top.

Standard playbooks tell you how much revenue the company generates.

They can't tell you whether the company built what it's selling.

So why did I title this "forget about the ARR"?

All I've been doing is defending the importance of financial and technical diligence. It is important.

But in order to evaluate an AI company, you have to forget about the spreadsheet and open the codebase.

The spreadsheet shows ARR and growth. It doesn't show whether any of it is defensible. Forget everything you think you know about AI company evaluation, because you haven't been taught how to do it.

You've been taught how to value SaaS companies, and you're applying the same playbook to a fundamentally different product category. That's what's holding your diligence back.

Quick disclaimer: I embed AI-native engineers into PE-backed firms.

I'm not a disinterested observer. I'm someone who's opened hundreds of codebases behind pitch decks that said "proprietary AI." Some were real. Most were not. I don't have all the answers on valuation. But I've learned a few things about what's behind the deck that might help if you're evaluating your next AI target.

Lesson 1: the pitch deck is a configuration file.

That sounds ridiculous. Let me show you.

A PE firm asked us to review an AI company before close. The pitch deck said "proprietary AI platform." The data room showed $4.2M ARR growing 40% year-over-year. The management presentation had 14 slides on their "foundational model architecture." Detailed diagrams. Proprietary terminology. Impressive.

We opened the codebase.

The "proprietary model" was gpt-4o with temperature set to 0.2. A system prompt. A React frontend. Six hundred lines of glue code any senior engineer could rebuild in a weekend.

They were asking 12x revenue.

The PE firm walked within the day.

I've reviewed a lot of codebases. The gap between pitch and product on this one was the widest I've seen in three years of reviews. No model. No training pipeline. No proprietary data. The entire product was an API call with a UI on top.

The ARR was real. The customers were real. Growing 40% a year. Real.

The AI was not.

You should not look at ARR and assume the technology behind it is defensible. You should open the codebase. The ARR is real, but the thing generating it might be 600 lines of code anyone can rebuild. "Anyone can rebuild it" is a fundamentally different valuation conversation than "proprietary AI platform."

Bain & Company arrived at the same conclusion from a different direction. They introduced a specialist engineering capability in 2023 for PE diligence. The team vibecodes functional replicas of target companies' software using AI code generation tools, then evaluates whether the product's competitive advantage lives in the code, the data, the workflow design, or the marketing deck.

Rebecca Burack, head of Bain's global PE practice, described the difference as "seeing something in 2D versus 3D." They've built hundreds of these prototypes. In at least one documented case, a PE firm pulled its bid on an analytics platform after Bain's prototype showed the core logic was reproducible.

We do the same thing from the other side. They build replicas to test reproducibility. We open the codebase to test whether there's anything proprietary to reproduce.

The company we reviewed had zero proprietary model logic. Their competitive moat was a system prompt and an API key.

That's not IP. That's a configuration file with a monthly bill attached.

The firms I know who evaluate AI companies correctly are not staring at the spreadsheet while they work. They're reading the code. The spreadsheet shows up as a secondary artifact. A byproduct of the technical reality, not a substitute for it.

This is not a feel-good LinkedIn point about "going deeper." I want you to change your relationship with how you evaluate AI acquisitions.

When you're fixated on financial metrics, every evaluation gets filtered through "what's the ARR growth rate" and "what multiple are comparables trading at." That filter consistently points you toward deals that look good on paper and collapse under technical scrutiny.

When you're focused on what's actually in the codebase, you make decisions based on "what did they build" and "can someone else build it for less." Those decisions may kill a deal in week one, save you from a write-down in year three, and look brilliant in year five. If it's unclear what I mean: the PE firm that hired us for that code review made one phone call after reading our findings. They walked the same day. Saved themselves a 12x multiple on 600 lines of glue code. Do you think they care about the lost deal fees? Probably. But the cost of acquiring a wrapper at platform pricing would've been harder to stomach than the cost of one more diligence workstream.

Lesson 2: the four blind spots your playbook doesn't cover.

The only diligence framework that matters for AI acquisitions: what risks does standard diligence miss?

Four gaps. The overlap between them is where deals look great on paper and destroy value after close.

PE firms I talk to have a financial diligence playbook and a commercial one. Almost none have a technical playbook built for AI codebases.

Gap 1: Model dependency.

If the target runs on a third-party model (OpenAI, Anthropic, Google), you need to know what happens when that provider raises pricing, deprecates the model version, or a competitor ships equivalent capability at 20% of the cost. Valutico calls this "agentic substitution risk" in their 2026 buyer's framework. I call it the thin-wrapper problem. If you can swap the underlying model in an afternoon, the "proprietary AI" is worth the React frontend it's wrapped in.

The company we reviewed scored a 10 out of 10 on this axis. Zero proprietary model logic. Zero fine-tuning. Their moat was a system prompt and an API key. That's a monthly subscription with a nice UI, not intellectual property.

Gap 2: Training data liability.

You acquire an AI company, you inherit how they sourced their training data. The datasets they scraped. The copyrighted material they ingested without permission. Most data rooms I've seen contain nothing on training data provenance.

Bartz v. Anthropic settled for $1.5 billion covering 482,460 books, roughly $3,100 per work after fees. Anthropic was required to destroy the pirated libraries and derivative copies within 30 days of final judgment. The settlement covers past conduct only; it doesn't establish future licensing frameworks.

EU AI Act penalties for prohibited AI practices run to 7% of global turnover or €35 million, whichever is higher. Other violations carry penalties up to 3% or €15 million. Six countries have already issued conflicting interpretations of whether AI training on copyrighted material qualifies as fair use. That creates a liability patchwork that any cross-border acquisition inherits.

You sign the purchase agreement, you sign for all of it.

Gap 3: Talent concentration.

In a traditional SaaS company, losing two engineers is a setback. In an AI company, the model pipeline often lives in three people's heads. I've seen shops where one ML engineer's departure would mean the company can't retrain or update its core product. Standard HR diligence counts headcount. It doesn't map who holds institutional knowledge about the model architecture, whether the training pipeline is documented, or whether a new hire could run it.

One engineer. No docs. 18-month earn-out. You're not buying a company. You're buying a countdown timer.

Gap 4: Integration failure rates.

RAND cited estimates that more than 80% of enterprise AI initiatives fail to deliver intended business value, roughly double the failure rate for comparable IT projects without AI. RAND's own study is qualitative, built on 65 interviews with data scientists and engineers. Treat the number as directional rather than precise.

Separately, MIT found roughly 95% of generative AI pilots returned zero measurable P&L impact. That measures financial return, not technical failure, and the study isn't peer-reviewed. But the direction is consistent.

Acquiring an AI company and plugging it into a portfolio company's stack is a multi-quarter engineering project carrying that same probability distribution. The product won't work on day one. It's an integration project with an 80%+ historical base rate of failure to deliver value.

These four gaps don't show up in the data room. They don't appear in the financial model. They don't surface in customer reference calls. You find them by opening the codebase and talking to the engineers who wrote it.

Be painfully honest with yourself: what does your current diligence process actually evaluate versus what should it evaluate for AI? Where do those two answers diverge?

That's your gap. Everything else is bullsh*t.

Lesson 3: this applies to every AI deal, but most firms won't do it.

A quick caveat: not every AI company is a wrapper. Companies with proprietary models, proprietary training data, documented pipelines, and distributed expertise exist. Acquiring those companies is a fundamentally different proposition.

I don't mean that as a hedge. I mean it statistically, practically, and honestly.

The nature of the current market is that there are companies that built something real and companies that wrapped someone else's model in a frontend and called it proprietary. Neither necessarily has bad ARR. Neither necessarily has bad growth. They look identical in the spreadsheet.

The difference shows up in the codebase. Most firms never look.

KPMG's 2025 Technology M&A Survey found that 66% of dealmakers acknowledge technical and AI debt as a concern during deal planning, but only 33% prioritize investigating it during pre-deal evaluation. Thirty-three points between acknowledging a risk and actually checking for it. KPMG recommends seven diligence dimensions evaluated through an AI-specific lens. Traditional frameworks, they note, "fail to evaluate AI-specific signals."

If you're reading this far into an article about AI diligence, you're probably at least curious about whether your own process has these gaps. Curiosity is enough. You don't need a dedicated AI diligence team or a PhD in machine learning. You need someone who can open the codebase and tell you what's behind the pitch deck.

So open it.

Quick sidebar: many PE firms acknowledge the risk of AI technical debt during deal planning. They discuss it in IC meetings. They reference it in memos. Then they close without a single code review. When anyone suggests adding a technical workstream, the answer is "we're on a tight timeline."

Every deal is on a tight timeline. But the "tight timeline" crowd treats technical diligence like a luxury when it's actually the most asymmetric check in the entire process.

If you knew what was behind the pitch deck before you signed the LOI, you would not sleep. You'd rewrite the deal structure overnight.

Am I admitting that we run these reviews for PE firms as a paid service, and this article effectively argues for the exact thing we sell?

Yes. I'm admitting that. The PE firm that hired us for that code review spent a fraction of what the transaction fee would've been. They saved themselves a 12x multiple on 600 lines of glue code.

But hypothetically, if a firm were to add this one workstream to their process, even once, they'd dramatically reduce their risk. They'd have technical evidence to back up their thesis, and real data to inform their pricing. That's not a luxury. That's the cheapest insurance in the deal stack.

The "tight timeline" crowd acts like it's all-or-nothing. Full 90-day technical diligence or nothing. That's a false binary. A focused code review takes days, not months. You can scope it strategically. The timeline is real. The excuse is not.

Lesson 4: spend money on diligence to save money on deals.

This one took the market a while to internalize.

When you're in deal mode, every diligence dollar feels like friction. You've paid for the financial model. You've paid for commercial analysis. You've paid for legal review. Adding another workstream feels like slowing down a process that needs to close. This instinct persists even after you've seen bad outcomes. You'll want to cut corners, shorten the scope, compress the timeline. This is setting yourself up for a write-down.

When you're in operator mode, diligence spend is deployment, not cost. The goal is to deploy capital where it reveals the truth, not to minimize the diligence bill.

That means hiring the engineering team that can actually read the codebase even though they cost more than a checkbox audit. It means paying for the assessment that takes three extra days, because those three days tell you whether you're acquiring a platform or a prompt. Invest in the evaluation that breaks your thesis before you invest $50 million based on it, because all returns from diligence capability are exponential. The only thing that's not exponential is the cost of one more workstream.

Firms that lost money on AI acquisitions (I mean this in outcome, not just valuation decline) optimized for closing fast. Firms that preserved capital optimized for evaluating correctly. One depletes portfolio value. The other compounds it. If you can't wrap your head around this one, come back to it after you've seen the inside of one AI codebase that didn't match its pitch deck.

I'm not saying be slow. Understand the difference between diligence spending and diligence investing: one adds to deal cost, the other subtracts from portfolio risk.

Bain understood this. They built an entire engineering capability to vibecode replicas of acquisition targets. That's a meaningful investment. But the alternative, closing on a deal where the core product is reproducible in a weekend with an API key and a frontend framework, costs a lot more.

Lesson 5: the asymmetric bet of walking away.

Most PE firms avoid walking away from deals with strong financial metrics. Some walk from obviously bad deals. Very few walk from deals that look good on paper but fail under technical scrutiny. Another truth that sounds obvious until you check the data: the biggest risk in AI M&A is acquiring something you haven't evaluated.

A smart walk is asymmetric: limited downside, unlimited upside.

Walking away from an AI wrapper with $4.2M ARR growing 40%? Asymmetric. Worst case, the company turns out to be real and you missed it. Best case, you deploy that $50M into something defensible instead.

Closing at 12x revenue without opening the codebase? Not asymmetric. You can inherit a $1.5 billion training data liability. You can acquire a product that one ML engineer can shut down by leaving. You can buy 600 lines of glue code for $50 million. That's not investing. That's gambling, and not the kind with favorable odds. There are more disciplined ways to deploy capital.

Thoma Bravo is learning this lesson at scale. The firm faces a projected $5.1 billion equity loss on Medallia, acquired for $6.4 billion in 2021. Orlando Bravo told CNBC: "We made a mistake." Peak-cycle pricing and overaggressive growth assumptions. The restructuring has become a cautionary reference for the entire 2021-2022 take-private cohort: Anaplan at $10.7B, Coupa at $8B, Citrix at $16.5B. All struck at peak multiples that no longer hold.

PitchBook reports PE software platform buyouts fell to 41% of deal value in 2026, the lowest level in a decade and a 30-percentage-point drop from the prior year. Only seven software platform transactions exceeded $100 million through May 2026. US software deal value is running at about one-quarter of 2025's $156 billion pace. Investors are shifting capital to add-ons and growth equity, which now represent 45% of deal value.

The pullback isn't just rates and credit tightening. Buyers are losing confidence that they can tell what they're buying.

Meanwhile, AI M&A deals grew 90% year-over-year in Q1 2026. 266 deals closed in Q1 alone. Nearly half of all tech deals now carry an AI component, up from one in four in 2024. Private-market AI companies command 8-15x revenue multiples versus 4-6x for comparable traditional SaaS.

But the premium is diverging: companies with proprietary data and embedded AI attract buyers. Pure wrappers without proprietary data, workflow, or distribution advantages are struggling to attract serious interest.

The game is finding deals where you can afford to be wrong about the growth trajectory, but being right about the technology creates defensible value. "Afford to be wrong" varies firm to firm. "Defensible value" varies deal to deal.

Most firms never walk because they're optimizing for the wrong metric: capital deployment instead of capital preservation. You're playing offense in a market that increasingly rewards defense. Evaluate more carefully before committing, because the other side of the table knows you're not checking.

Lesson 6: score it before you sign it.

The most counterintuitive thing I've learned from ten years of code reviews: don't sign before you score.

But Mark, I thought you said do diligence. Shouldn't you already know the answer by the time you get to the LOI?

Every firm I've worked with has a version of this story. They had a deal that looked clean. Financial diligence cleared. Commercial diligence cleared. References checked out. They signed. Then they discovered the technical reality. The thing they paid 12x for turned out to be worth the frontend it was wrapped in.

The temptation to sign on strong financial metrics is enormous. When the ARR is growing 40%, when gross margins look right, when customer references are positive, adding another diligence workstream feels irrational. It feels irresponsible to slow down.

This could not be further from the truth.

If you're evaluating something that's working financially, the best move is almost always to open the codebase before you sign. Let the technical reality inform the valuation. Don't interrupt the diligence with a closing deadline. You can take shortcuts on timeline management. But never skip the code review entirely.

Valutico published a four-axis framework for evaluating AI vulnerability in M&A. We run a version across our engagements. Score each axis from 0 (low risk) to 10 (high risk):

Axis 1: Model Dependency. 0 = fully proprietary model. 10 = pure API wrapper. What percentage of core functionality depends on third-party model APIs? What's the contractual relationship with each provider, including pricing terms, usage limits, and termination clauses? If the primary provider raised prices 5x tomorrow, what happens to margins? Can the team demonstrate the product running on an alternative model within 48 hours? What proprietary fine-tuning, training data, or model modifications exist beyond prompt engineering? Red flag: the team describes "prompt engineering" as their moat. That's a configuration file with a monthly bill attached.

Axis 2: Data Moat. 0 = irreproducible dataset with clear provenance. 10 = no proprietary data. What proprietary datasets does the company own, and can they demonstrate legal provenance for each source? Is training data provenance documented anywhere in the data room? Could a competitor with access to public data reproduce the company's model performance within six months? Do data licensing agreements survive a change of control? Red flag: no training data documentation. After Bartz ($1.5B, 482,460 books), this is a liability you inherit on signing.

Axis 3: Talent Concentration. 0 = knowledge distributed and documented. 10 = single point of failure. How many people can retrain or update the core model? Is the training pipeline documented and reproducible by someone who didn't write it? What happens to the product roadmap if the top two ML engineers leave after their earn-out? Can a new hire run the full model update cycle from documentation alone? Red flag: one engineer, no docs, 18-month earn-out. You're buying a countdown timer, not a company.

Axis 4: Substitution Risk. 0 = mission-critical with unique workflow. 10 = commodity automation. Does the product automate a task that an off-the-shelf AI agent could replicate tomorrow? What would it cost a well-funded competitor to rebuild the core functionality? Does the product's value come from the model, the data, the workflow, or the distribution? Red flag: core functionality is reproducible in a weekend with an API key and a frontend framework. We know. We've seen the codebase.

Score each axis. More than one score above 8 means the asking price needs to reflect the risk profile, not the growth rate. Two or more axes above 8 and the deal structure, escrow holdbacks, earnouts tied to technical milestones, acqui-hire pricing, should reflect it. Run this before your LOI, not after.

This applies to capability, too. Don't cash in your deal flow for a comfortable close too early. Keep evaluating. Keep building the playbook.

The gap between firms that run surface-level technical checks and firms that run full AI diligence isn't 2x. It's the difference between owning a platform and owning a prompt. The gap between having an AI playbook and not having one isn't incremental. It's the gap between Thoma Bravo's Medallia and walking away from 600 lines of glue code.

Invest in the capability. Build the playbook. Don't stop improving it until every deal in your pipeline has been scored.

Alright Mark, just drop the Limestone Digital technical diligence pitch. We all know you wrote this as an X article because the algorithm is boosting long-form and you embed AI engineers into PE firms.

Fine. The takeaway is:

Stop looking at the spreadsheet and start looking at the codebase. What is the actual technology behind the ARR, and can someone rebuild it in a weekend?

Run the four-axis scorecard on every AI deal: model dependency, data moat, talent concentration, substitution risk. Be honest about whether your current diligence process covers any of it. If it doesn't, add the technical workstream. Invest in evaluation: your scorecard, your team, your process. Score it before you sign it. Don't close early on strong metrics alone.

The spreadsheet shows ARR and growth. It doesn't show whether the company built what it's selling.

The diligence is not the friction. It's the filter.

Point that filter at the codebase, and the real valuation follows.

Every deal I've worked where the firm ran full technical diligence ended one of two ways: they walked away knowing why, or they closed knowing what they bought.

Three years ago we started opening codebases behind pitch decks that said "proprietary AI."

We're still finding wrappers.

P. S. If you're actually interested about what we do at @LimestoneHQ, feel free to explore our website. We embed AI-native engineers into existing codebases. Half of our clients are PE-backed.

https://limestonedigital.com or Book a Call

Mit einem Klick speichern

Virale Artikel mit YouMind per KI tief lesen

Speichere die Quelle, stelle gezielte Fragen, fasse die Argumentation zusammen und verwandle einen viralen Artikel in wiederverwendbare Notizen in einem einzigen KI-Arbeitsbereich.

YouMind entdecken
Für Creator

Verwandle dein Markdown in einen sauberen 𝕏-Artikel

Wenn du eigene Langtexte veröffentlichst, wird die 𝕏-Formatierung von Bildern, Tabellen und Codeblöcken mühsam. YouMind macht aus einem ganzen Markdown-Entwurf einen sauberen, sofort postbaren 𝕏-Artikel.

Markdown zu 𝕏 testen

Mehr Muster zum Entschlüsseln

Aktuelle virale Artikel

Mehr virale Artikel entdecken