← All insights

AI in deals, part 2 of 8

What an AI build actually cost, and what independent analysis showed.

Cleaner than two team-built codebases. Still needs people on it.

I ran static analysis on my own AI build, using the same tools I would point at a target. Here is what came back.

I’ve spent years assessing software development, DevOps and R&D financials on deals, but I’m not a developer. I wanted to see the work from the other side.

So I built a production application by directing an AI coding assistant. Not a single tool but a chain: diligence report, then the IC paper, then that diligence carried into integration or carve-out planning, then synergies tracked against it. A wrong read in diligence doesn’t surface in diligence. It surfaces as a synergy that never lands, long after anyone traces it back.

932 commits. 189 pull requests. 3,122 passing tests. And 42% of every line changed was rules, documents, enforcement or tests. Even in a domain I know well, that specification wasn’t sitting there to reuse. It had to be written.

The analysis compared the code against two open-source codebases built by teams, at the same size.

Line by line, mine was cleaner than both: fewer security issues, less than half the lint warnings, less duplication. So code quality is not why it still couldn’t be trusted without a person on it.

Structurally, it was worse: five times the lines per file, ten times the complexity, four files doing too much. Tests were 24% of the code against their 42% and 40%, on a build whose own rules set the test gates explicitly.

Comparison table. This build: 1 person, 115 person-days, 110 security issues per 10k lines, 144 lint warnings, 58% near-duplicate functions, 241 average lines per file, median complexity 10, 4 files doing too much, 24% test code. Team build A: about 150 people, 1,200 person-days, 137, 726, 70%, 45, 1, 0, 42%. Team build B: about 80 people, 2,900 person-days, 225, 311, 64%, 49, 1, 0, 40%.
Static analysis. Same tools, same configuration, comparable codebase size. Directional, not a benchmark.

I’ve diligenced this before. It’s the in-house app built by a self-taught developer, just produced faster.

Reliable at the level a check can see. Unreliable at the level nothing checks. The thing implementing your specification is also the thing deciding when it has been met.

This isn’t one build. Faros AI tracked over 10,000 developers: teams using AI heavily merged 98% more changes, spent 91% longer in review, and delivered no faster at company level. The work moved from writing code to checking it.

So in diligence, “how much of your code is AI-written?” is the wrong question. Ask whether delivery speed and failure rates improved after adoption, and who reviews the output.

Then price that reviewer. My build needed product, architecture and QA, and finishing it needs an engineer on structure. In a target’s plan those are permanent roles: someone who reads the output in the function that owns the risk, and someone who keeps the rules and checks current when the model underneath changes. One at $180K loaded is about $2.2M of enterprise value at 12x, in a case that usually counts model cost only.

The cost is never just AI. It’s AI plus the people who check it.

Pricing AI in a deal you’re working on? Tell me about it, or read more insights.