I ran static analysis on my own AI build, using the same tools I would point at a target. Here is what came back.
I’ve spent years assessing software development, DevOps and R&D financials on deals, but I’m not a developer. I wanted to see the work from the other side.
So I built a production application by directing an AI coding assistant. Not a single tool but a chain: diligence report, then the IC paper, then that diligence carried into integration or carve-out planning, then synergies tracked against it. A wrong read in diligence doesn’t surface in diligence. It surfaces as a synergy that never lands, long after anyone traces it back.
932 commits. 189 pull requests. 3,122 passing tests. And 42% of every line changed was rules, documents, enforcement or tests. Even in a domain I know well, that specification wasn’t sitting there to reuse. It had to be written.
The analysis compared the code against two open-source codebases built by teams, at the same size.
Line by line, mine was cleaner than both: fewer security issues, less than half the lint warnings, less duplication. So code quality is not why it still couldn’t be trusted without a person on it.
Structurally, it was worse: five times the lines per file, ten times the complexity, four files doing too much. Tests were 24% of the code against their 42% and 40%, on a build whose own rules set the test gates explicitly.