DORA 2026 Report Summary: ROI of AI-Assisted Development

Key Takeaways (TL;DR)

  • The J-Curve: DORA's 2026 ROI report models an initial productivity dip before AI pays off; its sample calculator assumes a 15% dip for three months, driven by learning and the "verification tax" (Google DORA, 2026).
  • Productivity Disparity: Stanford's study of nearly 100,000 developers found AI gains of 30–40% on simple greenfield tasks but only 0–10% on complex work in existing codebases (Stanford, 2025).
  • PR Review Bottleneck: LinearB's analysis of 8.1 million PRs reveals AI-generated code waits 4.6x longer for first review, expanding code review queues (LinearB, 2026).
  • Measured vs Perceived Speed: METR found experienced developers were 19% slower with early-2025 AI tools while believing they were 20% faster. Its 2026 follow-up points to a speedup, but the result is not statistically significant (METR, 2025–2026). See my breakdown of the METR data.
  • System First: AI accelerates existing organizational habits; fix review pipelines, automated testing, and deployment CI/CD before adding generation tools.

Google's DORA team published The ROI of AI-assisted Software Development (Google DORA, 2026). In my technical advisory work with scale-ups, I observe that AI coding assistants amplify underlying team habits. However, if an organization suffers from slow code reviews or fragile test suites, AI acceleration compounds technical debt rather than reducing it. Consequently, fixing engineering pipelines prior to rolling out AI tools is essential for positive ROI.


What does the Google DORA 2026 AI report reveal about engineering ROI?

DORA 2026 AI Report refers to Google DORA's The ROI of AI-assisted Software Development, a framework and calculator for estimating the return on AI coding tools, built on a decade of DORA research into software delivery performance (Google DORA, 2026).

+------------------------------------------------------------------------------------+
|                         AI PRODUCTIVITY BENCHMARKS                                 |
+--------------------+--------------------------------+------------------------------+
| Work Type          | Measured Effect (Source)       | Primary Obstacle             |
+--------------------+--------------------------------+------------------------------+
| Simple Greenfield  | +30% to +40% (Stanford)        | None (Low verification tax)  |
| Complex Existing   | 0% to +10% (Stanford)          | Context & Architecture Debt  |
| Code Review Queue  | +91% review time (Faros AI)    | Senior Reviewer Bottleneck   |
+--------------------+--------------------------------+------------------------------+

Stanford's data shows AI accelerates greenfield coding but yields 0–10% gains on complex work in existing systems. DORA's report explains why the average hides so much: early gains are eaten by learning time and the verification tax, the effort spent checking that AI-generated code is correct, secure, and fits the architecture. Furthermore, generating more code without automated verification increases maintenance overhead and lead time to production. In my experience, technology leaders who measure lead time for changes before deploying AI tools avoid common adoption pitfalls. As a result, engineering teams that establish automated CI testing prior to agent deployment capture maximum velocity improvements.

Citation Capsule: Google DORA AI ROI Benchmark


Why is code classified as a technical liability rather than an asset?

Software engineering research has long estimated that maintenance accounts for well over half of total software lifecycle cost, far more than initial development.

[Code Generation] ----> (Unverified PR Volume) ----> [Lifetime Maintenance Debt]

Generating additional code without robust verification scales operational liabilities. Specifically, when building systems under regulatory frameworks like GDPR or the EU AI Act, unverified code expands security attack surfaces and compliance risk. As a result, review the guide on GDPR LLM RAG architecture traps for more details.


What breaks when engineering teams add AI coding assistants to broken systems?

Faros AI's telemetry from more than 10,000 developers across 1,255 teams found that high-AI-adoption teams merged 98% more PRs, while PR review time rose 91% (Faros AI, 2025). LinearB's analysis of 8.1 million pull requests across 4,800 teams adds that AI-generated PRs wait 4.6x longer for a first review (LinearB, 2026).

+------------------------------------------------------------------------------------+
|                         AI SYSTEM BREAKDOWN PATTERNS                               |
+--------------------+---------------------------------------------------------------+
| System Layer       | Breakdown Mechanism                                           |
+--------------------+--------------------------------+------------------------------+
| 1. Code Review     | AI code waits 4.6x longer for initial human review            |
| 2. CI/CD Pipeline  | Flaky integration tests cause deployment failures to spike    |
| 3. Architecture    | Unenforced patterns lead to codebase drift and entropy        |
+--------------------+---------------------------------------------------------------+

Specifically, unverified AI code overwhelms senior reviewers. To fix this, engineering leaders must adopt the shift-up engineering model to move product context into machine-consumable repository schemas.

AI's Impact on Engineering Metrics Bar chart showing AI impact on engineering metrics based on Faros AI and LinearB data. AI's Impact on Engineering Metrics PRs merged +98% Review time +91% Wait for 1st review 4.6x longer Sources: Faros AI AI Productivity Paradox (2025); LinearB 2026 Software Engineering Benchmarks
Generation scales. Absorption doesn't.

The same pattern shows up in Linear's workspace data: coding agents tripled weekly pull requests, yet teams ended up working more, not less. I break that down in the AI Jevons Paradox.


How does the AI productivity J-Curve affect engineering teams?

AI Productivity J-Curve refers to the initial drop in delivery performance during AI tool adoption before net gains materialize. DORA's ROI calculator uses a 15% dip lasting three months as its sample default (Google DORA, 2026).

[Tool Adoption] ----(Dip: learning + verification tax)----> [Recovery] ----> [Net Gains, if the system absorbs them]

Controlled studies show how hard the recovery is to measure. METR's randomized trial found experienced open-source developers took 19% longer with early-2025 AI tools, while believing they were 20% faster (METR, 2025). A follow-up experiment from late 2025 estimated an 18% speedup for returning developers, but with a confidence interval of -38% to +9%; METR itself calls the data unreliable because 30–50% of developers skipped tasks they did not want to do without AI (METR, 2026). I break down what METR's 2026 data actually shows in a separate post.

The AI Productivity J-Curve Illustrative line chart of the AI productivity J-Curve described in DORA's ROI report. Not plotted from measured data. The AI Productivity J-Curve Illustrative, based on the J-curve model in DORA's 2026 ROI report
Both lines dip. Only one recovers. The difference is engineering foundation strength.

In my experience, teams that fix CI/CD pipelines early recover faster. If your PR queue has grown since you rolled out coding assistants, the tool isn't the problem; the review and delivery system around it is. To put a number on the dip for your own team, see how to calculate AI coding assistant ROI. That system is what I fix as an interim or fractional CTO: book a 30-minute teardown and bring your review latency and deployment numbers. For more on how AI shifts engineering work, read the AI Jevons Paradox and Shift-Up Engineering.

Updated 30 September 2026: corrected source attributions for the Stanford, Faros AI and METR figures, and the description of METR's 2026 follow-up.


Frequently Asked Questions

Does AI make software engineering teams faster?

It depends on the work. Stanford's study of nearly 100,000 developers found 30-40% gains on simple greenfield tasks, but only 0-10% on complex work in existing codebases.

What is the DORA J-Curve?

The DORA J-Curve describes the temporary drop in team velocity during AI tool onboarding before workflows adapt and net productivity gains emerge.

What should engineering leaders fix before deploying AI coding tools?

Engineering leaders should streamline code review pipelines, stabilize CI integration test suites, and establish automated deployment templates.

Why does AI-generated code wait longer for review?

LinearB data shows AI-generated code waits 4.6x longer for first review because human senior developers must verify unfamiliar code blocks without clear context.