THE KEY ANSWER

Measure the time from task start to a working change and the quality after deployment. Compare similar tasks and account for review, fixes, and tool learning. Do not extrapolate the results of a single study to every team.

01

What do studies say, and what do they not settle?

The METR study published in July 2025 covered 16 experienced developers and 246 tasks in open-source projects familiar to them. Under the tested conditions, using contemporary AI tools increased work time by an average of 19%. This is the result of a specific experiment, not a universal assessment of all tools available in 2026.

It is worth treating this as an argument for measurement, rather than for the simple thesis that “AI helps” or “AI harms.” A different task, user experience, repository, and tool version can change the outcome. A feeling of acceleration is also insufficient. A team needs data on its own work and a clear distinction between perception and a completed change.

Context and references: METR: Measuring the Impact of Early-2025 AI

02

Measure the entire delivery flow

Record the time for implementation, waiting for review, fixes, testing, and deployment. If code is created faster but the review queue grows, the bottleneck has shifted to another place. It is also worth looking at the size of changes and the number of returns to the same problem. The number of accepted suggestions does not answer these questions.

A good subject for comparison is task classes: bug fix, new form, dependency migration, analysis of an existing module. Do not compare simple screens built with AI to difficult incidents handled without it. Note the context and limitations of the trials. The goal is a decision about the way of working, not a ranking of individuals.

03

Where might hidden costs arise?

Generated code must be understood, checked, and later maintained. Long changes make assessment difficult, and solutions ill-suited to the architecture create further exceptions. Establish what a small, reviewable portion of work should look like. A tool can help prepare a test or explain a module before it starts rebuilding it.

Also check the quality of the task description. A lack of acceptance criteria causes further iterations regardless of who writes the code. A good experiment might be improving the context and instructions in the repository, rather than purchasing another subscription. Document architectural decisions so that they are useful for both humans and tools.

04

How to conduct a fair team experiment?

Choose comparable tasks, establish usage rules, and allow time for learning. Collect quantitative results and short notes on where the tool helped or hindered. Check the result after deployment, because a maintenance issue may appear later than the savings during writing.

Based on this, establish specific practices: what the team uses AI for, what always requires control, and how changes are divided. Do not assume identical benefit in every area. Repeat the assessment after a significant change in the tool or process. The most useful outcome is a better way of delivering the product, not an impressive percentage on a slide.

WHERE TO START

Bring this into your project.

  • Compare similar task classes.
  • Include review, fixes, and deployment.
  • Observe quality and maintenance after the change.
  • Treat measurement as process improvement, not as an evaluation of people.

Choose one thing your process is missing today. It's a useful topic for your first conversation with the team.

QUESTIONS AND ANSWERS

Frequently asked questions.

Is the number of lines of code a good metric?

No. More code may mean higher maintenance costs. More important is completing the task with expected quality and without unnecessarily increasing complexity.

Does the METR result mean that AI is slower?

Only in the described conditions of the study from early 2025. It does not prove this for all teams, tasks, and later models. Therefore, the article proposes measurement in your own environment.

Sources and context

Prepared by the ALGOV team. Current as of September 8, 2026. Examples describe possible scenarios, not results from client projects. How we create our guides.

YOUR SITUATION IS UNIQUE

Let's put these insights to work.

Describe the task, your data and what gets in the way today. Together we'll decide which first step can test the solution's value.

Discuss your idea ↗Explore our service: Product development and support