Developer productivity metrics for AI-assisted teams.
No single number measures developer productivity: teams start from DORA for delivery, SPACE for five human dimensions or the DX Core 4. On an AI-assisted team, add what each engineer's coding agents cost and whether their sessions end in merged pull requests.
13/14 AI Adoption55 PRs Merged10.3h Median AI Time
13/14
AI Adoption▲ 3
55
PRs Merged▼ 8
10.3h
Median AI Time▼ 3.5h
57.2h
Build Hours▲ 20.3h
Agents Ran
Planning43
Implementation49
60300
Aug 14Aug 21Aug 28Sep 4
AI Plan Stack
Top Dev46
p256
80400
Aug 14Aug 21Aug 28Sep 4
About this sample
Sample team week · Earlier trend points are illustrative. *Model-estimated activity, including overlapping automation; not engineer hours. Deltas use the email’s unrounded totals.
↓ Scroll to explore
USED BY ENGINEERING TEAMS AT
What are developer productivity metrics?
Developer productivity metrics describe how well an engineering team turns work into software that ships. No one number captures it: the authors of the SPACE framework write that productivity “cannot be measured by a single metric or dimension.” The frameworks most teams start from:
DORA: software delivery performance (change lead time, deployment frequency, failed deployment recovery time, change fail rate and deployment rework rate).
SPACE: five dimensions from GitHub and Microsoft Research (ACM Queue, 2021). They are satisfaction and well-being, performance, activity, communication and collaboration, and efficiency and flow.
DX Core 4: DX's framework that encapsulates DORA, SPACE and DevEx in four dimensions (speed, effectiveness, quality and business impact).
How AI coding changes what to measure
Coding agents make a change cheaper to write, so activity counts rise whether or not more ships. And how fast AI feels is a poor guide. In METR's randomized study of 16 experienced open-source developers working on 246 issues (July 2025), developers took 19% longer when allowed to use AI tools, and afterwards still believed AI had sped them up by 20%.
In February 2026 METR said its follow-up could no longer measure the effect reliably: 30% to 50% of developers said they held back tasks they didn't want to do without AI. Its raw estimate for newly recruited developers was 4% less time per task, anywhere from 15% less to 9% more.
DX's AI Measurement Framework measures AI code assistants and agents on three dimensions: utilization, impact and cost. DX measures them with system data and surveys together; Zest supplies the system side for AI coding sessions.
How to measure AI coding impact per engineer
Zest is AI coding analytics: its plugins and extension report each session from the coding agents your engineers run, and the Zest GitHub App joins those sessions to what merged.
Utilization: which coding agents, models, skills and MCP tools each engineer uses, across Claude Code, Cursor, Codex, GitHub Copilot Chat, Grok Build and Hermes.
Cost: input and output tokens per session, with an estimated cost from model list prices.
Impact: the pull request each session led to, and PR throughput (merged pull requests per week, averaged over four weeks).
Speed: PR cycle time (median and p75 hours from opened to merged, week over week), through Ask Zest or Zest's MCP server.
From AI coding metrics to a better next week
Each evening Zest writes every engineer's standup from that day's sessions and pull requests, so nobody types a status update.
Once a week (Fridays by default) the Team Standup lists the AI stack worth sharing and up to three moves to ship more, so the practices behind one engineer's merged work can spread to the rest of the team.
What Zest does not measure
Zest doesn't calculate DORA metrics or run developer-experience surveys. It doesn't ingest deployments or incidents, and it doesn't capture Copilot inline completions. Keep your CI/CD, incident and survey tools for delivery and sentiment, and use Zest for the AI coding work that feeds them.
TRY IT ON YOUR TEAM
See what your team ships with AI.
Starter is free (3 seats, no credit card). Pro is $49 a month with 5 seats and Max is $149 a month with 10 seats — priced per workspace, not per engineer, with extra seats at $19 a month.
What are the most important developer productivity metrics?+
There isn't one. DORA's metrics cover software delivery; SPACE adds satisfaction, collaboration and flow; the DX Core 4 groups them into speed, effectiveness, quality and business impact. On an AI-assisted team, add AI utilization, cost and impact per engineer.
How do you measure AI's impact on software development, and its ROI?+
Measure the work, not the impression of it: which AI tools and models each engineer uses, what each session costs, and whether it ends in a merged pull request. Those are the two sides of any AI coding ROI number.
What does a 10-engineer team pay?+
$144 a month on Pro, for the whole workspace. Pro adds Team statistics, PR Throughput, Session & Token intelligence; Max (10 seats included, $149) adds Agentic product management, AI Feature + Ticket analytics, Cloud agent + Hermes.