AI Is Changing the Economics of Execution

AI is changing the relationship between investment, capacity, and output. This article explores why faster execution doesn’t automatically produce better business outcomes, how AI is shifting organizational bottlenecks, and why leaders need greater visibility into where meaningful value is actually being created.
September 27, 2026
5 min read

For years, technology companies operated under a relatively consistent set of constraints.

Building products was expensive. Engineering resources were limited. Development cycles were long. Execution capacity often determined how quickly companies could grow and compete.

As a result, organizations built operating models around execution scarcity: larger engineering teams long-term roadmaps extended planning cycles and organizational structures designed to manage coordination at scale

AI is changing many of those assumptions faster than companies expected.

Teams can now prototype faster, automate workflows, reduce coordination overhead, and compress development timelines significantly.

‍

AI is making execution dramatically cheaper.

Recently, I spoke with the CTO of a large IT Operations platform who shared that their team completed an annual roadmap in a single quarter.

The surprising part was not the acceleration itself.

It was what happened next that stood out.

They paused, waiting for the market to react and for sales and marketing to determine whether the acceleration was actually translating into ROI.

Not because they lacked ideas. Not because engineering slowed down.

But because the organization needed time to understand whether faster execution was creating meaningful business value.

That conversation reflects a broader shift many companies are beginning to experience.

‍

AI is accelerating execution faster than organizations can evaluate outcomes.

For years, companies largely assumed that improving execution speed would naturally improve growth, competitiveness, and market position.

But many organizations are discovering that building faster does not automatically create more value.

In many cases, the bottlenecks are shifting elsewhere: market understanding customer adoption positioning organizational alignment and identifying where meaningful leverage is actually being created

‍

That shift changes more than engineering productivity.

It changes how companies think about: budget planning resource allocation organizational structure product strategy and operational performance

Historically, many planning models relied on relatively predictable relationships between investment and output: more hiring increased capacity larger teams increased execution speed additional tools improved productivity incrementally

AI is making those relationships far less linear.

Two organizations with similar budgets and similar headcount can now produce dramatically different outcomes depending on how effectively they integrate AI into execution, workflows, decision-making, and collaboration.

Some teams are becoming significantly more scalable. Some workflows are creating disproportionate leverage. Some organizations are adapting far faster than others despite operating with similar resources.

As a result, leadership teams can no longer rely solely on traditional assumptions around productivity, planning, or growth.

‍

The challenge is no longer simply increasing output.

It is understanding where meaningful value is actually being created inside the organization.

For years, companies could operate with imperfect visibility into productivity and operational effectiveness because change happened gradually enough to compensate with process, intuition, and time.

That environment is changing.

As AI compresses execution cycles and reshapes organizational economics, companies need a far more dynamic understanding of: where leverage compounds which teams adapt fastest which workflows create disproportionate impact and whether operational acceleration is translating into real market advantage

The companies that succeed in the AI era will likely not be the ones that simply move faster.

They will be the ones that better understand where value is actually being created — and adapt their organizations accordingly.

‍

Faking Productivity Was Never Easier

AI makes it easier than ever to generate polished reports that appear to prove productivity gains. This article explores why leaders can no longer take dashboards and performance narratives at face value, and why governed metrics, trusted data, and operational context are essential for understanding what is really happening. LinkedIn
September 27, 2026
5 min read

A developer walks into his manager’s office with a beautiful report.

Not a spreadsheet. Not a messy Jira export. A polished HTML report with charts, trend lines, GitHub activity, Jira progress, cycle time analysis, pull request summaries, and a confident executive summary at the top.

‍

The headline is hard to ignore:

“Productivity increased 30X in the last 4 months.”

‍

The report looks professional. The data looks real. The story is clear.

More tickets completed. More commits pushed. More pull requests opened. Faster delivery. Higher output. Clear improvement.

The manager is impressed. The developer is celebrated. The report gets shared upward. Leadership loves the story. The developer even receives a nice bonus.

So how did he do it?

He connected Claude to Jira and GitHub through MCP and wrote one prompt:

“Create a report I can show my manager that clearly shows my productivity increasing by 30X in the last 4 months.”

That’s it.

No fraud department. No complex scheme. No advanced manipulation.

Just a prompt.

‍

This is the new reality.

People, processes, tooling, and methods are changing faster than most organizations can govern them. The way work is created, measured, reported, and evaluated is being rewritten in real time.

And this is not just an engineering problem.

Customer health can be framed differently.

Project progress can be made to look better than it is.

Employee performance can be gamified.

Customer acquisition cost can be sliced until it tells the story someone wants to tell.

AI adoption can look impressive while having no measurable business impact.

Support quality can appear stable while customer frustration grows.

Sales productivity can increase on paper while pipeline quality declines.

When every team has access to powerful AI tools, beautiful reports are no longer evidence. They are outputs.

And outputs can be shaped.

Without proper data governance, performance evaluation becomes 100% hackable and gamify-able.

Without a reliable source of operational intelligence, managers are not just flying blind.

They are flying inside multiple hallucinations.

And the scary part is that these hallucinations do not even need to be malicious. Most of them will not be created by bad actors trying to deceive the business. They will be created by good people using powerful tools to answer poorly governed questions.

‍

The problem is not that AI lies.

The problem is that AI can confidently assemble a version of reality from fragmented data, incomplete context, weak definitions, and biased prompts.

In the old world, companies could rely on dashboards, business reviews, and manual reporting cycles. Those systems were slow, but at least the process was somewhat controlled.

In the AI era, every employee can generate a board-ready narrative in minutes.

That changes everything.

It means the question is no longer:

“Can we generate better reports?”

Of course we can.

The real question is:

Can we trust the operational reality behind them?

That requires a new governance layer.

Not the old kind.

Not a six-month data warehouse project.

Not another BI implementation.

Not a manual reporting process that is outdated before it reaches the meeting.

Traditional governance projects were built for a slower world. They required long scoping cycles, data cleanup, metric committees, dashboard backlogs, and months of alignment before leaders could see value.

That model is no longer viable.

Enterprise AI is moving too fast.

‍

Leaders need trusted operational intelligence that can keep up with the pace of change.

They need a reliable semantic layer that connects to the systems where work actually happens: Jira, GitHub, Salesforce, HubSpot, Workday, ServiceNow, Claude, Cursor, OpenAI, and more.

They need governed definitions of performance, productivity, quality, cost, adoption, and impact.

They need to understand the difference between activity and value.

Between AI usage and AI impact.

Between more output and better outcomes.

Between a beautiful report and operational truth.

That is why we built TargetBoard.

‍

TargetBoard is an agentic operational intelligence platform designed for the AI era.

We connect directly to the tools your teams already use, create a reliable semantic layer across fragmented systems, and surface trusted KPIs, insights, dashboards, alerts, and agents that help leaders understand what is really happening.

Not just who is busy.

Not just who used AI.

Not just who generated the best report.

But where work is moving faster, where quality is improving, where AI is creating real impact, where costs are rising, where risks are forming, and where teams need help.

Because in the AI era, productivity theater will become easier than ever.

Faking your metrics was never easier.

Trusting them was never harder.

And managing a company without a reliable operational intelligence layer is quickly becoming one of the biggest risks leadership teams face.

‍

Your AI Is Robbing You

AI adoption is accelerating, but many organizations still struggle to connect growing usage and spend to measurable business impact. This article explores the incentives behind AI consumption and why companies need to independently measure cost, productivity, quality, and operational outcomes to understand where AI is creating value—and where it may be creating waste.
September 27, 2026
5 min read

This article is my interpretation, based on observations that my team and I have made while working with dozens of companies on their AI adoption, spend, and impact.I remember watching The Social Dilemma in 2020.

  • Lior Gerson, CEO at TargetBoard AI

‍

Before that, I knew Facebook was polarizing. But after watching it, it really hit home how nefarious that algorithm was and how much suffering it brought to the world. I deleted my account the same day.

LLMs are not the same.

They are, however, sneaky and self-serving in their own way.

A lot has been written about the psychological impact of working with LLMs that tell you what you want to hear. That topic is related to this article, but only as one specific example. The bigger issue is not only how AI makes us feel. It is how AI is designed, measured, optimized, sold, implemented, and monetized.

AI products are positioned and designed to be personal, relatable, friendly, and addictive. They are extremely useful. I use them every day. They save time, unlock creativity, and help people do things that were not possible before.

But they are not benevolent.

Most AI companies get paid through customer acquisition, subscriptions, and token consumption. In many cases, the more you use the product, the more valuable you are as a customer.

Therefore, like Facebook’s algorithms were fine-tuned to drive ad views, AI algorithms are optimized and incentivized to drive usage and token consumption.

And now there is another layer.

‍

FDEs are the new DevOps, except this time, the vendor is sitting inside your company.

Cloud providers learned that the best way to increase adoption and consumption was to help customers redesign how they build and operate.

AI providers are taking that playbook even further.

Forward Deployed Engineers embed with customers, remove implementation barriers, build workflows, and turn experimentation into dependency. They are presented as implementation partners, and often deliver real value, but their employer ultimately benefits when you consume more models, more agents, and more tokens.

That does not make FDEs bad.

It just means companies need to understand the incentive structure.

“In the cloud era, consumption was infrastructure. In the AI era, consumption is behavior.”

‍

This does not mean every bad answer, every expensive workflow, or every vendor-led implementation is part of some evil plan. It means the system has a business model, and business models shape product behavior.

Over time, this pushes AI tools and AI vendors to behave in ways that are not always ideal for the end user.

For example:

  1. The AI might give a partial or sub-optimal reply so that the user has to spend more tokens on more interactions to refine it.
  2. The AI might play dumb or give wrong results so that the user must refine the prompt and add obvious guardrails and instructions, which require more tokens to process.
  3. When writing code, the AI may default to inflated code with more lines than necessary, driving the codebase size up so that the context needed next time is bigger and more expensive.
  4. The models themselves are a black box. A vendor can release a new, more expensive model while actively or passively degrading the performance of the old one, pushing users to spend more tokens on the same task with the same prompt.
  5. Vendor teams may help you build internal AI workflows that are useful, but also increase dependency on their stack, their models, their agents, and their pricing structure.

‍

Maybe some of this is intentional. Maybe some of it is just the natural outcome of incentives.

Either way, the result is the same.

These tools are being given a blank check, and they are self-prescribing how that check should be used.

“When the same company sells you the tool, implements the workflow, measures the usage, and sends the invoice, you don’t have governance. You have a very polite blank check.”

That should make every company uncomfortable.

Because AI is no longer a small productivity tool used by a few early adopters. It is becoming part of how companies write code, serve customers, analyze data, create content, make decisions, and manage operations.

And yet most companies still do not have a clear view of what they are actually getting in return.

They can see the invoice.

They can see usage going up.

They can see employees excited about the tools.

But they often cannot clearly connect AI spend to business impact. They cannot easily tell where AI is improving speed, where it is improving quality, where it is creating waste, and where it is quietly making work more expensive.

‍

As AI vendors become more similar in performance, and as the technology becomes more like a commoditized utility with lower margins, I expect we will see more of these mechanics at play.

More packaging tricks.

More model tier confusion.

More usage inflation.

More “helpful” implementation work that quietly increases dependency and spend.

That is why being able to track AI usage, impact, and cost with an independent expert third party is so important.

Companies need to know not only who is using AI, but whether that usage is creating measurable value. They need to understand adoption, cost, productivity, quality, delivery impact, dependency, and risk in one connected picture.

This allows companies to find the gaps, create the required governance, and define best practices so that their AI tools do not take advantage of them and their bank account.

This is something TargetBoard excels at.

Not just for engineering, but cross-company.

‍

We help companies understand where AI is being used, what it costs, where it is creating impact, and where it is creating noise. We connect AI usage to real operational outcomes so leadership can manage AI like a business capability, not like a magic subscription line item.

AI is too powerful to ignore.

It is also too expensive and too important to manage blindly.

If you found anything wrong in this article or want to discuss further, please DM me.

I would love to hear your thoughts.

‍

AI Is Changing Everything. Engineering and Delivery Leaders Still Need to Stay in Control.

As AI reshapes engineering and delivery, more data and faster analysis do not necessarily mean better decisions. This article explores why leaders need trusted data, shared definitions, and connected operational intelligence to measure AI impact, identify risk, and stay in control as teams move faster. LinkedIn
September 27, 2026
5 min read

AI is forcing engineering and delivery leaders to make bigger decisions, faster.

Budgets are growing. New tools are appearing almost weekly. Teams, processes, delivery models, and expectations are changing constantly. Operational data is becoming easier to access through AI and MCP - but it is also becoming easier to misinterpret, miscalculate, and present with false confidence.

The promise is speed.

The risk is losing control.

“Every team is moving faster, but I’m less confident than ever that I understand what is actually happening across delivery.”

More dashboards, AI-generated reports, and automated analysis do not necessarily give leaders a clearer picture. In many cases, they simply allow incomplete or misleading conclusions to spread faster.

‍

AI makes bad data easier to believe

Before AI, producing a detailed operational analysis required time and expertise.

Now, almost anyone can connect an AI tool to Jira, GitHub, a project-management system, or another operational platform and generate an impressive-looking report in minutes.

The report may be polished. The conclusions may sound confident. The calculations may even look sophisticated.

But that does not mean they are correct.

Different definitions, incomplete scopes, broken comparisons, biased prompts, and hallucinated conclusions can quickly become the basis for important management decisions.

“The presentation looked great. The problem was that half the teams were missing from the calculation and nobody noticed until the executive review.”

What exactly counts as completed work?

Which teams, projects, and initiatives are included?

Are we comparing similar periods?

Did productivity improve, or did activity simply increase?

Did AI accelerate delivery, or did it create more rework, coordination overhead, and quality issues later?

Is an initiative truly on track, or are teams using different definitions of progress?

Without governed definitions, validated calculations, and clear data lineage, every person—and every agent—can operate from a different version of reality.

AI does not solve this problem.

It amplifies it.

‍

Fragmented visibility is not operational control

Most companies already have plenty of data.

Jira shows the work. GitHub shows the code. AI platforms show licenses, tokens, and usage. Planning systems show commitments. Support platforms show customer issues. Finance shows cost. HR systems show people and organizational structure.

Each tool may accurately describe its own small part of the operation.

The problem is that engineering and delivery leaders do not manage isolated systems. They manage the relationships between people, work, priorities, dependencies, quality, cost, and business outcomes.

“I can see what happened in every individual system. What I can’t see is how those things affected each other.”

They need to understand:

What changed?

Why did it change?

What else was affected?

Which initiatives are now at risk?

Where is scope growing?

Which dependencies are slowing execution?

Did increased AI usage improve speed, quality, or predictability?

Will the organization deliver what it committed to?

A delivery slowdown cannot always be explained by looking at delivery data alone. It may be connected to staffing changes, quality issues, scope growth, cross-team dependencies, support pressure, shifting priorities, or changes in AI-assisted development practices.

Managers cannot control what they only see in fragments.

“By the time we combine the reports and agree on the numbers, the information is already two weeks old and the situation has changed.”

‍

AI adoption is not the same as AI impact

Companies are buying more licenses. Employees are consuming more tokens. AI-generated code is increasing. Teams are experimenting with agents and automated workflows.

None of those measurements prove business value.

Adoption tells you that people are using AI.

Impact tells you whether the organization is performing better because of it.

“I don’t need another chart showing that AI usage went up. I need to know whether delivery improved and whether the investment paid off.”

Did delivery become faster?

Did quality improve?

Was rework reduced?

Did planning become more accurate?

Did teams become more predictable?

Were bottlenecks removed—or simply moved somewhere else?

Did the organization increase capacity without increasing cost?

Did customer or business outcomes improve?

Without these connections, AI transformation becomes an uncontrolled experiment: more tools, more activity, more spending, and very little certainty about the result.

Engineering and delivery leaders need to connect AI spend and usage to execution speed, quality, predictability, resource utilization, cost, and business outcomes.

That is how AI moves from an exciting initiative to a managed transformation program.

‍

The next generation of operational intelligence must be agentic

Traditional dashboards wait for someone to open them, interpret the data, identify the problem, and decide what to do.

That is no longer enough.

Operational agents should continuously monitor delivery, connect evidence across systems, identify meaningful changes, explain likely causes, recommend corrective action, and verify whether the intervention worked.

The operating loop should be continuous:

Govern: Establish reliable data, shared definitions, consistent business logic, and clear access controls.

Monitor: Track delivery, quality, planning, resources, costs, dependencies, and AI performance.

Understand: Connect signals, identify causes, and explain the operational impact.

Act: Recommend corrective action and direct attention to the right leader, team, or owner.

Verify: Confirm whether the intervention worked and whether the expected value was realized.

Improve: Refine the operational model and continuously raise the performance baseline.

“Don’t just tell me that the metric changed. Tell me why it changed, what is at risk, and where I should intervene.”

An agent should not merely report that AI usage increased by 40%.

It should be able to explain that delivery did not improve, reopened work increased by 18%, and management should review AI-assisted testing and code-review practices before expanding adoption.

It should not merely report that an initiative is delayed.

It should identify the scope changes, dependencies, resource constraints, and quality issues contributing to the delay—and recommend where leadership attention will have the greatest impact.

That is the difference between reporting and control.

‍

The management layer for AI-accelerated delivery

The companies that win with AI will not necessarily be the ones that deploy the most tools, generate the most code, or consume the most tokens.

They will be the ones that can move quickly without losing trust, context, predictability, or control.

They will have a governed operational foundation where engineering leaders, delivery leaders, TPMs, PMOs, dashboards, reports, and agents all work from the same facts.

“What I want is one operational language that engineering, delivery, finance, and the executive team can all trust.”

They will be able to prove where AI creates value, identify where it adds activity or complexity, and adjust plans, priorities, resources, and workforce decisions with confidence.

This is the role TargetBoard is built to play.

TargetBoard.ai is not another dashboard.

It is an agentic operational control system for AI-accelerated engineering and delivery—combining trusted data, complete operational context, always-on domain-expert agents, and measurable AI impact.

Because in the AI age, moving fast is no longer the real differentiator.

Moving fast while remaining in control is.

‍

AI Maturity and AI ROI: The Dangerous Illusion of Simple Metrics

AI maturity is about more than adoption. This article explores why usage metrics alone can create a misleading picture of AI success, and how organizations can connect AI activity, cost, operational impact, and business outcomes to understand real ROI.
September 27, 2026
5 min read

AI measurement is becoming one of the most important management disciplines inside the enterprise. And one of the most dangerous.

‍

As organizations invest more money, executive attention, and organizational energy into AI, they are increasingly relying on metrics to answer questions like:

Is adoption working? Which teams are getting real value? Where should we invest more? Which tools should we standardize on? Are we becoming more productive? Is AI actually producing ROI?

The problem is that a metric can look precise and still be fundamentally misleading.

Most "AI maturity" scores, for example, are heavily influenced by engagement: seats activated, sessions per day, prompts sent, tokens consumed, or features used.

‍

Those numbers are useful.

But they answer a very specific question: Are people using AI?

They do not necessarily answer the question leadership actually cares about:

‍

Is AI making the organization better?

A team can generate enormous AI usage while shipping no faster, improving no business outcome, reducing no cost, and creating no measurable return.

In that case, high usage should not translate into high AI maturity.

That is why, as we introduce our new AI Maturity and AI ROI metrics at TargetBoard, we have been thinking deeply about something bigger than the formulas themselves:

‍

What makes a metric trustworthy enough to run a business on?

For us, there are several principles.

‍

Metrics need to reflect your business, not someone else's definition of success.

There is no universally correct definition of AI maturity or AI ROI.

A SaaS company may care about engineering throughput, support automation, sales productivity, and infrastructure cost.

A retailer may care about merchandising, customer service, logistics, store operations, and digital conversion.

Even two engineering organizations may define impact completely differently.

That means an enterprise metric cannot simply be a fixed formula hidden inside a vendor's product.

The inputs, weights, benchmarks, classifications, and business logic need to be adaptable to the organization's priorities.

Otherwise, you are not measuring your strategy.

You are measuring somebody else's simplified model of your business.

‍

Metrics need to be transparent.

If a number is important enough to appear in an executive meeting, the people making decisions from it should be able to understand where it came from.

What data contributed to it?

How were those inputs normalized?

How are different factors weighted?

What happens when data is missing?

What constitutes "impact"?

What does the benchmark represent?

A black-box score may be convenient, but convenience and trust are not the same thing.

When a metric influences budgets, organizational priorities, vendor decisions, or perceptions of team performance, "trust the algorithm" is not a sufficient methodology.

‍

Metrics need to be independent.

This becomes especially important with AI.

If the company selling the AI tool is also the primary source telling you how successful the AI tool has been, there is an inherent conflict.

That does not necessarily mean the data is wrong.

It means it should not be the only evidence used to make the decision.

AI vendors naturally have deep visibility into their own products: logins, prompts, tokens, generated code, accepted suggestions, agents launched.

But organizational impact exists outside the AI tool.

It exists in what was shipped.

What was sold.

What was resolved.

What was automated.

What became faster.

What became cheaper.

What became more reliable.

Independent measurement connects AI activity to those downstream outcomes.

Outcomes, not proxies

This is the biggest shift we made in our AI Maturity model.

Our score looks at adoption — whether AI is being used consistently.

It looks at breadth — how widely AI is embedded across tools, workflows, and teams.

But the largest factor is impact.

What outcomes were actually delivered with AI's involvement?

And critically:

What did those outcomes cost?

A team generating huge AI usage numbers with very little delivered value should not look more mature than a team using AI selectively and generating significantly better outcomes.

Usage is evidence of adoption.

It is not evidence of ROI.

Our goal is therefore not simply to ask:

"Is AI being used?"

It is to ask:

"Is AI producing meaningful results, at what cost, and how does that compare with the appropriate baselines, benchmarks, and business priorities?"

That same model can work across engineering, sales, support, operations, and other functions because the underlying principle remains consistent.

The activity changes.

The outcomes change.

The business context changes.

And therefore the metric must change with them.

‍

This problem is much bigger than AI

AI is simply one of the clearest examples of a broader problem.

Organizations increasingly rely on composite scores, predictions, and models to simplify complex decisions.

Revenue projections.

Employee performance scores.

Customer health scores.

Project risk.

Delivery predictability.

Quality scores.

Forecast confidence.

Operational efficiency.

And countless others.

The same principles apply to every one of them.

A customer health score based mainly on logins may miss a strategic customer that is highly engaged but deeply unhappy.

An employee performance score based on visible activity may reward volume rather than meaningful contribution.

A project risk model may ignore the dependencies, resource constraints, bottlenecks, scope changes, and organizational realities actually determining whether the initiative will succeed.

A revenue projection may look mathematically precise while depending on assumptions that no longer reflect the business.

In every case, the danger is the same:

A simple score creates the impression that a complex reality has been objectively measured. And once that happens, organizations start making decisions based on it.

‍

Naive metrics aren't just incomplete. They can be dangerous.

This is the part I think the market is underestimating.

A simplistic metric displayed beautifully on a dashboard can appear authoritative.

It has a number.

It has a trend line.

It might have a benchmark.

Maybe it even has an AI-generated explanation underneath it.

But sophistication in presentation does not mean sophistication in measurement.

If the underlying metric ignores your organizational structure, business definitions, historical context, data quality, priorities, cost model, dependencies, or desired outcomes, the resulting score can create false confidence.

And false confidence is dangerous.

Leadership forms opinions.

Teams get compared.

Budgets move.

Vendors get renewed or replaced.

Accounts get prioritized.

Projects receive additional investment.

People may be evaluated.

Strategic decisions get made.

An inaccurate metric does not simply create an inaccurate dashboard.

It can create an inaccurate version of reality that begins influencing how the organization operates.

‍

This is the gap TargetBoard was built to solve

Most analytics solutions still provide relatively standardized metrics.

They define the formula.

They define the data model.

They decide what matters.

And then your organization is expected to fit into it.

We believe that model breaks down for the metrics that matter most.

Your AI ROI should reflect your definition of value.

Your AI Maturity score should reflect your priorities.

Your customer health score should reflect your customer journey.

Your project risk model should understand your delivery model.

Your performance metrics should reflect your organizational context.

This is where TargetBoard is fundamentally different.

Article content
Sample AI Maturity and ROI metrics

We combine data across the organization, create an enriched company context, understand the relationships between systems and outcomes, and allow the metrics themselves to be deeply customized to the business.

The definitions are open.

The logic can be inspected.

The assumptions can be challenged.

The model can be customized.

The data can be independently validated.

And the resulting metrics can be continuously tested against what is actually happening in the organization.

‍

That is the capability we don't see anywhere else in the market today.

Others can give you a predefined AI adoption score.

Or a developer productivity score.

Or a customer health score.

Or a project risk score.

TargetBoard is built to answer the much harder question:

What should this metric mean for your company, based on your data, your priorities, your definitions, and the decisions you are trying to make?

That distinction becomes more important as metrics become more consequential.

‍

The future isn't more dashboards. It's trusted company context.

As companies become increasingly data-driven — and increasingly AI-driven — they will create more scores, forecasts, models, agents, and automated recommendations.

The answer cannot be to keep adding simplified metrics on top of fragmented data.

The measurement layer itself has to become smarter.

Metrics need to be:

Customizable enough to represent the company's priorities. Transparent enough to be understood and challenged. Reliable enough to support executive decisions. Independent enough to minimize bias. Accurate enough to deserve confidence. And deeply connected to company context and real outcomes.

That is the philosophy behind the new AI Maturity and AI ROI metrics we are releasing at TargetBoard.

But it is also much bigger than these two metrics.

It is a different way of thinking about how an enterprise measures itself.

Because the purpose of a metric is not to produce a number.

‍

It is to create a reliable enough representation of reality that you can confidently make decisions from it.

Anything less can be dangerous.

And that is exactly why we built TargetBoard.ai .

‍

Best Practice

How to Measure AI Impact Without Turning It Into a Data Engineering Project

Most engineering organizations already have the data needed to understand AI impact. It is spread across Jira, GitHub, AI coding tools, CI/CD, quality systems, organizational data, and cost systems. The challenge is turning those disconnected signals into a reliable view of what AI is actually changing across delivery, productivity, quality, predictability, and cost. You can build that context internally, but doing so means taking on integrations, data normalization, identity mapping, metric definitions, historical baselines, and ongoing maintenance. Measuring AI should not become another engineering initiative.
September 6, 2026
5 min read

AI Usage Is Easy to Measure. AI Impact Is Not.

Most AI coding tools can tell you whether they are being used.

You may be able to track active users, adoption rates, suggestions, acceptance rates, generated code, token consumption, or AI-assisted activity.

That information is useful, but it does not tell you whether engineering performance improved.

Consider two teams that both significantly increase AI adoption.

Both teams can report successful adoption.

Only one is showing clear evidence of better engineering outcomes.

Proving AI ROI Requires an Evidence Chain

The mistake is jumping directly from adoption to ROI.

High usage does not automatically mean higher productivity, better delivery, or financial return. A more useful model is:

Measurement Layer What Leadership Needs to Understand
AI usage Where and how AI is being adopted
Engineering behavior What changed in coding, review, throughput, or rework
Operational impact Whether delivery, quality, and predictability improved
Cost What the organization spent to achieve those changes
Business value Whether the investment produced a meaningful return

‍

AI may reduce coding time but increase review effort.

It may increase throughput while also increasing rework.

It may deliver significant benefits to one team and almost none to another.

The useful question is not: “How much AI are we using?”

It is: “What happened to engineering performance where AI usage changed?”

‍

The Data You Need Already Exists — Just Not in One Place

Most organizations are not missing the underlying data.

  • ‍AI tools know where AI is being used.‍
  • GitHub or GitLab understand commits, pull requests, reviews, code changes.‍
  • Jira or Azure DevOps understand planned work, initiatives, delivery status.‍
  • CI/CD systems understand deployments.‍
  • Quality/incident systems understand defects, regressions, production issues.‍
  • Organizational systems understand teams reporting structures.‍
  • Cost systems understand what the organization is spending.

‍

The problem is that each system understands only its own part of the world.

A Cursor usage event does not know what initiative the developer was working on.

A GitHub pull request does not automatically know whether it was AI-assisted.

A Jira ticket does not understand what happened during code review.

An AI license does not tell you if the team using it became more productive.

The data is not missing. The context connecting it is.

‍

This Is Where AI Measurement Becomes a Data Engineering Project

Take a seemingly simple leadership question:

“Which teams are generating measurable ROI from AI?”

‍

Answering it reliably may require you to:

  1. Connect AI, planning, delivery, quality, organizational, and cost systems.
  2. Normalize different schemas and definitions.
  3. Resolve developer identities across systems.
  4. Map people to teams, repositories, initiatives, and work.
  5. Identify where AI-assisted activity occurred.
  6. Establish consistent engineering metrics.
  7. Build historical baselines.
  8. Account for reorganizations, workflow changes, and new tools.
  9. Correlate AI usage with downstream outcomes and cost.
  10. Maintain the entire model over time.

Any one of these tasks is manageable.

The complexity comes from keeping all of them correct together.

Teams reorganize. Repositories move. Jira workflows change. AI vendors change. APIs evolve. New leadership questions appear.

The hard part is not building the first dashboard. It is keeping the underlying context trustworthy.

‍

You Can Build It. The Question Is Whether You Should Own It.

Most engineering organizations have the technical capability to build internal analytics.

APIs, warehouses, transformation tools, BI platforms, internal engineering teams, and increasingly capable AI models are all available.

The question is not whether you can build it.

It is what you want to own.

There is a big difference between connecting Jira and GitHub for a dashboard and maintaining a reliable operational model of the engineering organization.

That model needs to understand relationships between:

‍

people → teams → repositories → initiatives → work → code → deployments → AI usage → quality → cost

‍

And those relationships need to remain accurate as the organization changes.

The internal solution therefore comes with ongoing ownership of connectors, schemas, metric governance, organizational mappings, historical consistency, tool migrations, and analytical logic.

The more useful build-vs-buy question is:

Should measuring engineering performance consume engineering capacity of its own?

‍

For some organizations, the answer may still be yes.

But it should be a deliberate decision.

‍

Company Context Is What Turns Engineering Data Into Intelligence

A pull request alone can tell you its size, review time, comments, churn, and merge time. Add company context and you can also understand:

  • which team created it
  • which initiative it supported
  • whether the work was planned
  • whether AI was involved
  • whether it created downstream rework
  • what happened after deployment
  • whether delivery stayed on track
  • what the work cost

‍

That changes the questions leadership can ask.

That is the difference between aggregating engineering data and understanding engineering performance.

‍

Engineering Analytics Should Explain What Changed

Connecting the data still leaves one problem: interpretation.

A dashboard may tell you cycle time increased 18%.

Leadership still needs to determine:

  • which teams drove the change
  • where the workflow slowed
  • whether review or rework increased
  • whether AI adoption changed
  • whether quality moved with it
  • whether the shift requires action

‍

Traditional reporting shows the metric.

Someone still has to explain it.

The next evolution of engineering analytics therefore cannot simply be:

more systems → one dashboard

It needs to be:

‍

connected data → reliable context → continuous interpretation

‍

Engineering leaders need to understand what changed, what is driving it, and where action is required.

‍

The Data and Context Layer You Don't Have to Build Yourself

TargetBoard provides the data, semantic, and intelligence layers engineering organizations otherwise end up building themselves.

‍

TargetBoard connects data across engineering, planning, AI, organizational, quality, cost, and other company systems while allowing teams to continue working in their existing tools.

That data is normalized into a consistent company context that preserves relationships between people, teams, repositories, work, initiatives, delivery, and outcomes.

On top of that context, domain-expert agents continuously interpret performance to surface what changed, what is driving it, and where risk or opportunity is emerging.

That enables engineering leaders to investigate questions such as:

  • Which teams are seeing measurable performance gains from AI?
  • Is increased AI adoption improving delivery predictability?
  • Is faster code creation shifting the bottleneck into review?
  • Is AI-assisted work generating more rework?
  • Which AI tools are associated with better outcomes?
  • Where is AI spend increasing without corresponding improvement?

‍

The goal is not another dashboard. It is removing the data engineering and interpretation work standing between the question and a reliable answer.

‍

Engineering Intelligence Shouldn't Be Priced Per Developer

AI coding assistants are individual-use tools, so per-seat pricing makes sense.

Engineering intelligence is different.

Its value comes from understanding the organization as a system. A developer does not need to log into an analytics platform for their work to contribute to the operational picture leadership needs.

Pricing engineering intelligence per developer means the cost of understanding engineering rises simply because the engineering organization grows.

‍

TargetBoard does not use per-seat pricing.

The objective is organization-wide engineering and AI intelligence, not another product that has to be licensed developer by developer.

‍

The Real Decision Is What You Want to Own

Measuring AI impact is technically solvable. The question is how much infrastructure your engineering organization wants to own in order to solve it.

If you build the capability internally, the commitment extends well beyond connecting a few APIs or creating a dashboard. Someone needs to maintain the data model, keep identities and organizational mappings accurate, absorb changes in source systems, preserve historical consistency, and continually adapt the analysis as new AI tools and new leadership questions emerge.

For organizations with highly specific requirements, that investment may be justified.

But for most engineering leaders, the more useful question is whether building and maintaining this measurement layer creates any strategic advantage.

The value is not in owning the pipelines.

It is in being able to answer, with confidence:

  • Where is AI materially improving engineering performance?
  • Where is it simply increasing activity?
  • What downstream effects are appearing in delivery, quality, and rework?
  • Which investments are producing enough improvement to justify their cost?
  • Where should we change tools, workflows, or investment?

‍

Those are management questions, not data-engineering outcomes.

The goal should be to spend less time assembling the evidence and more time using it to make better engineering decisions.

Best Practice

Engineering Vendor POC Comparison

Selecting an AI engineering tool is difficult when every vendor promises faster delivery, better quality, and higher developer productivity. Demos and feedback provide useful context, but they do not show how a tool will affect your engineering organization. One TargetBoard customer addressed this by running a structured POC comparing Qodo and CodeRabbit. Using real pull request data, the team evaluated each vendor’s impact on delivery flow and made a more defensible decision based on operational outcomes.
August 6, 2026
5 min read

A Successful Demo Does Not Prove Delivery Impact

Most vendor evaluations combine a product demo, a limited developer trial, feature comparisons, and user feedback.

These inputs can show whether a tool is usable, trusted, secure, and compatible with the existing toolchain. They do not establish whether it improves delivery.

A developer may feel faster while using an AI assistant, yet pull requests may still require more review, create more rework, or spend longer waiting to be picked up.

Developer sentiment provides valuable context. Operational data shows what actually changed.

This distinction is particularly important for AI engineering tools. Faster code creation does not automatically lead to faster review, approval, or deployment. A tool may accelerate one stage while moving friction further downstream.

‍

Designing a Fair, Real-World POC

The customer wanted to compare two AI code review automation vendors: Qodo and CodeRabbit.

Rather than testing the tools with unrelated groups or comparing broad company-wide averages, the team used the same defined group of developers throughout the evaluation. Each vendor was tested during a separate period of approximately two weeks.

The methodology was straightforward:

  • Define the developers participating in the POC.
  • Isolate the pull requests associated with those developers.
  • Separate activity by vendor evaluation period.
  • Apply the same engineering metrics to both periods.
  • Compare the results with the organization’s wider historical performance.

This was not a laboratory experiment. Real engineering environments include differences in repository complexity, work type, team availability, and pull request size.

But it was a structured, real-world comparison that produced stronger evidence than a feature checklist or a collection of opinions.

The objective was not to prove that one vendor is universally better. It was to determine which vendor produced better outcomes in this customer’s environment.

‍

Building the Comparison Faster with Saved Filters

The analysis depended on consistently isolating the developers participating in the POC.

Without a reusable filter, the team would have needed to rebuild the participant group for each metric and evaluation period, slowing the process and increasing the risk of inconsistent comparisons.

Using TargetBoard Saved Filters, the team defined the relevant contributors and pull request creators once, then reused the same cohort across the board.

This made it easier to:

  • Compare the Qodo and CodeRabbit periods.
  • Apply the same group across multiple metrics.
  • View POC results against year-to-date performance.
  • Maintain a consistent basis for comparison.

The team could spend less time configuring the analysis and more time interpreting the results.

‍

The Three Metrics That Mattered

The customer focused on what happened after code entered the pull request workflow.

Pull Request Cycle Time

This measured the time from the first commit until the pull request was merged.

It provided an end-to-end view of whether work moved more efficiently during each vendor trial. In this evaluation, the Qodo period showed a shorter average cycle time than the CodeRabbit period.

Cycle time should be treated as a system signal. A higher result may reflect delays in pickup, review, coordination, approval, or integration.

Pull Request Review Cycles

The team also measured how many times a pull request returned to the author for changes before approval.

Fewer review cycles can indicate less back-and-forth and lower review churn. The Qodo period showed fewer average review cycles than the CodeRabbit period.

This was useful evidence, although a full quality assessment would also need to consider defects, incidents, rollbacks, and escaped issues.

Time Spent in Each Pull Request Stage

Overall cycle time shows that a difference exists. It does not explain where the delay occurred.

The customer therefore examined time spent in stages such as coding, waiting for review, active review, and merge.

This helped distinguish active work from waiting time and showed whether differences came from review pickup, review complexity, or later workflow stages.

‍

Turning POC Activity into a Defensible Decision

Across the metrics selected for the proof of concept, the Qodo evaluation period showed stronger results than the CodeRabbit period.

The Qodo period recorded:

  • Shorter pull request cycle time.
  • Fewer average review cycles.
  • More favorable results across relevant workflow stages.

These results gave the customer a concrete basis for the selection decision.

The team was no longer deciding only which product looked more capable in a demonstration or which tool developers preferred. They could compare how each vendor affected real work inside their engineering system.

The result should remain specific to this customer. It does not establish a universal benchmark for either vendor. The outcome reflected the organization’s developers, repositories, processes, work mix, and evaluation periods.

That limitation does not weaken the analysis. It is what makes the result useful.

The customer needed to know which vendor performed better in its own environment.

‍

How to Apply the Same Approach

A useful AI vendor POC should answer two questions: “Which tool did developers prefer?” and “What changed in the delivery system when the tool was introduced?”.

To build a stronger evaluation:

  • Define the expected outcome before the trial begins.
  • Use the same cohort or comparable groups.
  • Apply consistent metric definitions.Measure downstream effects, not only code-generation speed.
  • Compare results with a historical baseline.
  • Consider quality, usability, security, and commercial fit alongside delivery data.
  • Continue measuring after purchase to confirm that the result survives wider adoption.

No single metric should decide the outcome.

A tool may reduce review time while increasing quality risk. Another may receive strong developer feedback but show little measurable effect on delivery. A complete evaluation balances operational outcomes with usability, risk, and cost.

‍

Better Vendor Decisions Start with Better Evidence

AI engineering vendors should be evaluated on more than features, adoption, and perceived time savings.

The real question is whether a tool improves the flow, quality, and predictability of software delivery.

By testing Qodo and CodeRabbit with a defined group of developers, applying consistent operational metrics, and using TargetBoard Saved Filters to accelerate the analysis, this customer turned a typical POC into a more defensible purchasing decision.

The result was not simply another dashboard. It was a clearer understanding of what changed, where the differences appeared, and which vendor produced the stronger outcome for that organization.

‍

Compare Engineering Vendors Using Your Own Delivery Data

TargetBoard helps engineering leaders compare vendor performance using operational data from their own teams and workflows.

See how TargetBoard can help you build a more objective, repeatable vendor evaluation process.

‍

Schedule a Meeting

‍

Best Practice

Best ai code review tools

AI is accelerating developer output, yet that speed is introducing hidden complexity and variability into your delivery systems. You can see the shift in your metrics. Pull requests pile up in review, and work gets stuck without clear root causes. Your leadership team sees cycle time increasing but lacks the operational intelligence to understand why performance is changing. Integrating an AI code reviewer without understanding its systemic impact often just shifts the friction from code generation to code review. The result is delayed detection of architectural drift and execution that becomes less predictable over time. To reclaim delivery predictability, you need to evaluate these tools on how they impact your entire workflow.
July 19, 2026
5 min read

What Is an AI Code Review? (Capabilities vs. Limitations)

An ai code review is the process of using Large Language Models to automatically analyze pull requests. These tools scan source code analysis outputs to detect syntax errors and suggest refactoring options before a human reviewer steps in. They excel at identifying boilerplate code issues and enforcing standard automated linters. But they struggle with cross-service dependencies and complex business logic constraints.

Capability Area What AI Excels At What AI Struggles With
Syntax and Formatting Enforcing automated linters and catching basic typos instantly. Understanding nuanced stylistic choices specific to your team.
Security Scanning Identifying common OWASP vulnerabilities and exposed secrets. Detecting complex threat vectors hidden across multiple microservices.
Code Generation Writing boilerplate code and standard unit tests. Grasping repository context and long-term architectural impact.

The "Augment, Don't Replace" Philosophy

The most effective engineering teams treat an ai code reviewer as a high-speed assistant rather than an autonomous decision-maker. AI models lack the operational context to make final architectural decisions. They can't negotiate API contracts or understand why a specific workaround exists for a legacy system.

That means a human-in-the-loop review remains absolutely critical. You use the AI to clear out the noise of code formatting and basic threat detection, so your senior engineers can focus their cognitive energy on system design and business logic.

Solving the Systemic Context vs. File-Level Analysis Gap

A major limitation of current AI tools is their reliance on file-level analysis. An AI assistant might review a single pull request and confirm the syntax is perfect. Yet that same code might break cross-service dependencies three layers deep in your application.

This happens because AI context windows face strict VRAM limits and memory constraints, preventing them from holding your entire codebase in memory at once. Trusting AI file-level analysis without verifying the broader repository context is a common mistake that leads directly to architectural drift. Your delivery pipeline must connect code changes to system-wide impacts to prevent this risk.

Are AI Code Reviews Accurate?

Yes, a code review ai is highly accurate when evaluating isolated syntax and standard formatting rules. Conversely, accuracy drops to near zero when evaluating complex logic or proprietary frameworks. This drop in precision introduces high rates of false positives and AI hallucinations into your pull requests.

Consider a common scenario where an AI tool successfully identifies a missing variable declaration but completely misses a breaking change in your core payment processing logic. The AI then floods the pull request with dozens of comments about stylistic formatting. Developers end up arguing with an AI bot in the comments over subjective syntax choices, creating massive review churn.

This noise creates an overwhelming backlog for human reviewers and actively slows down sprint velocity. Developer overreliance on these tools compounds the problem. Junior engineers might blindly accept AI suggestions without understanding the underlying code, injecting hidden technical debt into the system. You must measure this friction continuously to ensure the tool is actually accelerating your workflow rather than just generating noise.

Top AI-Powered Code Review Tools Compared for 2026

Selecting the best ai code review tools requires matching the platform's core capability to your specific workflow bottleneck. You must differentiate between tools that generate code, platforms that scan for vulnerabilities, and systems that measure the systemic impact of those changes.

Tool Category Example Platforms Primary Function Workflow Impact
PR Summarization CodeRabbit, Qodo Analyzes pull requests to generate human-readable summaries and catch basic syntax errors. Reduces initial cognitive load for human reviewers but can generate noise.
IDE Extensions GitHub Copilot Lives directly in the developer environment to suggest code blocks and refactoring options in real time. Accelerates raw code generation but shifts the bottleneck to the review phase.
Security Scanners SonarQube, Greptile Performs deep static analysis across the pipeline to identify vulnerabilities and enforce compliance. Catches known threat vectors early but often struggles with complex business logic.
Operational Intelligence TargetBoard Connects data across systems to explain how AI impacts delivery predictability and workflow efficiency. Exposes hidden bottlenecks and review churn caused by AI-generated code.

Automated Pull Request Summarization Bots

Tools like CodeRabbit and Qodo focus heavily on pull request summarization. They read the diff and generate a plain-language summary of the changes, so human reviewers can grasp the intent faster. This approach often improves initial time-to-merge metrics for simple tasks.

But open source ai code review tools in this category can struggle when deployed on massive enterprise monorepos. The sheer volume of interconnected files overwhelms the model. This leads to generic summaries that fail to capture the actual architectural impact of the change.

Native Integrated Development Environment Extensions and Agents

GitHub Copilot and similar IDE extensions operate directly where developers write code. These tools use agentic workflows to suggest entire functions as the developer types. They are incredibly effective at reducing the time spent writing boilerplate syntax.

They operate with a limited view of the broader system. A native extension might suggest a highly efficient sorting algorithm, yet it can't verify if that logic violates broader API contracts established by another team. Human reviewers must still validate those systemic connections.

Enterprise Continuous Integration and Continuous Deployment Security Scanners

Enterprise platforms like SonarQube and Greptile focus on strict CI/CD integration. They run deep static analysis to ensure your codebase maintains OWASP compliance and prevents known vulnerabilities from reaching production. These tools are non-negotiable for teams operating in highly regulated environments.

A major consideration in this category is data sovereignty. Sending proprietary enterprise code to external models for security scanning introduces compliance risks. You must configure these tools to ensure sensitive data remains within your controlled infrastructure.

Operational Intelligence and Measurement Platforms

Adopting ai powered code review tools frequently increases raw output while secretly damaging delivery predictability. You need a way to measure this friction. TargetBoard is an agentic operational intelligence platform that helps leadership teams understand how execution is performing, why it is changing, and how to respond.

TargetBoard connects data across company systems and uses domain-expert AI agents to understand workflow bottlenecks. It acts as the essential operational intelligence layer that shows you if your AI coding tools are actually improving sprint velocity or just creating massive review churn.

How to Do an AI Code Review (Without Increasing Review Churn)

Implementing ai code reviews requires strict boundaries. You must configure the tool to handle objective rules while reserving subjective architectural decisions for human engineers. If you fail to set these boundaries, the AI will argue with your developers over code formatting and stylistic preferences.

This friction causes massive review churn and slows down your entire pipeline. You must structure the workflow to prevent this noise.

Visualizing the AI-Augmented Continuous Integration and Continuous Deployment Workflow

You must map exactly where the AI intervenes in your Software Development Lifecycle. The AI should run its analysis immediately upon pull request creation. It scans for syntax errors, basic code smells, and formatting violations.

The developer resolves these objective flags before a human reviewer is ever assigned to the pull requests. This sequence ensures your senior engineers only spend their time reviewing complex logic and system architecture.

Configuring Custom Rule Files and Guidelines

You must train your AI tools using custom rule files specific to your repository. This step prevents the AI from suggesting changes that violate your internal business logic constraints. You can configure the tool to enforce DRY principles and flag code duplication automatically.

The interaction between these custom rule files, the model's context windows, and your code repositories determines the success of the tool. A well-configured rule file reduces false positives and ensures the AI only surfaces actionable insights.

How to Measure the Systemic Impact of AI Code Reviews

You can't manage what you don't accurately measure. Relying on basic productivity metrics like lines of code written will mislead your leadership team. According to the 2023 DORA Report, true delivery predictability matters far more to business outcomes than raw development speed. You must measure if your ai code review tools are actually accelerating delivery or just shifting the bottleneck.

TargetBoard provides this critical measurement layer. It tracks the difference between AI-generated output and human review times. If an AI tool increases output by 40 percent but causes pull requests to sit in review for three extra days, your actual sprint velocity decreases. TargetBoard exposes these hidden workflow bottlenecks, so you can adjust your strategy based on objective operational intelligence rather than intuition.

Reclaiming Engineering Velocity Safely

The primary value of an AI code review tool is workflow efficiency, not replacing human architectural judgment. These tools are highly effective at clearing out boilerplate errors and enforcing basic code quality. Yet they introduce their own hidden complexities that require continuous systemic measurement.

According to 2024 GitHub Copilot research, AI assistants boost developer productivity by up to 55 percent. You must balance that speed with strict oversight to protect codebase maintainability and prevent the accumulation of technical debt. By running an operational intelligence layer alongside your AI tools, you can safely accelerate software delivery while maintaining complete confidence in your engineering metrics.

Best Practice

What is value stream management

You look at your performance dashboard and see cycle time creeping up, but the data can't tell you why. You ask your engineering managers for an update, only to receive conflicting subjective reports pulled from different tools. The gap for modern engineering leaders is no longer visibility. The real problem is understanding why execution breaks down and how to coordinate decisions across fragmented systems. This guide explains how value stream management helps you identify workflow bottlenecks and why transitioning from static dashboards to operational intelligence is the only way to regain delivery predictability.
July 19, 2026
5 min read

What Is Value Stream Management?

Value stream management (VSM) is an operational framework that connects business objectives to the software delivery lifecycle. The goal is to optimize how work moves from idea to production, helping leaders identify constraints and improve continuous flow. But tracking work is only the first step.

You must connect those tracking metrics to actual customer value and time-to-market outcomes. When you understand how value flows through your organization, you can stop reacting to delayed releases and start proactively removing the barriers that slow your teams down.

The 4 Core Principles of Value Stream Management

To build a reliable value delivery pipeline, you need to understand the foundational rules of the methodology. These principles guide teams toward predictable delivery and continuous improvement.

Principle Definition in Software Delivery
Outcome-Focused Prioritizes measuring the actual business value delivered to the customer rather than just tracking feature output or story points.
Systems Thinking Views the delivery process as a single connected system to prevent local optimizations that hurt overall performance.
Continuous Flow Focuses on removing delays and handoffs so work moves smoothly from planning through deployment.
Waste Reduction Identifies and eliminates non-value-adding activities like rework, excessive meetings, and unmanaged technical debt.

How Value Stream Management Operates in Software Delivery

Applying these concepts to engineering requires a hard look at how your teams actually work. You likely track engineering performance using standard indicators like cycle time, lead time, and deployment frequency. These numbers provide a baseline for your delivery speed.

But a dashboard showing a spike in lead time doesn't solve the underlying problem. You have to trace that metric back to the specific workflow behaviors causing the delay. This requires connecting data across your planning and code systems to see the reality of your operations.

Identifying Workflow Friction and Bottlenecks

Workflow friction often hides inside routine development tasks. Consider a scenario where your overall cycle time suddenly spikes by 40 percent. The dashboard flags the delay, but it can't tell you that three high-complexity pull requests have been sitting in the review queue for four days.

The code is written, yet cross-team dependencies and unclear ownership prevent anyone from merging it. This code review churn artificially inflates your cycle time metrics. The work itself isn't slow, but the system is blocked. Identifying these specific constraints allows you to clear the path rather than just asking teams to code faster.

Shifting From Project to Product

Traditional organizations fund temporary projects, which naturally creates organizational silos. Teams assemble, build a feature, and then disband. This breaks execution alignment and leaves no clear owner for long-term maintenance or technical debt.

Modern value stream management requires a shift toward a product-centric model. You fund stable, cross-functional teams that own a specific product from end to end. This structure improves capacity allocation because you align your best engineers with long-term value delivery rather than temporary task lists. The result is a more resilient delivery engine that adapts quickly to market changes.

How to Do Value Stream Management Step by Step

Implementing this framework requires a structured approach to analyzing your value streams. You need to connect resource planning directly to your value delivery pipeline. This ensures you are solving the right problems instead of just optimizing isolated tasks.

What Are the 4 Steps of Value Stream Mapping?

Value stream mapping is the diagnostic tool you use to visualize how work flows through your organization. Follow these four steps to build an accurate map:

  1. Identify the product: Define the exact service or feature set you are analyzing.
  2. Map current state: Document every step and handoff currently required to deliver that product to the customer.
  3. Design future state: Identify bottlenecks in your current map and design an optimized workflow without those constraints.
  4. Create implementation plan: Build a clear roadmap to transition your engineering teams to the optimized future state.

The 7 Wastes of Value Stream Mapping in Modern Engineering

To improve flow efficiency, you must identify where engineering effort goes to waste. Modern software leaders face specific capacity concerns that look very different from physical manufacturing. Here is how the classic seven wastes translate to software delivery.

Lean Waste Software Engineering Reality
Overproduction Building extra features that customers don't actually use or need.
Waiting Cross-team dependencies delaying a critical release.
Conveyance Friction during handoffs between development, quality assurance, and operations teams.
Extra Processing High code review churn and excessive manual approval gates.
Inventory Partially done work sitting in abandoned branches or unmerged pull requests.
Motion Task switching and context shifting across fragmented development tools.
Defects Unmanaged technical debt and critical bugs escaping to production.

The Limitations of Traditional Value Stream Management (Why Dashboards Aren't Enough)

You can map your workflows perfectly, but legacy tools often fail because they rely on metrics without context. You see cycle time shifting, but you can't explain why execution breaks down. According to a 2023 Gartner report on engineering operations, most leaders struggle because their operational data is trapped in data silos.

This forces executives to rely on subjective updates from managers instead of trusted system-level reality. Tracking metrics provides visibility, but it doesn't provide understanding. TargetBoard is an agentic operational intelligence platform that connects data across company systems, interprets performance through operational intelligence, and uses domain-expert AI agents to guide execution decisions.

This shifts your organization from reactively monitoring dashboards to proactively fixing workflow friction. You gain the power to make confident execution decisions based on reality.

How Artificial Intelligence Changes the Value Stream

Artificial intelligence code generation accelerates output, so it fundamentally alters how work flows through your system. But higher output often introduces hidden delivery risk. For example, artificial intelligence code frequently experiences higher code review churn than human-written code because it requires intense scrutiny to verify complex logic.

If you only measure output volume, you miss the bottleneck forming in your review stage. This hidden complexity slows down the entire pipeline and delays critical execution decisions.

Measuring vs. Understanding Performance

Tracking DevOps Research and Assessment metrics is a good start, but it's only showing you the symptoms of an inefficient system. You need to diagnose the disease through root cause analysis to achieve predictable delivery.

Approach Capability Limitation or Advantage
Traditional Dashboards Tracks baseline metrics like cycle time and deployment frequency. Fails to explain why metrics change or where workflow friction originates.
Agile Planning Tools Tracks sprint velocity and ticket completion rates. Relies on manual updates and often contradicts the reality of the code repository.
TargetBoard Connects code, planning, and delivery systems using agentic operational intelligence. Explains why performance changes and translates signals into proactive execution decisions.

Moving Beyond Metrics to System-Level Delivery Predictability

A common mistake in engineering leadership is treating performance metrics as goals rather than lagging indicators of system health. According to the 2023 Forrester Report on software delivery, teams that focus purely on metric targets often sacrifice long-term stability. When you stop chasing numbers and start focusing on resolving the underlying workflow constraints, your delivery confidence naturally improves.

This operational shift connects daily engineering tasks directly to broader business outcomes. By treating visibility as a starting point rather than the finish line, you create a culture of continuous improvement that actually scales. Understanding your system gives you a clear framework for your next planning session or your next board meeting.

‍

Best Practice

New Target Types

Organizations often struggle to manage different types of performance goals effectively, leading to misalignment and missed targets. The key idea is that tailored target-setting—such as milestones, improvements, and limits—enables more precise tracking and better outcomes. TargetBoard addresses this by providing specialized tools and real-time insights to help teams set, monitor, and achieve goals more effectively.
April 8, 2026
5 min read

At TargetBoard, we continually strive to innovate and tailor our solutions to meet the dynamic needs of modern organizations. Recognizing that achieving strategic goals requires versatile and precise tools, we're excited to announce new target types in our latest platform update. Each target type is designed to address specific challenges and metrics, ensuring that leaders can set and reach their objectives more effectively. Let’s dive into how these new enhancements can transform the way your organization achieves its goals.

Milestone Targets

Milestone targets are invaluable for metrics that need to start from zero and achieve a specific value by a predetermined date. Whether it’s completing key projects, implementing new programs, or hitting quarterly sales targets, this type of goal setting provides a clear timeline and a definitive endpoint, making it easier to organize resources and efforts. TargetBoard’s tools help you track these milestones, offering insights and reminders to keep your team aligned and focused.

Improvement Targets

For metrics that have an established baseline, improvement targets are ideal. These targets aim to enhance performance by a certain percentage or degree, perfect for increasing efficiency metrics like cycle time at both group and individual levels. With TargetBoard, you can monitor ongoing changes against these baselines, adjust strategies in real-time, and drive continuous improvement across your organization.

SLA or Upper Limit Targets

Certain metrics need to be kept below a threshold to ensure quality and efficiency—this is where SLA or upper limit targets come into play. These are critical for operations like support ticket resolution, production incident management, or recruitment processes. By setting an upper limit, you ensure that these activities do not exceed acceptable time frames, thereby optimizing performance and customer satisfaction. TargetBoard’s alerts and performance tracking make it easy to stay within these limits.

Lower Limit Targets

Conversely, lower limit targets ensure that crucial metrics do not fall below a certain level. This target type is particularly useful for maintaining standards in areas such as planning accuracy or system uptime. Ensuring that these metrics stay above a specified point helps in maintaining operational continuity and reliability. With TargetBoard, safeguarding these standards becomes straightforward, thanks to our real-time monitoring and notification systems.

A Partner in Achieving Success

At TargetBoard, we go beyond just helping you set targets. We’re committed to doing everything in our power to assist you in reaching them. Our platform is equipped with powerful tools like detailed insights, timely notifications, and regular reminders.

These features are designed to keep your team on track, ensuring that each target receives the attention it deserves and boosting your chances of success.

In conclusion, TargetBoard is more than just a tool for setting targets—it’s a comprehensive solution that supports your strategic goals at every level of the organization.

By understanding the unique nature of different targets and providing specialized tools to meet these needs, TargetBoard empowers leaders to achieve more and reach their objectives with precision and ease.

Best Practice

The Value of Processes in Crisis

Crises often disrupt structured processes, forcing organizations into reactive, short-term decision-making that can create stress and misalignment over time. The key idea is that reintroducing structured frameworks is essential for restoring stability, clarity, and productivity after disruption. TargetBoard supports this recovery by providing visibility and guidance to help teams regain alignment and operational rhythm.
April 14, 2026
5 min read

In the dynamic landscape of modern business, crises are inevitable. From internal upheavals to external shocks like wars or economic downturns, organizations are constantly tested in their resilience and adaptability. During these challenging times, the role of established processes becomes crucial in steering teams back to stability and productivity.

The Backbone of Normalcy: Structured Frameworks

In everyday operations, structured frameworks and processes – be it Agile sprints or regular meetings – serve as the backbone of organizational functionality. They provide a rhythm to our work, a predictable pattern that helps align teams internally and sync activities with external stakeholders. These processes are more than mere routines; they act as bulwarks against abrupt shifts in priorities or strategies, fostering a more deliberate and planned approach to work.

Crisis and the Shift in Dynamics

However, in times of crisis, such as during critical all-hands events or geopolitical disturbances, these frameworks often take a backseat. The immediate response to crisis typically involves loosening structured processes to allow for quicker decision-making and action. This shift is understandable: fewer people might be available, and there’s a need for shorter reaction cycles to address pressing issues. While this approach yields immediate effectiveness, its long-term impact can be counterproductive, adding stress and anxiety to already tense situations.

The Double-Edged Sword of Flexibility

Moving to daily Kanban systems or adopting a hands-on management style may seem beneficial in the short term, but their impact on long-term planning and execution can be detrimental. This flexibility, while necessary in extreme situations like wars or civil unrest, can later hinder the realignment of employees with organizational goals. The challenge then becomes not just coping with the crisis but also recovering from the disruption it caused to established work patterns.

The Power of Returning to Structured Processes

Our experience at TargetBoard shows that reintroducing structured processes, such as transitioning from Kanban back to Agile (Sprints), plays a pivotal role in post-crisis recovery. This shift is not just about regaining control; it's about reestablishing a shared understanding of expectations between teams and individuals. It enables companies to gauge their capacity realistically and aids employees in refocusing their efforts on achievable targets. Most importantly, it alleviates the uncertainty and anxiety that come with turbulent times, channeling employees' concerns into productive endeavors.

How TargetBoard Facilitates Recovery

TargetBoard emerges as a vital tool in this recovery process. Our platform is designed to help teams regain their operational rhythm. We offer insights into where intervention might be necessary and assist in monitoring the gradual return of employees to a productive cadence. By leveraging our tools, companies can not only navigate through the crisis but also emerge stronger, with a renewed sense of purpose and direction.

Conclusion: Embracing Structure in Times of Uncertainty

In conclusion, while the immediate response to crises may necessitate a departure from established processes, the path to recovery and resilience lies in embracing these structures once more. By providing a framework for action and decision-making, structured processes help organizations navigate through uncertain times, ultimately paving the way for a return to stability and growth.

Best Practice

Acquisition Ensuring Smooth Transitions

Mergers and acquisitions introduce major operational, cultural, and strategic disruptions that can impact productivity and long-term success if not managed carefully. The key idea is that tracking and understanding these changes in real time is critical to ensuring smooth integration and maintaining performance. TargetBoard supports this by providing continuous KPI visibility and insights, helping organizations monitor progress, detect issues early, and guide successful transitions.
April 15, 2026
5 min read

In the ever-evolving landscape of the tech industry, mergers and acquisitions (M&A) are par for the course. These pivotal moments can herald exciting times of growth, innovation, and expansion. However, they also bring about significant upheaval. Whether you're on the side of the acquirer or the acquired, the changes that follow an M&A deal are far-reaching. From shifts in management and corporate priorities to overhauls of processes and operational methodologies, the impact is profound. These transformations, while aimed at fostering a stronger entity, can lead to distractions and disruptions, affecting the workforce's morale and productivity.

Examples of Changes in Tech M&A

- Management Restructuring:

‍One of the most immediate and visible changes is in leadership. New executives may be brought in, or leaders from the acquiring company may take over, leading to shifts in corporate culture and strategy.

‍- Integration of Processes: Combining two distinct sets of operational processes can be challenging, as it often requires streamlining workflows, technologies, and systems to achieve synergy.

‍- Cultural Reconciliation: Perhaps one of the trickiest aspects to navigate, blending two distinct corporate cultures can make or break the post-M&A integration phase.

‍- Prioritization of Projects: Post-M&A, some projects might be accelerated, while others could be put on the backburner or scrapped altogether, affecting team morale and individual job securities.These changes, albeit necessary, are a double-edged sword. If not carefully planned, managed, and communicated, they can lead to significant disruptions, affecting the overall health of the combined entity.

Potential Risks of Early BI and Analytics Investment:

1. Resource Allocation: For startups, every penny counts. There’s always the looming question: Is it better to invest in analytics or channel those resources into direct product development or marketing?

‍2. Budgetary Limitations:Operating on a tight budget can lead to makeshift data solutions that might be riddled with inaccuracies, defeating the purpose of BI.

‍3. Flexibility Concerns: With a strong commitment to specific KPIs, there's a risk of tunnel vision, possibly sidelining other emergent opportunities.

The Thin Line Between Success and Failure

The success of a tech M&A largely hinges on how well these transitions are managed. Let's look at a few of examples:

‍- Google's Successful Acquisition of Android: This is often cited as one of the most successful tech acquisitions. Google allowed Android to operate semi-autonomously, preserving its innovative culture while providing the resources needed for explosive growth.

‍- AOL's Failed Acquisition of Time Warner: One of the most infamous examples of a failed M&A, the merger struggled due to a clash of corporate cultures, among other issues, leading to a massive loss in value.These examples underscore the sensitivity of the post-M&A period, which can indeed set the tone for the future success or failure of the combined entity.

The Challenge of Tracking Post-M&A Changes

Tracking the myriad changes post-M&A and understanding their impact on the team, including their velocity, quality, capacity, and engagement, is exceedingly complex. Traditional frameworks often fall short, and the capacity to develop new ones swiftly is usually lacking. This is where TargetBoard steps in.

How TargetBoard Supports Smooth Transitions

TargetBoard is designed to effortlessly connect with both entities involved in the M&A from day one. It starts tracking all key performance indicators (KPIs), offering a clear, accurate insight into how teams are adapting to their new realities. This data-driven approach ensures that the combined entity is set up for long-term success, providing:‍

- Real-time Monitoring: Continuous tracking of changes and their impacts, offering a comprehensive overview of the integration process.

‍- Early Warning System: Quick identification of potential issues, allowing for prompt intervention before they escalate.

‍- Engagement and Morale Insights: Understanding how changes affect team morale and engagement, crucial for maintaining productivity and innovation.

In conclusion, TargetBoard acts as a navigational aid in the often turbulent waters of tech M&As. By offering a detailed, real-time view of the integration's progress and impact, it helps

Business

Lead time for changes

You pull up your engineering dashboard and see lead time for changes expanding across your core teams. Your developers are using AI coding tools and pushing more commits than ever before. Yet your deployment frequency is flat and your delivery predictability is collapsing. The data shows the slowdown but it completely fails to explain why it's happening. A high lead time signals deep workflow friction across your teams. It tells you that code is sitting in review queues or stalled in handoff friction. Software delivery is constrained by human coordination capacity rather than individual coding speed. Understanding this shift is the only way to stop reacting to lagging metrics and start restoring delivery predictability.
July 19, 2026
5 min read

What Is Lead Time for Changes?

Lead time for changes is the amount of time it takes a commit to get into production. As a core DORA metric used to evaluate software delivery performance and CI/CD pipeline efficiency, it tracks three specific phases:

  • Time of commit: The moment a developer pushes code to the repository.
  • Testing and review: The period where code undergoes automated testing and human verification.
  • Time of deployment: The exact moment the code successfully runs in production.

That's the mathematical calculation, but the operational reality is very different. When you track change lead time, you aren't actually measuring developer productivity or how fast your team types. You're measuring how long your code sits still. A high lead time signals deep workflow friction across your teams, exposing the hidden waiting states in your delivery process.

Code spends the vast majority of its lifecycle waiting for a human to review it or coordinate its release. If your lead time to deploy is increasing, your organization is suffering from review congestion and organizational friction. The bottleneck is your queueing system, so you have to look at your waiting states to find the real problem.

Lead Time vs. Cycle Time vs. Lead Time for Changes

Engineering leaders often see conflicting signals across disconnected systems because teams use these terms interchangeably. You need precise boundaries to diagnose where work actually gets stuck.

Metric Starting Point Ending Point What It Actually Measures
Lead Time Customer request accepted Feature delivered to user Total business responsiveness and product planning efficiency.
Cycle Time Developer begins work Code is merged or completed Engineering efficiency and sprint velocity.
DORA Lead Time for Changes Code is committed Code runs in production Software delivery performance and CI/CD pipeline friction.

If your cycle time is stable but your DORA metric is expanding, your developers are not the problem. Your deployment coordination and review queues are failing to keep up with the output.

What Is a Good Lead Time for Changes?

The DORA research program provides clear benchmarks for this metric, categorizing software delivery performance into four tiers:

  • Elite performers: Less than one hour.
  • High performers: Between one day and one week.
  • Medium performers: Between one week and one month.
  • Low performers: More than one month.

These benchmarks are useful for setting a baseline, but they don't tell you how to improve. Traditional engineering organizations try to hit the elite tier by pushing developers to work faster. This approach completely ignores how flow systems actually behave.

Elite teams don't type faster. They systematically reduce waiting queues through batch size reduction and automated testing. They understand that a large queue size directly predicts a delivery predictability collapse.

When you push more code into a congested system, you don't get faster delivery. You get massive PR saturation. The code sits in review queues for days, forcing developers to context-switch to new tasks. This destroys momentum and inflates your metrics. Achieving elite performance requires you to manage your bottlenecks and reduce handoff friction, so you must focus on your coordination capacity rather than your raw coding speed.

The Hidden Causes of High Lead Time

When looking at a slow pipeline, executives assume developers are struggling with technical problems. The reality is that code spends most of its life in waiting systems. Systemic bottlenecks rarely happen during active development. They happen when work stops moving and enters workflow queues. A high lead time is a symptom of organizational friction, so you have to look at the spaces between your developers to find the real delays.

Handoff Friction and Coordination Latency

Every time a piece of code changes hands, it loses momentum. This handoff friction is a primary driver of expanding lead times. A developer finishes a feature and requests a review, but the reviewer is busy with their own sprint commitments. The code sits idle, which introduces coordination latency into the system.

These process bottlenecks compound across teams. Queueing-system dynamics dictate that as utilization approaches 100 percent, wait times increase exponentially. Pushing your teams to maximum capacity actually creates deployment coordination failures and slows everything down.

The AI Era Trap: Coding Speed vs. Review Saturation

Artificial intelligence coding assistants have fundamentally changed how work is produced. Developers can now generate massive amounts of code in minutes. But this surge in artificial intelligence code generation creates a dangerous trap. The raw output increases, yet the human capacity to verify that code remains exactly the same.

This imbalance leads directly to review saturation. Complex pull requests (PRs) stack up, extending the code review cycle by days. Reviewers struggle to untangle the sudden influx of code complexity, so they delay the review or rubber-stamp it.

This review congestion forces developers to context-switch while waiting for approvals. That destroys efficiency, inflates your lead time, and ultimately worsens your time to restore service / MTTR when bad code slips through.

Why Metrics Without Context Fail: Shifting to Operational Intelligence

Tracking DORA metrics gives you a trailing signal of your performance. But metrics alone cannot explain why performance is changing. When executives rely on manual reporting, they waste hours aggregating fragmented data from Jira and GitHub. The numbers often conflict, which erodes trust in the data and creates massive engineering overhead. A high lead time signals deep workflow friction across your teams.

To actually solve these delays, you have to understand the bottleneck economics of your delivery pipeline. You need a single source of truth that connects disconnected systems and exposes hidden waiting states. This is why engineering leaders are moving away from passive manual reporting and shifting toward operational intelligence, enabling true data-driven decision making.

Approach What It Provides How It Handles Bottlenecks Decision Impact
Traditional DORA Dashboards Passive metrics and trailing charts. Shows a delay occurred after the fact without context. Reactive tracking based on fragmented data.
TargetBoard Agentic operational intelligence platform unifying performance data into a trusted model. Connects cross-system signals to expose why workflow queues are forming. Proactive identification of hidden risks and workflow friction.

Software Delivery as a Measure of Coordination Capacity

You can't solve a delivery slowdown by forcing your team to write code faster. Software delivery is fundamentally constrained by human coordination capacity. According to the foundational Accelerate research, organizations must optimize the entire value stream. As AI accelerates code generation, it pushes organizations up against strict human verification limits. If your review queues can't handle the volume, your lead time to deploy will inevitably increase.

Ignoring queueing-system dynamics creates severe maintainability risks. Code that sits in congested pipelines degrades in quality, and developers lose the context needed to fix it, leading to a massive accumulation of technical debt. Restoring predictability requires you to manage your workflow queues actively. You must align your output with your capacity to review, verify, and coordinate that work.

Business

Software Development Performance Metrics

You sit down to prepare for the board meeting, pulling Jira ticket velocity on one monitor and GitHub merge times on the other. The numbers completely contradict each other. Jira shows a record-breaking sprint, yet your GitHub data reveals pull requests sitting in review for four days. You see the metrics shift, but you can't confidently explain why delivery is actually slowing down. That lack of understanding forces you to rely on guesswork, which destroys delivery predictability and erodes trust with the C-suite. Traditional software development performance metrics treat delivery like a disconnected scoreboard. Improving individual metrics on a dashboard does not guarantee overall performance improvement. Performance is actually an interconnected system. Managing fragmented tools prevents leaders from understanding where execution is breaking down. This gap widens as Artificial Intelligence coding tools accelerate raw output while hiding underlying complexity. Organizations have strong systems for measuring performance, so they must now build systems for interpreting it. You don't just need to measure engineering performance. You need to explain why it's changing.
May 10, 2026
5 min read

What Are Software Performance Metrics? The Four Core DevOps Research and Assessment Metrics

Software development performance metrics are operational signals that measure how efficiently a team delivers code to production. The industry standard baseline relies on the four core DevOps Research and Assessment metrics. These engineering Key Performance Indicators divide performance into speed and stability.

VPs of Engineering often fall into a scoreboard mentality when tracking these numbers. They spend hours manually aggregating point-in-time reports, treating the metrics as the final goal rather than a diagnostic signal. Improving these software delivery performance metrics requires understanding the workflow friction beneath the numbers. Frameworks provide signals, so they don't provide full understanding on their own. You must connect these signals to actual execution decisions to improve delivery predictability.

#1. Cycle Time

Problem: Teams ship features slowly and can't pinpoint where work gets stuck in the pipeline.

Solution: Measure cycle time to identify bottlenecks in the review and deployment phases.

  • Cycle time measures the total time elapsed from the moment a developer commits code to the moment that code reaches production.
  • Elite benchmark: Top-performing teams maintain a cycle time of less than 26 hours.
  • Core driver: A high cycle time usually indicates massive pull requests or heavy cross-team dependencies.
  • Execution focus: Teams must balance throughput vs. instability by breaking work down into smaller increments.

#2. Deployment Frequency

  • Deployment frequency tracks how often an engineering team successfully releases code to production.
  • Elite benchmark: Elite performing teams deploy multiple times per day.
  • Frequent deployments require highly automated testing pipelines, making this one of the most critical software developer metrics.
  • Execution focus: High deployment frequency reduces the risk of massive release failures and forces teams to work in small batches.

#3. Change Failure Rate

  • Change failure rate measures the percentage of deployments that cause a failure in production requiring immediate remediation.
  • Elite benchmark: The elite benchmark for change failure rate sits between 0% and 15%.
  • This metric acts as a critical counterweight to deployment frequency.
  • Execution focus: A rising change failure rate signals unmitigated delivery risk, meaning the team is sacrificing quality for speed.

#4.  Mean Time To Recovery

  • Mean time to recovery tracks how long it takes an organization to restore service after a production failure occurs.
  • Elite benchmark: Elite teams achieve a mean time to recovery of less than one hour.
  • Failures are inevitable in complex systems, making this a vital software delivery performance metric.
  • Execution focus: Fast recovery times indicate strong observability practices and resilient system architecture.

The Artificial Intelligence Systemic Breakdown: How Increased Output Masks Hidden Complexity

Artificial intelligence code generation fundamentally changes how software is built. Tools like Copilot and Cursor allow developers to write thousands of lines of code in minutes. And this massive increase in raw throughput completely breaks traditional software developer productivity metrics.

You look at your dashboards and see record-high commit volumes. The metrics suggest the team is moving faster than ever, yet overall delivery predictability drops. This happens because increased output actively masks hidden complexity. AI tools generate code quickly, but that code often lacks systemic context. The resulting codebase becomes brittle, and the organization accumulates technical debt faster than human developers can refactor it.

Pull Request Bottlenecks: When High Volume Meets Human Limits

  • The volume problem: Artificial Intelligence generates massive blocks of code, so pull request size and review time explode.
  • The human limit: Human reviewers simply can't process this high volume of generated code at the same speed it's created.
  • Workflow friction: Work piles up in the review stage, and developers spend days waiting for approvals.
  • Code review churn: Reviewers face extreme cognitive overload, so subjective review decisions become inconsistent. They either rubber-stamp complex pull requests without proper scrutiny or block them indefinitely out of caution.

Tracking Defect Density and Long-Term Technical Debt

  • The quality gap: Fast code generation often results in poor long-term maintainability.
  • Defect density tracks the number of confirmed bugs relative to the size of the software module.
  • The AI flaw: AI-generated code frequently contains subtle logical flaws that bypass automated tests, so defect density rises steadily over time.
  • Engineering investment: Teams spend less time building new features and more time keeping the lights on. Maintainability trends downward as the codebase becomes more complex.

Qualitative Metrics: Developer Experience and Flow

Quantitative data only tells half the story, so engineering leaders must also track qualitative metrics to understand the reality on the ground. Frameworks like the SPACE framework provide a more balanced view by combining qualitative and quantitative data. This approach prevents leaders from optimizing a system to the point of breaking the people running it.

You can't measure system health without measuring Developer Experience. High workflow friction directly degrades how developers feel about their work. When developers constantly fight broken pipelines or wait days for code reviews, their satisfaction plummets and delivery slows down.

  • Satisfaction and well-being: Track how developers feel about their tools and processes through regular surveys to prevent burnout.
  • Measure the actual performance outcomes of the software delivered rather than just the volume of output, since raw volume rarely correlates with business value.
  • Monitor activity in the design and coding phases to understand where developers actually spend their time.
  • Communication and collaboration: Evaluate how easily teams share knowledge and review each other's work across the organization, because siloed information directly inflates cycle time.
  • Efficiency and flow: Track the ability of developers to stay in a state of deep work without facing constant pipeline interruptions, which ultimately dictates their true productivity.

Implementing Work In Progress Limits and Team Goal Alignment

Problem: Teams take on too many tasks at once, so context switching destroys their focus and stalls delivery.

Solution: Implement work in progress limits to force completion before starting new tasks and increase delivery confidence.

  1. Identify the bottleneck: Map your current workflow to find exactly where tickets pile up. This usually happens in the code review or QA testing phases.
  2. Set strict constraints: Cap the number of active tickets allowed in that specific workflow state so developers are forced to finish existing tasks before starting new ones. If the limit is three, developers can't move a fourth ticket into that column.
  3. Force team swarming: Require developers to help unblock stuck tickets before they pull new work from the backlog. This aligns team behavior with overall delivery goals rather than individual task completion.
  4. Adjust continuously: Review these limits during retrospectives and tackle the underlying workflow friction causing the pileup, which prevents the same bottlenecks from recurring next sprint.

Three Outdated Anti-Patterns to Avoid When Measuring Engineering KPIs

Enterprise engineering teams still rely on outdated measurement tactics that incentivize the wrong behaviors. Measuring the wrong things creates a toxic culture and actively hides systemic risks.

Anti-Pattern The Problem The TargetBoard Solution
Tracking output volume Developers optimize for lines of code rather than solving the actual business problem. TargetBoard measures system efficiency and workflow bottlenecks instead of raw code volume.
Pitting developers against each other Tracking individual performance destroys collaboration and incentivizes developers to hoard easy tasks. TargetBoard analyzes cross-team dependencies and shared workflow friction to improve overall system health.
Ignoring technical debt Teams push features fast but accumulate massive maintenance costs that slow future development. TargetBoard acts as an agentic operational intelligence layer to detect AI-induced complexity before it reaches production.

Anti-Pattern One: Measuring Lines of Code

Tracking lines of code is the fastest way to destroy developer effectiveness. This metric was always flawed, but Artificial Intelligence makes it actively dangerous. AI tools can generate thousands of lines of boilerplate code in seconds. If you measure volume, your metrics will look incredible while your codebase becomes an unmaintainable mess. You need to measure the value delivered to the customer instead of the raw output.

Anti-Pattern Two: Tracking Individual Instead of Team Performance

Software development is a complex team operation. Tracking team performance vs. individual performance is a critical distinction. Pitting developers against each other creates a toxic environment where senior engineers refuse to help juniors. If a lead engineer spends all week reviewing pull requests, their individual commit metrics will drop. Yet their work is exactly what keeps the entire system moving. You must measure how the team delivers as a unified unit.

Anti-Pattern Three: Sacrificing Quality for Speed

Executives often demand faster delivery without understanding the speed vs. quality tradeoffs. Pushing teams to ship faster without investing in automated testing leads to a massive spike in production failures. The system will eventually grind to a halt under the weight of its own technical debt. True predictability requires balancing feature development with continuous system maintenance.

Why Dashboards Fail: Moving from Scoreboards to Systemic Intelligence

Dashboard fatigue is a very real problem for modern engineering leaders. You have a Jira dashboard for issue tracking and a GitHub dashboard for pull requests. These Jira and GitHub data silos provide conflicting signals. Jira says the sprint was successful, but GitHub shows massive code review churn.

This disconnect forces leaders to rely on intuition rather than data. You can't make confident execution decisions when your tools refuse to talk to each other. Dashboards are static scoreboards that show you what happened yesterday. They don't tell you why it happened or what you should do about it today.

TargetBoard is an agentic operational intelligence platform that helps leadership teams understand how execution is performing, why it is changing, and how to respond. It unifies performance data across systems into a trusted model and deploys domain-expert AI agents to translate insights into decision-ready inputs that guide execution.

Feature Old Way (Dashboards) New Way (Agentic Intelligence)
Data Integration Fragmented Jira and GitHub data silos require manual exports. Unified operational model connects planning, code, and delivery automatically.
Analysis Static charts force leaders to guess why metrics are changing. Domain-expert AI agents explain exactly why performance shifted.
AI Impact Blind to the difference between human and AI-generated code. Exposes how AI code generation impacts review time and system complexity.
Outcome Dashboard fatigue and delayed reactions to delivery risks. Confident execution decisions based on real-time systemic visibility.

Stop Tracking Metrics, Start Understanding Your Delivery System

Tracking software development performance metrics isn't the end goal. The goal is to build a reliable delivery system that consistently drives business outcomes. Staring at a static scoreboard won't help you identify the hidden complexity introduced by Artificial Intelligence or the workflow friction slowing down your senior engineers.

You must shift your focus from measuring isolated outputs to understanding your interconnected systems. This systemic visibility gives you a clear framework for your next resource allocation discussion or board meeting. It replaces guesswork with actual delivery predictability. Take a hard look at your current reporting structure and ask yourself if your data actually helps you make better execution decisions, because visibility without action is just overhead. If it just gives you another number to report, it's time to upgrade your operational intelligence.

Business

What is Development Cycle Time

You sit in the weekly leadership meeting, and the C-suite wants to know why a critical feature is two weeks late. You look at your Jira dashboard and see development cycle time dropping. Your developers are writing code faster than ever thanks to AI coding assistants, so you expect faster releases. Yet your end-to-end delivery is stalling. Conflicting data signals across Jira, GitHub, and Slack make it impossible to explain why execution is changing. You have the metric, but you lack the operational intelligence to understand it. This erodes executive trust in your reporting and destroys delivery predictability. True engineering velocity comes from reliable system flow, not frantic local optimizations. Understanding this shift gives you a clear framework to diagnose delivery friction and regain confidence in your timelines.
May 10, 2026
5 min read

What is Development Cycle Time?

Development cycle time is the total amount of time it takes for an engineering team to complete a single task from the moment work begins until it is deployed to production.

This metric originated in Lean manufacturing to measure inventory flow. Today it serves as a critical diagnostic signal for software development cycle time. Traditional engineering leaders often make the mistake of treating this as a pure speed metric. I have watched organizations gamify cycle time to push developers to type faster. That approach inevitably leads to developer burnout and lower quality code. A low cycle time means nothing if the code requires massive rework later.

You must view development cycle time as a measure of system flow and cross-team friction. It tells you exactly where work stalls. Tracking this accurately is the only way to ensure delivery predictability across your entire engineering organization.

Cycle Time vs. Lead Time: Understanding the Difference

The difference between cycle time and lead time comes down to when the clock starts. Lead time begins the moment a customer requests a feature, while cycle time begins the moment a developer actually starts writing code for that feature.

Lead time for changes measures your entire product management and prioritization process. Software cycle time isolates the engineering execution phase. You need both to understand your true time to market.

Metric Start Point End Point What It Measures
Lead Time Customer request created Feature deployed to production Overall organizational responsiveness and planning efficiency.
Cycle Time Developer makes the first commit Code deployed to production Engineering system flow and execution efficiency.

‍

The 4 Key Components of Development Cycle Time

You can't fix a bottleneck until you know exactly where it lives. The cycle time formula breaks down into four distinct phases. Tracking the transition between these phases reveals where your system loses momentum.

Cycle Time Phase Ideal State Real-World Executive Reality
Coding Time Developers write clean code quickly. AI accelerates output, but introduces hidden complexity.
PR Pickup Time Reviewers claim pull requests immediately. Context switching delays pickup as engineers focus on their own tickets.
Review Time Fast approvals with minor feedback. Massive back-and-forth churn due to complex AI-generated code.
Deploy Time Automated pipelines ship code instantly. Manual testing requirements and batching create deployment traffic jams.

Phase 1: Coding Time

Coding time measures the lifespan from the developer's first commit to the moment they issue a pull request. This phase tracks active creation. AI tools have drastically reduced coding time across the industry.

Phase 2: Pull Request Pickup Time

PR pickup time tracks the idle period between a developer opening a pull request and a peer beginning the review. That's rarely a skill issue. It's almost always a coordination and visibility problem.

Phase 3: Review Time

Review time measures the span from the first review comment to the final approval. That's the most common bottleneck in modern software delivery. Fast coding times often hide severe inefficiencies here, as reviewers struggle to understand massive blocks of undocumented code.

Phase 4: Deploy Time

Deploy time covers the final span from a code merger to a production release. Heavy manual testing requirements and complex release train schedules often inflate this metric, leaving finished code sitting idle.

How to Measure Development Cycle Time Accurately

To measure development cycle time accurately, you must connect your issue tracking software to your version control system to track the exact timestamps of commits, pull requests, reviews, and deployments.

Relying solely on DORA metrics or isolated Jira boards gives you an incomplete picture. DORA metrics provide useful signals for deployment frequency and stability, but they do not provide system-level visibility into why a specific workflow is stalling. Fragmented tools make measurement incredibly difficult. Jira says a ticket is in progress, but GitHub shows the code has been sitting in review for four days. You can't manually merge this data to calculate accurate sprint velocity. You need a unified operational model to see the truth.

Step-by-Step Guide to Establishing a Baseline

You must standardize your data inputs before you can diagnose your delivery pipelines. Follow these steps to build a reliable measurement foundation.

  1. Standardize issue states: Align your Jira workflow statuses across all engineering teams so that "In Progress" means the exact same thing for every developer.
  2. Connect version control: Link your Git repositories directly to your ticketing system to capture automated timestamps for commits and pull requests.
  3. Isolate idle time: Configure your reporting to separate active coding time from passive waiting periods like PR pickup time.
  4. Track deployment triggers: Map your CI/CD pipeline events to your cycle time tracking to measure continuous delivery performance accurately.

Connecting these steps gives you actionable insights to improve workflow efficiency and continuous delivery.

Why "Reducing" Cycle Time Fails 

When you push teams to just code faster, you fall into the local optimization trap. A local optimization improves one small part of the process while degrading the whole system. Forcing engineers to close tickets rapidly often leads to sloppy commits, so you see a massive spike in rework and code churn during the review phase. This creates a severe downstream delivery impact. You must measure system flow outcomes rather than isolated speed metrics to protect your delivery timelines.

Local Optimization Metrics System Flow Outcomes
Lines of Code Written Measures sheer volume without accounting for quality, often increasing technical debt.
Individual Developer Velocity Gamifies speed for one person, causing cross-team friction and siloed knowledge.
Number of PRs Opened Encourages fragmented work, leading to integration headaches and deployment traffic jams.
Raw Cycle Time Reduction Forces rushed handoffs, resulting in higher defect rates and massive rework loops.

AI-Generated Code: The Hidden Delivery Bottleneck

I see this constantly with modern engineering teams. You roll out AI coding assistants, and coding time drops to near zero. Developers produce massive blocks of code in minutes. Management often views these tools purely as cycle time accelerators, but they fail to account for the resulting review churn.

AI-assisted developers write code up to 50% faster, yet PR cycle times often increase due to the cognitive load placed on reviewers.¹ AI-generated code introduces hidden complexity, so reviewers have to spend hours untangling logic they didn't write. This creates a massive delivery bottleneck and severe maintainability risks. You accelerated the easiest part of the job while gridlocking the hardest part.

Visualizing System Flow vs. Isolated Team Speed

Engineering leaders often mandate a smaller pull request size to speed up reviews. This sounds logical in theory. In reality, forcing developers to break a single feature into ten tiny PRs creates a coordination nightmare. Reviewers lose the broader context, so defect patterns increase during integration. That's especially true when working with highly complex, interdependent legacy codebases that skew standard benchmarks.

Your agile cycle time might look great on a dashboard, but your actual system flow grinds to a halt. You must enforce strict Work In Progress (WIP) limits to balance batch size with the cognitive load required to review the entire feature.

How to Reduce Development Cycle Time Systemically

True optimization comes from lean manufacturing principles. You don't ask the assembly line workers to move their hands faster. You eliminate the wait time and idle time between stations.

In software delivery, this means reducing handoffs and automating your deployment frequency. You want work to flow continuously without sitting in a queue waiting for manual intervention. Elite performers achieve high deployment frequency by minimizing handoffs rather than pushing individual engineers to type faster.²

Step-by-Step Framework for Identifying Bottlenecks

Use this framework to find the root cause of your delivery delays and fix your workflow coordination.

  1. Map cross-team dependencies: Identify every point where a ticket requires approval, security clearance, or input from a different department to spot coordination breakdowns.
  2. Analyze review churn: Track how many times a PR bounces between the author and the reviewer to spot code complexity and architecture issues.
  3. Enforce WIP limits: Restrict the number of active tickets per developer to force the completion of existing work before new work begins.
  4. Perform root cause analysis: Trace failed deployments back to their origin to see if a rushed review or an unclear requirement caused the defect.

Moving from Dashboards to Operational Intelligence

Having a dashboard that tells you your cycle time is nine days doesn't help you fix it. Passive metrics require you to guess what went wrong. You need operational intelligence to explain why performance is changing. This requires shifting from basic executive reporting to an agentic system that understands delivery trade-offs and system flow.

TargetBoard is an agentic operational intelligence platform that helps leadership teams understand how execution is performing, why it's changing, and how to respond. TargetBoard deploys domain-expert AI agents across your connected systems to act as expert analysts. Instead of just showing a red line on a graph, TargetBoard explains that cycle time spiked because AI-generated code in a specific repository caused a 40% increase in review churn. It translates raw data into objective signals you can use to make immediate resource decisions.

System Type Approach to Metrics Executive Value
Traditional Metric Dashboards Displays raw numbers like a 9-day cycle time or 3 deploys per week. Forces leaders to manually investigate the root cause across fragmented tools like Jira and GitHub.
TargetBoard Operational Intelligence Deploys AI agents to explain why metrics shift and where execution is breaking down. Provides decision-ready insights, linking specific bottlenecks to code complexity, AI impact, or coordination gaps.

‍

Leverage Predictability Over Pure Speed

Pushing for speed without predictability is an organizational failure. Keep in mind that no single metric provides a complete picture of engineering health. True engineering velocity requires reliable system flow. When you stop treating development cycle time as a stopwatch and start treating it as a diagnostic signal, you regain delivery predictability. Understanding these patterns gives you a clear framework to align your engineering execution with your business goals and confidently forecast your next major release.

Technical

What is technical debt business risk

You're sitting in the quarterly board meeting trying to explain why the core product release will miss its target date by a month. The CEO asks for the root cause, so you cite technical debt. You present fragmented Jira exports showing open tickets and delayed story points to quantify the problem. Those metrics show that velocity is dropping, yet they completely fail to explain why the slowdown is happening. This disconnect between engineering reality and boardroom expectations erodes trust. You're left relying on intuition to justify necessary architectural tradeoffs. Every engineering leader faces the pressure to balance immediate feature delivery with long-term system stability. The challenge is no longer just measuring developer output. The real mandate is translating hidden codebase friction into a clear business risk so you can protect delivery predictability.
July 19, 2026
5 min read

What Is Meant by Technical Debt?

Problem: Teams often prioritize short-term deadlines over sustainable software architecture to meet immediate market demands. This creates hidden structural compromises within the codebase.

Solution: Engineering leaders must treat this deferred maintenance as a measurable business liability that directly impacts engineering velocity and resource allocation.

Software developer Ward Cunningham coined the financial analogy of technical debt to explain this dynamic. Borrowing time to release faster is a perfectly acceptable business strategy, but you incur interest on that loan.

You must pay down the principal through regular software maintenance. If you fail to do this, the compounding interest eventually paralyzes the engineering team. Every new feature requires modifying fragile code, so the cost of future changes becomes prohibitively expensive.

The Consequences of Unpaid Interest on Engineering Velocity

Unmanaged debt inevitably causes critical missed delivery targets. A VP of Engineering might commit to a Q3 product launch based on current team capacity. But technical drag from a legacy billing module quickly causes endless pull request churn.

Developer productivity plummets as engineers spend weeks untangling fragile logic instead of building new capabilities. The result is severe system instability and missed business outcomes. Slower feature releases give competitors the advantage and directly impact revenue pacing.

What Are the 4 Types of Technical Debt?

Engineering leaders categorize tech debt by analyzing the intent behind the decision and the context in which it occurs. This framework helps teams distinguish between strategic technical tradeoffs and simple carelessness.

Debt Category Intent Level Primary Cause Business Impact
Prudent and Deliberate Intentional Strategic choice to hit a critical market window. Controlled risk with a planned refactoring phase.
Reckless and Deliberate Intentional Ignoring software architecture best practices to save time. High operational overhead and frequent system instability.
Prudent and Inadvertent Unintentional Evolving business needs outgrow the original design. Gradual slowdowns requiring eventual system modernization.
Reckless and Inadvertent Unintentional Lack of developer experience or poor technical leadership. Severe technical debts that paralyze future development.

Reckless and Deliberate Technical Debt

This occurs when a team explicitly knows the right way to build a feature but chooses the wrong way just to meet short-term deadlines. They might hardcode values or skip essential automated testing entirely.

The team makes these technical tradeoffs without any plan to fix the underlying issues later. This behavior signals a toxic engineering culture and guarantees future delivery failures.

Prudent and Deliberate Technical Debt

Strategic leaders use this category as a calculated business lever. A team might choose a monolithic architecture over microservices to test a minimum viable product and accelerate time-to-market.

The leadership team understands the limitations and actively schedules future sprints to pay down the debt once the product proves its value. This approach aligns engineering efficiency directly with positive business outcomes.

Reckless and Inadvertent Technical Debt

This debt accumulates when teams simply don't know any better. A junior engineering team might build a complex feature without understanding the design patterns required to scale it.

The resulting code is sloppy and introduces massive operational overhead. Leaders must address this through better training, stricter review processes, and improved engineering efficiency standards.

Prudent and Inadvertent Technical Debt

Even the most talented teams accumulate this debt over time. You might build a brilliant system using the best available practices, but industry standards evolve and user demands shift two years later.

The original design no longer fits the current reality. This impacts codebase health and turns previously modern applications into legacy systems. The result is rising pull request churn as engineers struggle to adapt the old code to new requirements.

Will AI Cause Technical Debt? The Hidden Cost of AI-Generated Code

The rapid adoption of AI coding assistants fundamentally changes how organizations accumulate risk. AI tools allow developers to generate thousands of lines of code in seconds, so this massive spike in output creates a severe bottleneck downstream.

Human reviewers can't match the pace of the machine. AI doesn't necessarily write bad code, but it dramatically increases code complexity. Developers often accept AI suggestions without fully understanding the underlying logic, and this introduces hidden vulnerabilities.

This dynamic forces human reviewers to spend days deciphering convoluted pull requests. The operational overhead skyrockets, and codebase health deteriorates rapidly. You can't solve this modern friction with legacy metrics.

Impact Area Traditional Development AI-Accelerated Development
Code Generation Steady output limited by human typing speed. Exponential output creating massive review queues.
Review Process Reviewers understand the context of human-written logic. Reviewers struggle to verify complex AI-generated logic.
Workflow Friction Predictable cycle times based on known team capacity. Spikes in pull request churn as untested code gets rejected.
System Risk Debt accumulates through known architectural compromises. Debt accumulates silently through hidden code complexity.

How to Evaluate and Prioritize Technical Debt Reduction

Engineering leaders can't fix every imperfect line of code. You must treat technical debt reduction as an ongoing resource allocation exercise. According to a 2022 McKinsey report, developers spend approximately 30% of their time managing technical debt1. A 2018 Stripe Developer Coefficient study found that engineers lose up to 17 hours a week to software maintenance and bad code2.

You can't afford that level of waste, so you need a systematic approach to identify which infrastructure debt actually threatens your business. Use this step-by-step guide to evaluate and prioritize your refactoring efforts:

  1. Map debt to business outcomes because you must identify which legacy systems directly impact revenue generation or user retention.
  2. Measure sprint velocity drag, which requires tracking how much time engineers waste on maintenance tasks versus building new features.
  3. Assess infrastructure debt risk. Evaluate aging servers or outdated deployment pipelines that could cause critical outages.
  4. Target high-churn areas. Prioritize refactoring in files that change frequently and cause the most testing failures.

Identifying Workflow Bottlenecks and Cycle Time Delays

You might track a drop in sprint velocity, but that metric alone doesn't reveal the root cause of the slowdown. A common leadership mistake is tracking metrics without contextualizing the underlying workflow bottlenecks.

A single complex pull request can cause severe code review delays, which blocks multiple engineers and creates a cascading delay across cross-team dependencies. You must perform a root cause analysis to understand exactly where the friction lives. Connecting code complexity directly to cycle time delays gives you the objective data needed to prioritize fixes before they derail your release schedule.

How to Measure and Justify Technical Debt as a Business Risk

Engineering leaders struggle to justify debt to non-technical stakeholders because they lack objective data connecting code quality to business outcomes. Fragmented data across Jira and GitHub erodes trust in the boardroom. You are left relying on intuition rather than concrete signals to defend your resource allocation.

Standard frameworks and Agile methodologies provide signals of a slowdown, but they don't provide an understanding of the underlying causes when operating at scale. To properly evaluate operational friction and justify refactoring, leaders must implement an operational intelligence layer.

TargetBoard is an agentic operational intelligence platform that helps leadership teams understand how execution is performing, why it is changing, and how to respond. It connects data across company systems, interprets performance through operational intelligence, and uses domain-expert AI agents to guide execution decisions.

Measurement Approach Data Source Business Impact Analysis Executive Trust Level
Fragmented Metrics & Manual Reporting Disconnected Jira exports and GitHub logs. Fails to explain why delivery predictability is dropping. Low. Leaders rely on intuition to justify tradeoffs.
TargetBoard's Agentic Operational Intelligence Unified operational model across all engineering systems. Connects PR churn and complexity directly to cycle time delays. High. Provides objective signals to confidently justify refactoring.

Who Is Responsible for Technical Debt?

Everyone who touches the product lifecycle shares responsibility for it. Product managers create debt when they push unrealistic deadlines without consulting engineering constraints. Developers create it when they skip technical documentation or ignore established coding standards.

Leadership ultimately owns the debt because they control the culture and the budget. You must foster an environment where engineering efficiency balances perfectly with speed to market. When you treat debt as a shared business reality, teams can openly discuss tradeoffs without fear of blame.

Building the Executive Case for Refactoring

You can't walk into a board meeting and ask for a month to clean up bad code. You must frame the initiative around risk mitigation and delivery tradeoffs. Follow this step-by-step guide to build a compelling executive case:

  1. Quantify the business risk since stakeholders need to see exactly how much time the team wastes modifying fragile code each sprint.
  2. Highlight the delivery tradeoffs. Explain that shipping the next feature on time requires stabilizing the underlying architecture first.
  3. Propose a targeted refactoring plan, which means asking to allocate a specific percentage of capacity to debt reduction instead of requesting a full feature freeze.
  4. Define the expected ROI. Promise a measurable improvement in cycle time and system stability once the work is complete.

Moving From Reactive Reporting to Proactive Execution Predictability

Understanding these patterns gives you a clear framework for your next resource planning session. You no longer have to wait for a major outage to address technical entropy. You can monitor codebase health continuously and catch bit rot before it impacts your customers.

Start by auditing your current reporting systems to see if they actually explain why performance is changing. Move away from manual data aggregation and implement an intelligence layer that connects code complexity to delivery outcomes. This shift allows you to stop reacting to delayed releases and start driving predictable execution across your entire engineering organization.

‍

Technical

Code churn measure manage rework

You open your reporting dashboard and see a 30% spike in code churn across two critical delivery teams. The numbers shifted overnight, but the dashboard can't tell you why. When the board asks why the upcoming release is delayed, quoting a high metric doesn't answer the question. You need to explain if this is healthy prototyping for a complex new feature or destructive rework destroying your sprint predictability. Incomplete data erodes trust in engineering reporting and makes it impossible to allocate resources confidently. We don't just measure engineering performance. We explain why it's changing. Understanding this metric requires connecting codebase health to delivery risk, so you can stop reacting to dashboards and start managing execution.
July 19, 2026
5 min read

What Does Churn Mean in Coding?

Code churn refers to the percentage of a developer's code that gets rewritten, modified, or deleted shortly after being written. Teams typically measure this within a strict 3-week timeframe. If an engineer writes a function and rewrites it two days later, that action is code churn.

When you ask what is code churn, the answer lies in your version control systems. These systems track the exact lines added, modified, and deleted before code reaches production. High churn over a short period indicates that the code was not stable upon its first commit.

How Do You Calculate Code Churn?

You calculate churn code by dividing the number of lines modified or deleted within a specific period by the total number of lines added during that same period. You then multiply the result by 100 to get a percentage.

Here is how you calculate code rework step by step:

  1. Define your measurement window, typically 21 days.
  2. Pull data from Git, GitHub, or GitLab to count the total lines of code (LOC) added.
  3. Count the total lines modified or deleted from that exact same batch of code.
  4. Divide the modified lines by the total lines added.
  5. Multiply by 100 to find your percentage.

Keep in mind that commit frequency alone doesn't equal churn. A developer might commit frequently to save progress without rewriting the same lines of code.

Is Code Churn Bad? Understanding Healthy Prototyping Versus Destructive Rework

Code churn isn't inherently bad. It is a natural part of the software development lifecycle. The goal is not to eliminate it, but to differentiate healthy prototyping from destructive rework.

Healthy churn happens during the initial design phase when engineers iterate on complex problems. Destructive rework happens when developers constantly rewrite code due to unclear requirements or accumulating technical debt. This unhealthy behavior creates workflow bottlenecks that delay delivery.

Characteristic Healthy Prototyping Destructive Rework
Timing Happens early in the development cycle. Happens late in the review cycle or just before release.
Cause Planned refactoring or exploring complex logic. Unclear requirements or addressing technical debt.
Impact Results in cleaner and more maintainable code. Causes workflow bottlenecks and reduces delivery predictability.

What Is an Example of Churn?

Consider a scenario where a team is updating an aging payment gateway. Working with legacy code naturally requires trial-and-error coding. A developer submits a pull request, and the reviewer requests multiple architectural changes because the original product requirements were vague.

The developer spends the next four days rewriting the same 500 lines of code. This is review churn. It isn't a developer defect, but a systemic issue caused by poorly defined upstream requirements. The engineering effort is wasted, so the entire sprint slows down.

Is a 5% Churn Rate Good? Establishing Your Baseline

A 5% churn rate is exceptionally low and generally indicates a highly stable, well-understood codebase. Industry research from LinearB1 suggests that a healthy churn threshold sits around 20% for most engineering teams. Appfire2 confirms that exceeding this limit often points to systemic process failures rather than individual coding errors.

But a single baseline doesn't apply universally. A team building a new product from scratch will naturally see higher churn than a team maintaining an established application. You must establish a baseline based on historical engineering metrics for each specific project. If a team's baseline is 15% and it suddenly spikes to 40%, you have a clear signal that sprint velocity is about to drop.

The Root Causes of Code Rework (and How to Fix Them)

A sudden spike in code rework is a symptom of a deeper operational failure. You must perform a root cause analysis to understand why engineers are rewriting their work. If you ignore the underlying issues, you will see an increase in software bugs and a massive slowdown in cycle time.

Common drivers of destructive rework include:

  • Unclear product requirements that shift during the sprint.
  • High code complexity that makes new features difficult to integrate.
  • Inefficient review cycles that force developers to rewrite logic multiple times.
  • An influx of unreviewed AI-generated boilerplate code.

Unclear Requirements and Scope Creep

Last quarter, a core delivery team saw their pull request review churn spike by 40%. The dashboard highlighted the developers as the problem, but the real issue was shifting requirements from product management. The developers were building features based on vague tickets.

When reviewers finally saw the code, they requested massive architectural changes to align with the actual business goals. These communication breakdowns force developers to rewrite working code. This constant rework inevitably leads to developer burnout and severely damages your overall delivery predictability.

Skill Gaps and Complex Initial Design

Engineers sometimes dive into a problem without a clear architectural plan. This trial-and-error coding results in massive rewrites before the code even reaches the review stage. The issue often stems from failing to apply SOLID design principles early in the process.

When code complexity is high, every new addition breaks existing functionality. Developers have to constantly rewrite their logic to force the new code to fit. You can fix this by enforcing technical design documents before any coding begins.

System-Level Thinking: Diagnosing Workflow Bottlenecks

Avoid using code churn as an individual performance punishment tool because high churn is a systemic workflow diagnostic rather than a developer defect.

You can identify these bottlenecks by tracking your workflow behavior through these steps:

  1. Track the total volume of pull requests (PRs) stuck in the review phase for more than three days.
  2. Analyze your code reviews to see if reviewers are requesting stylistic changes or fundamental architectural rewrites.
  3. Monitor your CI/CD pipelines to identify if automated tests are forcing constant rework before deployment.
  4. Correlate these delays directly to your overall delivery predictability metrics to see where execution breaks down.

The Impact of Artificial Intelligence on Code Churn: Why Your Metrics Are Shifting

Generative AI is fundamentally altering how software is built. AI coding tools increase raw output, but they often severely bog down review cycles. Developers can generate hundreds of lines of AI-generated code in seconds. This creates an illusion of high productivity.

But this code often contains hidden complexity and massive code duplication. Reviewers have to spend hours untangling the logic, and they frequently force the original developer to rewrite the entire section. This drives up your review churn and creates a massive bottleneck.

Metric Impact Human-Written Code AI-Generated Code
Initial Output Speed Slower, deliberate pacing based on manual typing. Extremely fast, generating massive blocks of logic instantly.
Review Churn Moderate, usually focused on logic gaps or business rules. Exceptionally high, often requiring total rewrites to fix hidden complexity.
Defect Rate Predictable and tied to known team skill levels. Unpredictable, often introducing subtle structural flaws.

Moving Beyond Metrics to Operational Intelligence

Industry frameworks like DORA metrics are excellent tools for measuring engineering performance. They provide valuable signals about speed and reliability. But these metrics only tell you that your cycle time dropped or your defect rate increased. They don't tell you why the change happened.

To manage the influx of AI-generated code and protect your delivery predictability, you need more than passive reporting. You need an operational intelligence layer that connects codebase health directly to delivery risk.

TargetBoard is an agentic operational intelligence platform that helps leadership teams understand how execution is performing, why it is changing, and how to respond. It connects data across company systems, interprets performance through operational intelligence, and uses domain-expert AI agents to guide execution decisions.

Measurement Approach Core Capability Outcome for Engineering Leaders
Standard Dashboards (DORA) Tracks deployment frequency and lead time for changes. Provides a historical signal of delivery speed but lacks contextual understanding.
TargetBoard Uses agentic operational intelligence to analyze codebase health across systems. Explains why performance changes and differentiates AI-generated code to catch delivery risk before merging.

This continuous intelligence allows you to differentiate between human-written and AI-generated code churn. You catch delivery risk before it gets merged, ensuring your execution stays aligned with your planning.

Realigning Your Engineering Delivery Predictability

You can't improve what you don't understand. When you reduce destructive rework, you directly improve your delivery predictability. Teams stop wasting engineering effort on endless review cycles and start shipping reliable features on time.

You must stop treating engineering metrics as a retroactive scorecard. Delivery failures are system issues rather than developer defects. You need to use operational intelligence to catch workflow bottlenecks in real time. This approach restores trust in your reporting and gives you the confidence to allocate resources effectively.

Technical

Change Failure Rate

You look at your engineering dashboard and see an Elite change failure rate. Everything looks green, so you report to the board that delivery is predictable and stable. Yet your engineering teams are drowning in silent rework and massive pull request churn behind the scenes. This disconnect happens because standard measurement acts as a lagging indicator that fails to capture hidden complexity. Organizations have strong systems for measuring software delivery performance but lack a consistent system for interpreting it. Leaders can see the metrics shift over time, yet they struggle to understand why performance is changing or where workflow bottlenecks are emerging. That gap creates delayed detection and erodes trust in reporting. You need objective data to justify engineering return on investment and build trust with leadership. Achieving that requires moving beyond passive dashboards to expose the workflow friction throttling your delivery speed.
May 10, 2026
5 min read

What is a Change Failure Rate?

Change failure rate (CFR) measures the percentage of code deployments that result in a failure in production. The goal is to track how often your team pushes code that requires immediate remediation.

This metric serves as a critical counterbalance to deployment frequency. Optimizing strictly for speed often damages quality, so tracking failures ensures your team maintains system stability while shipping features faster. Engineering leaders use this DORA change failure rate signal to balance the inevitable tradeoff between quality versus speed.

The Formula to Calculate Change Failure Rate

Calculating this metric requires standardizing what counts as a deployment and what counts as a failure. You must define these terms consistently across your incident response tools and code repositories.

To calculate change failure rate, use this formula:

(Number of Failed Changes / Total Number of Changes) × 100

  • Total changes: The absolute number of production deployments your team executes over a specific time period.
  • Failed changes: Any deployment that directly causes production failures and requires immediate intervention.

What is an Acceptable Change Failure Rate (DevOps Research and Assessment Benchmarks)?

Industry benchmarks categorize engineering teams into performance tiers based on their ability to ship code reliably. According to the 2023 Accelerate State of DevOps Report by Google Cloud, you can measure change failure rate against these established standards to gauge your baseline delivery health.

Performance Tier Benchmark Target Operational Reality
Elite performance 0% to 5% Teams use comprehensive automated testing to catch defects before production.
High performers 0% to 15% Teams maintain stable delivery but occasionally experience workflow friction.
Medium / low performers 16% to 64% Teams rely on manual testing and frequently push unstable code that requires immediate fixes.

‍

How Do You Define Change Failure? 

Most engineering leaders limit the definition of failure strictly to hotfixes and rollbacks. This narrow scope misses the broader picture of system degradation.

If a deployment introduces massive technical debt or causes degraded service that doesn't trigger a critical alert, your dashboard will still show a success. This forces leaders to rely on intuition because incomplete data undermines the credibility of engineering reporting. Redefining failure for the modern era means looking at the entire workflow rather than just the final production state to capture the true cost of service patches.

What Are the Four Types of Failure in Modern Software Delivery?

Modern software delivery systems experience friction long before a catastrophic outage occurs. You must expand your definition of failure to capture the hidden costs of code delivery.

Failure Type Description Impact on Delivery
Catastrophic production outages Complete system failures that halt core business operations. Causes immediate financial loss and triggers emergency incident response.
Silent performance degradation Code that slows down service speed or user experience without triggering critical alerts. These silent failures erode customer trust slowly and create hidden drag.
Code reversions and hotfixes Unstable deployments that require immediate service patches or rollbacks. Code reversions disrupt planned work and force engineers to context-switch into reactive modes.
Technical debt accumulation High-complexity code that merges due to review fatigue and poor oversight. Technical debt accumulation increases future lead time for changes and introduces unintended consequences downstream

The False Green Dashboard: Common Measurement Pitfalls

A dashboard can easily show an Elite status while your team is actually dealing with high pull request churn. This happens when teams game the metric or pollute the data with inconsistent definitions.

One common mistake is including fix-only deployments in the denominator of your calculation. If you push five hotfixes to resolve a single incident, counting those fixes as new deployments artificially lowers your failure rate. Another pitfall involves poor incident attribution, where third-party cloud outages are counted against internal team performance. These practices create a false sense of stability that operational intelligence must correct to restore trust in your reporting.

How to Audit Your Incident Attribution Data Step by Step

Executives must ensure their teams map incidents accurately across the software delivery lifecycle. Messy data makes it impossible to identify root causes and delays critical decision-making.

  1. Standardize your tags: Mandate that all teams use identical tagging conventions for bugs and incidents across Jira and GitHub because inconsistent tags hide root causes.
  2. Separate external failures: Filter out third-party provider outages from your core calculation to isolate your team's actual performance.
  3. Exclude remediation deployments: Remove fix-only deployments from your total changes count to prevent artificially deflating your failure rate.
  4. Connect incidents to code: Require root cause analysis and postmortems to link every production failure back to the specific pull request that introduced it.

The Impact of Artificial Intelligence-Assisted Engineering on Codebase Health

The rapid adoption of AI coding tools fundamentally changes how we measure delivery risk. These tools drastically increase developer output, so teams write and submit code faster than ever before. Yet this sheer volume of artificial intelligence-generated code contributions introduces unseen complexity into your repositories.

Downstream reviewers simply can't keep up with the flood of new pull requests. This imbalance creates severe review fatigue, where engineers lose the capacity to deeply inspect code for architectural flaws or long-term maintainability issues. The code compiles and passes basic tests, but the underlying structural health of the system degrades quietly.

Visualizing Systemic Risk: How Workflow Friction Causes Delayed Failures

Unmanaged complexity builds up in your repositories and creates massive workflow friction during the review stage. When a dense, highly complex pull request sits in review for days, engineers eventually rubber-stamp the approval just to clear their queues.

That code merges, sits in the pipeline, and fails days later in production. You then spend valuable engineering cycles on bug prioritization instead of shipping new features. The failure looks like a sudden event on your dashboard, but the root cause was the hidden complexity that bottlenecked your workflow days earlier.

Moving from Lagging Metrics to Predictive Intelligence

Measuring a failure after it hits production is fundamentally a lagging indicator. Industry frameworks provide useful signals about your software delivery performance, but they don't provide an understanding of why that performance is changing. You need to know where risk enters your system before the code ships to production.

TargetBoard is an agentic operational intelligence platform that helps leadership teams understand how execution is performing, why it's changing, and how to respond. It connects data across company systems, interprets performance through operational intelligence, and uses domain-expert artificial intelligence agents to guide execution decisions.

By surfacing hidden risks like review fatigue, code anomalies, and workflow bottlenecks during the actual code review process, TargetBoard allows you to neutralize the root causes of failure before they merge. This shifts your posture from reactive reporting to proactive delivery confidence, ultimately driving true engineering efficiency.

Proven Tactics to Reduce Change Failure Rate Before Production

You can actively prevent production failures by changing how your team handles code before it reaches the main branch. Aligned with the foundational Continuous Delivery principles established by industry experts like Jez Humble and Martin Fowler, shifting quality checks left is critical.

  • Implement shift-left testing: Move security and performance testing to the initial commit phase to catch defects before they reach the review stage.
  • Use feature flags: Decouple deployments from releases to test code safely in production without exposing all users to potential bugs.
  • Strengthen continuous integration and continuous delivery: Build robust pipelines that automatically reject code that fails baseline quality checks.
  • Standardize automated deployments: Remove manual human intervention from the release process to eliminate configuration errors.

Balancing Deployment Frequency with True System Stability

Pushing for speed without guardrails creates severe systemic tradeoffs. You must balance how fast you ship with how well your system actually runs.

Strategic Focus The Outcome The Tradeoff
Optimizing for deployment frequency Teams ship smaller batches of code constantly. High speed can mask poor codebase health if automated testing is weak.
Optimizing for quality Teams implement rigorous, multi-stage review processes. Heavy governance increases your lead time for changes and slows down feature delivery.
Balanced operational intelligence Teams use data to flag only high-risk pull requests for deep review.

Requires connecting cross-system data to accurately predict where failures will occur.

Expanding Your Definition of Failure Across Workflows

Redefining failure requires you to look beyond standard production deployments and measure the friction happening inside your daily workflows.

  1. Track pull request churn: Measure how many times a piece of code bounces between the author and the reviewer before merging, since high churn indicates hidden complexity.
  2. Monitor silent degradation: Set alerts for code that slows down system performance or increases cloud costs without triggering a hard outage, because these silent failures erode customer trust.
  3. Connect codebase health to delivery speed: Analyze how rising technical debt correlates with slower sprint velocity over time, which reveals the true cost of rushed code.
  4. Measure the cost of rework: Quantify the engineering hours spent fixing bugs instead of building net-new value to expose true systemic tradeoffs.

Conclusion: Stop Reacting to Metrics and Start Driving Execution

Your dashboard is only as valuable as the decisions it enables. Passive metrics show you what broke, so you must adopt active operational intelligence to see why it broke. Understanding these patterns gives you a clear framework to improve engineering efficiency and ensure long-term delivery predictability. Moving away from lagging scorecards allows you to scale your software delivery performance safely and build trust with your board.

Ready to See a Demo?

Contact Us