The Agentic Coding Payoff: Experimentation Velocity
Most teams measure AI impact on software delivery the same way: before and after comparisons on tasks that were already in the plan. A migration that used to take three weeks now takes two days. A bug that took a week to diagnose gets resolved in an afternoon. These are real wins.
But they’re measuring the floor, not the ceiling.
Speed on planned work is a productivity gain. The ability to attempt unplanned work is a strategic one. A team that ships 20% faster will be faster, sure. A team that runs ten experiments a quarter instead of two will, over time, find things the faster team never looked for. Both matter, but only one of them compounds over time.
In my last post, I argued that coding agents don’t replace platforms — they make building blocks, guardrails, and durable systems more important. That post was about what agents need to work well. This one is about what becomes possible once they do: a shift from delivery speed to experimentation velocity.
What Experimentation Velocity Looks Like
When the cost of trying something drops far enough, people try things they never would have planned.
Our engineering teams have been running structured agentic hackathons and sprints — short, intense blocks where engineers work on projects, with AI agents as full execution partners. What comes out of those sessions isn’t only incremental improvement on planned work. It’s entirely new bets that only exist because the cost of placing them collapsed.
Some of the work included:
- POC for managing human & agent interactions: Unblocked the architecture decision in days instead of weeks.
- Frontend analytic event migration project: Turned a discovery task into a full delivery.
- Backend service: 2-month schedule pull-forward on a foundational platform service.
- Product team: Effectively delivered one full quarter of roadmap ahead of schedule.
- Another product team: 6-week schedule pull-forward on a major feature, at reduced capacity.
- Yet another product team: A backlog item that had no path to delivery got shipped because AI created spare capacity.
These weren’t just shortcuts on existing work. They were bets no one would have made when each one cost weeks of effort.
Once we started looking for this pattern, we saw it happening more and more. Not because we gave people more time, but because each project went from a multi-week investment to an afternoon experiment. Discovery phases started producing working artifacts instead of just documents: working API scaffolding, configuration management tooling, reusable implementation foundations. The validation process itself became the starting point for building.
The Bigger Opportunity: Product and Growth Experimentation
The engineering examples are compelling, but they’re not the most interesting part of this shift.
The real unlock is what happens when product and growth teams start thinking in terms of experimentation velocity. When the question changes from “can we get this on the roadmap?” to “can we just try it?”
One unexpected outcome from our sprints was a cross-team POC that shipped a genuinely new product interaction model — a way for users to interact with AI directly within the screens they’re already using, without switching to a separate chat interface. Context-aware, embedded in the existing product, surfacing alongside whatever the user is already doing. The team built it in a single sprint using the existing backend architecture. It wasn’t on any roadmap. A few people had the idea, the cost of trying it was low enough to just go, and within days they had a working demo on two platforms. Within a week, that demo had reshaped how the product team was thinking about AI surface area and they’re now using that experiment as the foundation for new product plans. What started as an unplanned two-day bet is actively shaping the roadmap.
That sequence (have an idea, validate it quickly, commit if it works) is what high experimentation velocity actually looks like. It’s the sequence that matters most for product and growth teams, because it changes the economics of exploration.
Think about what product teams spend most of their time on: building cases for what to try next. Writing specs. Prioritizing backlogs. Negotiating for engineering capacity. Most of that overhead exists because trying things is expensive, so you need to be very sure before you commit resources. When trying things becomes cheap, the entire prioritization calculus changes. You don’t need a business case to place a small bet. You need a hypothesis and a day.
This is the difference between a product team that ships its roadmap 20% faster and one that tests three ideas for every one it used to test. The first team gets efficiency. The second team gets compounding discovery and some of those discoveries become the roadmap.
The New Bottleneck
Experimentation velocity is only valuable if your organization can act on what experiments reveal.
Once engineering capacity stops being the constraint, a new set of bottlenecks becomes visible and they don’t live in engineering.
Requirements that aren’t ready to act on. Design work that can’t keep pace with implementation. Product decisions that take longer to make than the feature takes to build. These are now the rate limiters. The constraint has moved from “can we build this?” to “do we know what to build, and can we decide fast enough to keep up with what we’re capable of executing?”
The teams pulling the most value from agentic workflows are the ones that treated this as a whole-system change, not just an engineering upgrade. They restructured how requirements get written: specific, testable, agent-ready. They compressed design cycles to match implementation speed. They pushed decision-making closer to the work, with explicit permission to move at 70% confidence and course-correct on real output rather than waiting for perfect plans.
What To Ask Your Teams
If you’re evaluating your AI investment primarily through a delivery lens (features shipped, velocity improved, cycle time reduced) you’re getting value, but you’re probably underestimating what’s available.
The better question is a pair: how much faster are we executing what we planned, and what did we try this quarter that we couldn’t have tried six months ago?
If the answer to the second question is a long list of experiments, prototypes, and bets that came out of nowhere, you’re seeing the real return. If the answer is just “we shipped the roadmap faster,” you’re leaving the more interesting half on the table.
Delivery speed is a good start. Experimentation velocity is the level up. The constraint has moved. The teams winning with AI aren’t just faster, they’re attempting more. That’s the metric worth tracking.