Measure Fluency, Not Maturity

Published on
August 25, 2026

This is the final post in our five-part series on scaling AI-assisted engineering. The first four were moves you make: structure the repos, build the knowledge layer, run bounded sessions, install shared rituals. This one is how you know any of it is working, and it's the piece that turns the whole program into something you can take to a board.

Maturity models measure you against a stranger

The standard way to assess an engineering organization's AI adoption is a maturity model: a ladder of levels, and an exercise in placing yourself on a rung. The problem is that the ladder was written without your codebase, your domain complexity, or how your teams are distributed. It tells you what good looks like in general. It has nothing to say about what good looks like for you.

DORA's research points at the better anchor. It defines elite performance as what your best teams are already doing. That's far more credible than a framework's idea of good, because your best teams are operating inside your exact constraints and still finding the balance. The question worth answering isn't "what level are we." It's "what are our most effective AI practitioners doing differently, and how do we spread it."

Don't measure against a framework. Measure against your best team.

What fluency actually measures

Fluency is a measurement across teams on three dimensions at once, because any one of them alone lies to you.

Code quality, tracked through signals like code drift and rework rate. Delivery stability, using the DORA metrics as the baseline: deployment frequency, lead time, change failure rate, and time to recovery. DORA's metrics are instruments, not a ladder. A maturity model tells you where you rank; DORA tells you what's actually happening, which is why the metrics survive here after the model didn't. And adoption patterns, read from token spend, session length, and model distribution, the same signals the earlier posts in this series were quietly producing.

Held together, those three give you a team-level view that connects AI investment to delivery outcomes. You find the teams achieving the best balance of all three, you study what they do, and you propagate it deliberately instead of waiting for it to diffuse on its own. That last word matters. Left alone, good practice diffuses slowly and unevenly. Measured and propagated, it moves on purpose.

Why measuring against one dimension misleads you

Any single number can be gamed, usually by accident. Adoption dashboards are the common trap: ninety percent of engineers using the tools tells you nothing about what the tools produced. Token spend on its own is just as misleading in the other direction. Cost going down looks like a win until you see it happened because a team started cutting corners.

This is why the three dimensions have to move together, and it's exactly the reading a CFO needs. Token costs falling while delivery stability holds or improves is the ROI story: you're spending less and shipping just as reliably or better. Token costs falling while stability drops is a warning, not a win. It means teams are trading quality for cost, and you want to catch that before it reaches production, not after. One number can't tell those two situations apart. The balanced view can.

Fluency is the thread through the whole series

Look back at the arc. Structured repositories made teams comparable, which is the precondition for measuring anything. The knowledge layer, bounded sessions, and shared rituals each threw off exactly the signals fluency reads: quality, stability, adoption. Measurement wasn't the last step bolted on at the end. It ran alongside everything from the start, and only now does it have a clean environment and real signals to work with.

That's the sequence, and the order is not arbitrary. Try to measure before the repos are structured and your comparisons are noise. Try to scale rituals before you can measure and you're propagating on faith. Each move makes the next one legible. Fluency is where they add up to a number you can defend.

The result you're actually after

The goal was never a maturity score. It was an engineering organization where AI-assisted work is disciplined, where the best teams' practices spread on purpose, and where you can stand in front of the people who funded the bet and show that spend went down while delivery held. That's what fluency measures. A maturity model can't give it to you, because it's measuring against someone else's idea of good instead of your own best work.

If your only evidence that AI is paying off is an adoption dashboard, you don't have proof, you have activity. Building the fluency measurement that connects AI investment to delivery outcomes is where a Team AI Upskilling engagement lands, and it's the payoff for everything in this series. Let's talk.

Frequently asked questions

What is AI fluency measurement? It's a team-level measurement of AI-assisted engineering across three dimensions at once: code quality, delivery stability, and adoption patterns. Rather than ranking you on a generic ladder, it identifies which of your teams achieve the best balance and lets you propagate what they do.

How is fluency different from a maturity model? A maturity model measures you against a generic framework written without your codebase, domain, or team structure. Fluency measures against your own best teams, which are already operating inside your real constraints, and spreads their practices deliberately instead of placing you on a rung.

What metrics does fluency use? Delivery stability from the DORA metrics (deployment frequency, lead time, change failure rate, time to recovery), quality signals like code drift and rework rate, and adoption signals like token spend, session length, and model distribution. The three are read together, never in isolation.

Why isn't an adoption dashboard enough? An adoption number tells you the tools are being used, not what they produced. Measured alone, it hides whether quality or stability moved. Fluency reads adoption alongside quality and stability so you can tell a genuine ROI win from a team quietly cutting corners.

This is Post 5 of Stride's five-part series, AI-Assisted Engineering at Scale. Start from the beginning: why structuring your repositories is the first move.

Share