Glowing network of connected nodes marked with code, branch, and checkmark icons, radiating from a bright central point over a dark circuit-board background

We Fixed Quality and Accidentally Got AI-Ready

Two earlier attempts to measure quality had failed. The third turned out to be the groundwork for AI-assisted development.

Key Takeaways
  • Team-owned tests, definitions of done, written context, and measurable delivery practices all matter more once AI starts writing code.
  • The program worked because teams scored themselves and backed every claim with evidence a reviewer could inspect.
  • The same evidence that helped humans review quality also created structured context an AI reviewer could use.
  • Early AI reviewer mistakes exposed team constraints that had never been written down anywhere.
  • AI adoption still needs honest measurement: the client's own data shows adoption climbing while efficiency gains aren't yet demonstrable.

The most useful AI-readiness work we did for one client started as a quality scorecard.

That sounds less exciting than a coding agent, and honestly it felt that way at first too. I expected the hard part to be getting the rubric right: enough detail to be useful, not so much that nobody wanted to touch it. The client had a real quality problem across a large engineering organization, and they needed a way to understand where teams stood. Later, when they started asking harder questions about AI-assisted development, I had to revise my own view of what we had been building. The evidence trail, boring as it was, turned out to be the useful part.

The environment

The client is a large retailer with a digital commerce platform built by a distributed engineering organization. Dozens of teams own React microfrontends and Java services, ship through GitLab CI/CD, and run on Kubernetes.

There was no single codebase, no single release train, and no one team that could just decide how quality worked for everyone else. Each team had its own service boundaries, pipeline habits, test suite (or lack thereof), staffing reality, and set of constraints. Some teams owned complete applications. Some contributed to other teams' repositories. Some worked inside vendor products that limited what they could automate.

Any standard that was supposed to matter across the organization had to survive contact with all of that, including teams that did not fit the clean version of the org chart.

The first two attempts

The client was candid about the history. They had tried to measure quality twice before, and neither attempt had stuck. We were the third consulting team to try.

They even wrote the lesson into the original scorecard. Paraphrasing only slightly:

We've done this before and we don't think we hit the mark. We want to do better, and we want to define our own future.

The first attempt was interview-based. A consulting firm talked to teams and asked, in effect, how good they were at testing and quality. Most teams thought they were doing reasonably well. The output was subjective and, in the client's own words, "not super helpful."

The second attempt went the other direction. A different consulting firm rated teams from the outside. The numbers looked more objective, but the process felt like finger-pointing and created tension.

Meanwhile, quality was owned by an external QA team, so defects showed up late, when they cost the most to fix. CI/CD configs had drifted team by team. There were no defined testing layers and no central quality strategy. And the microfrontend architecture that made modular development possible also made components a pain to test in isolation.

What changed

What became Organization Owned Quality (OOQ) kept the team ownership from the first attempt and the evidence from the second.

Teams had to provide evidence for each maturity level they claimed. They submitted artifacts: documented processes, work item examples, published pages, repeated patterns, repository evidence, pipeline evidence. "We usually do this" was explicitly weak evidence. A claim had to be backed by something a reviewer could inspect.

At the same time, teams scored themselves first. They brought the evidence forward. Reviewers verified it, challenged it where needed, and helped calibrate the standard. The adoption scale ran from Not Started to Full Adoption, and each action item included getting-started guidance. Teams were more willing to be honest about where they were when the same system gave them a next step.

The resulting scorecard had eleven pillars and 57 action items, covering documentation, test strategy, definition of done, team-owned automation, pipeline ownership, incident prevention, delivery metrics, and even hiring practice.

Some of the work was process, but a lot of it was plumbing. We built a reusable Playwright-based UI test SDK pattern so microfrontend teams could write and own their own tests. Docker packaging let the same automation run locally and in the pipeline. Standard GitLab CI/CD templates made builds, tests, and quality checks look more consistent from team to team.

The program's first year ran on a spreadsheet. It did what spreadsheets do in a program of this size: duplicated rows, stale tabs, manual cleanup, and a few cells nobody wanted to be responsible for breaking. The spreadsheet was also useful. It forced everyone to see the shape of the data before we hardened it into software.

We replaced it with a Rails and PostgreSQL application running on the client's Kubernetes platform. Today it tracks every team and a couple hundred users. It includes guidance on what belongs directly in the record versus what should stay linked from the source system, a structured place to capture team constraints, and notifications in the chat tool people already use.

The adoption dataset is exposed through a GraphQL API. Two of the original requirements were literally "make team details exportable for AI analysis" and "create endpoints to get the data programmatically or through AI." At the time, I treated that as a sensible extensibility requirement.

Two AI-assisted tools now sit on top of that foundation:

  • An OOQ Facilitator, a Microsoft 365 Copilot declarative agent that coaches teams through preparing a submission. It knows the pillar rubric and pulls evidence from the issue tracker and wiki using read-only access limited to each person's existing permissions. It grades evidence strength, assigns a confidence level, predicts the questions a reviewer is likely to ask, and hands back either a submission-ready write-up or a gap analysis with a plan to level up. It writes to nothing and doesn't score, so teams still own what they submit.
  • A Score Reviewer, an agent skill every reviewer uses. It's grounded in the pillar criteria and the client's own captured team context, and it applies the same criteria to every review, where reviewers used to rely on their own judgment across roughly 2,400 team-and-practice combinations. A human makes the final call.

The AI connection

About two years in, as the client started asking about AI-assisted development, we noticed that the things we wanted from teams for quality reasons were also the things an AI agent needed before it could safely change production code.

The quality practiceWhy it matters once AI is writing code
Team-owned automated test suitesAn agent can generate a lot of code very quickly. Without a suite the team owns and trusts, nothing tells you whether any of it works.
Test strategy defined before implementationGives the agent a written definition of "done and validated."
A published definition of done, enforced at mergeThis is the gate AI-generated code has to pass. If it's informal, "done" means something different to every reviewer and every agent.
Automated linting, warnings as errorsA linter checks every line the same way, whether a person or an agent wrote it.
Mutation testingCoverage percentage is exactly the metric AI-generated tests are best at gaming. Mutation testing asks whether your tests catch real behavior changes.
Team pipeline ownership, flaky-test preventionMore changes need a pipeline the team can triage itself, and a test suite nobody has learned to ignore.
Written process, ADRs, READMEsThe context an agent needs to ground itself.
Delivery metrics as an operating inputThe baseline you need to answer "did AI actually make us faster?" with data.

One of the very first items on the board, back in August 2024, was "document using Copilot with a test-first mindset." In May 2026 the team opened a work item to "look through all action items and figure out what needs to change to accommodate AI in the new world." They closed it that July, after revisiting the scorecard with AI-assisted development in mind.

Results

Since the measurement baseline was set, practices at Adopting or better across the organization grew from 41 to 820. Practices at Full Adoption went from 1 to 465. The organization now has 3,954 completed, evidence-backed practice reviews on the books.

In the most recent quarter, the program grew its footprint by about a third in new teams. The teams already on it improved 8.5%, and not a single team went backwards.

Definition of done is the most adopted practice, at 64% of team/practice combinations at Adopting or better, with test strategy consistency at 56%. They are also two of the controls AI-generated code has to pass before anyone should trust it. Most teams did not talk about them that way at first. They talked about fewer surprises in review, fewer arguments about what counted as done, and defects that stopped turning up after handoffs.

DORA metrics were one of the eleven pillars, and one of the least adopted. The organization stood up the instrumentation mid-program because the scorecard asked for it. The first readings:

MetricReadingDORA band
Change failure rate4.0%Elite
Mean lead time to change~2.2 daysHigh
Mean time to recover~65 hoursMedium

Change failure rate stayed between 4% and 7% in every month with meaningful deployment volume. Mean time to recover is still an obvious target for improvement.

Teams now write and maintain their own OOQ pages across at least nine separate wiki spaces: roadmaps, 90-day plans, squad-level test strategies, monitoring overviews, monthly operating reviews, repository inspection logs. When we started, no one could say how any given team ensured quality.

A single automated audit run processed 232 scores across 22 teams. It found adoption nobody had recorded and evidence that had gone stale.

What the Score Reviewer exposed

The first version of the Score Reviewer made confident, wrong calls.

It dinged one team for weak repository practices, when that team did not own any repositories at all; they only contributed to other teams' repos. It dinged another for missing linting, when that team worked inside a vendor product where they could not add a linter. It also assumed end-to-end testing happened in one environment when it actually happened in another.

None of those constraints were documented anywhere. They lived in the heads of people who had been around long enough to know the exception. In one review conversation, looking at a bad recommendation, the team asked why the correct answer required local folklore.

So we wrote the constraints down as structured, machine-readable context. From that point forward, a review, human or automated, could account for what a team could do. The team that had been penalized for undocumented constraints became the fastest-improving team in the organization the next quarter. We've written more about how context impacts AI agent output.

Reviewers now run the Score Reviewer against real submissions, partly to refine it and partly to find the next missing piece of context. I would not pretend every bad answer is secretly useful. Sometimes the model just misses. But enough of the misses have pointed at undocumented constraints that we started treating them as prompts for better context.

Measuring AI honestly

One of the client's teams also built a stage-by-stage framework to measure whether AI-assisted development is delivering. The framework is careful in ways I wish more AI ROI conversations were careful:

  • Every baseline is labeled Measured, Proxy, or Gut-feel, so nobody mistakes an estimate for a measurement.
  • They do not publish a median on fewer than ten samples.
  • Time, effort, and toil are reported separately, never added into one big flattering "savings" number.
  • They do not rank individual developers.

Their own findings say AI practice adoption is climbing fast, but efficiency gains are not demonstrable yet. For the broader picture, see what the data actually says about AI-assisted development.

Their framing line:

"Building a capability is not the same as getting value from it."

If you are about to roll out AI coding tools across a bunch of teams, write down what each team actually does and back it with evidence. Your agents will find the gaps quickly if you don't, the same way our Score Reviewer found the team with no repositories.

That's the kind of work we like at Leading EDJE: building quality culture and the technical plumbing that supports it, so AI-generated code has to clear the same bar as everyone else. If you need to increase quality ownership, standardize pipelines, shift testing left, or get honest about whether you're ready for the tools you're about to adopt, we'd love to talk about your quality program.

Frequently Asked Questions

Why did earlier quality measurement attempts fail?
What is Organization Owned Quality?
How does a quality program make an organization AI-ready?
What results did the program produce?
What did the AI Score Reviewer teach the team?