The Method

Measurement discipline is the product.

The AI exists. The models exist. What does not yet exist — in most engagements — is a measurement protocol rigorous enough to produce a conclusion a sceptical CFO will fund. That is what we built.

Why the method matters

Most AI pilots are not designed to produce a decision.

A pilot is typically run on a curated set of issues, with the AI model having access to more context than it would in production, and evaluated by the team that ran it. The result is an impressive number that does not survive the first follow-up question.

We designed the assessment around the opposite assumption: that every methodological choice will be interrogated by someone who wants the answer to be no. The blind protocol, the sourcing discipline, the independent audit option — each one exists to answer a specific objection before it is raised.

The four protocols

Each one closes a specific gap.

Protocol 01

The blind protocol

The AI model is given the ticket and the pre-fix codebase. Its answer is frozen and timestamped. Only then does it see the fix your engineers actually shipped. The gap between what it predicted and what happened is the measurement. No hindsight. No contamination.

Protocol 02

Source labelling

Every figure in the financial model carries a label: client data, Aevon measurement, or industry benchmark. There are no unlabelled lines. A CFO who cannot trace a number back to its origin cannot defend it. We make that trace explicit before they ask.

Protocol 03

Issue taxonomy agreement

Before measurement begins, we agree the issue taxonomy with your team. What counts as a maintenance issue, and what does not, is defined by you — not inferred by us. This is how we prevent the result being shaped by how we classified the inputs.

Protocol 04

Independent scoring

We invite your engineers to independently score a random sample of our findings — before the final report, not after. Their scores are included in the output. If their assessment differs from ours, that difference is reported, not smoothed over.

The 90-day timeline

Every checkpoint has a defined output.

Days 1–30
Discovery

Baseline and taxonomy

Map the maintenance workload. Agree the issue taxonomy. Define what AI-addressable means for this codebase. Capture the baselines — engineer hours by issue class — that the financial model will rest on. Output: agreed taxonomy document and baseline dataset.

Day 40
Interim

Early model preview

Full output model built on the first four weeks of data, timed to your board calendar. You see the shape of the final answer 50 days before it lands — including an explicit statement of what we cannot yet conclude. No surprises at the final milestone.

Days 31–60
Evidence

Full blind protocol measurement

Every issue in the dataset is scored under the blind protocol. Weekly findings reviews with your team. Independent scoring session with your engineers. Running dataset — you can see the numbers build in real time.

Days 61–90
Decision

Final package

Evidence base, financial model, and operating blueprint delivered as a complete package. Structured to be presented by your team without us in the room — because that is when it will be tested hardest.

The method is public. The engagement is the application of it to your specific codebase and history.

Talk to someone who runs engagements →