The AI exists. The models exist. What does not yet exist — in most engagements — is a measurement protocol rigorous enough to produce a conclusion a sceptical CFO will fund. That is what we built.
A pilot is typically run on a curated set of issues, with the AI model having access to more context than it would in production, and evaluated by the team that ran it. The result is an impressive number that does not survive the first follow-up question.
We designed the assessment around the opposite assumption: that every methodological choice will be interrogated by someone who wants the answer to be no. The blind protocol, the sourcing discipline, the independent audit option — each one exists to answer a specific objection before it is raised.
The AI model is given the ticket and the pre-fix codebase. Its answer is frozen and timestamped. Only then does it see the fix your engineers actually shipped. The gap between what it predicted and what happened is the measurement. No hindsight. No contamination.
Every figure in the financial model carries a label: client data, Aevon measurement, or industry benchmark. There are no unlabelled lines. A CFO who cannot trace a number back to its origin cannot defend it. We make that trace explicit before they ask.
Before measurement begins, we agree the issue taxonomy with your team. What counts as a maintenance issue, and what does not, is defined by you — not inferred by us. This is how we prevent the result being shaped by how we classified the inputs.
We invite your engineers to independently score a random sample of our findings — before the final report, not after. Their scores are included in the output. If their assessment differs from ours, that difference is reported, not smoothed over.
Map the maintenance workload. Agree the issue taxonomy. Define what AI-addressable means for this codebase. Capture the baselines — engineer hours by issue class — that the financial model will rest on. Output: agreed taxonomy document and baseline dataset.
Full output model built on the first four weeks of data, timed to your board calendar. You see the shape of the final answer 50 days before it lands — including an explicit statement of what we cannot yet conclude. No surprises at the final milestone.
Every issue in the dataset is scored under the blind protocol. Weekly findings reviews with your team. Independent scoring session with your engineers. Running dataset — you can see the numbers build in real time.
Evidence base, financial model, and operating blueprint delivered as a complete package. Structured to be presented by your team without us in the room — because that is when it will be tested hardest.
The method is public. The engagement is the application of it to your specific codebase and history.
Talk to someone who runs engagements →