Engineering report: how much faster is AI, measured

Every agency claims a multiplier. Almost none publishes what they measured, against which baseline, on how many projects. This page publishes the method first and the numbers second.

Engineering report

How much faster is AI actually? Our own numbers, with the method

Every agency claims a multiplier. Almost none publishes what they measured, against which baseline, on how many projects.

The claim, in full

AI-accelerated, senior-reviewed. On the code-writing share we are typically 2–4× faster than our own pre-AI baseline. Typical full-project savings: 30–50%.

That sentence is the only performance claim we make anywhere on this site. If you ever see a bigger number from us without a method next to it, treat it as a mistake and tell us.

Cadence

Quarterly

Baseline

Our own pre-AI hours

Sample

Every project, not the flattering ones

Method

How we measure

Six rules, fixed in advance, so a quarter with bad numbers still gets published.

Unit of measure. Engineer-hours per delivered feature of comparable complexity, grouped by task class.

Baseline. Our own tracked hours from projects delivered before the pipeline changed. Not an industry average, not a competitor.

Sample. Every project in the period, including the ones that went badly. We state N per task class.

Exclusions. Research spikes, projects whose scope changed mid-flight, and any task class with fewer than three comparable data points.

Who measures. The delivery lead, from the tracker, not from memory.

Reporting. Multipliers to one decimal; ranges rather than point estimates wherever N is under ten.

Results

By task class

The first edition is in preparation: measurements are being collected on current projects and will be published with N per class.

Task class
StageAI in the loop
Where AI helpsExpected effect
Boilerplate, CRUD, admin screens
Heavily
The biggest measurable win
Data model and migrations
Partly
Drafted, then reviewed line by line
Business logic
Partly
Drafted with AI, rewritten by the owner
Tests
Heavily
Generated coverage, human-reviewed assertions
Third-party integrations
Rarely
Their docs lie; the debugging is human
Requirements and scoping
No
Trade-offs are judgement calls
UI and UX decisions
No
Generated layouts look fine and convert badly
Code review
No
A human is accountable for what merges
Limitations

What did not get faster

If we ever publish a number without this section, stop trusting the number.

Discovery, requirements and design decisions did not compress at all.

Integrations against poorly documented third-party APIs sometimes took longer, because generated code confidently invents endpoints that do not exist.

Client decision latency dominates calendar time on roughly a third of projects and is entirely outside our control — no tooling changes that.

Task classes with small samples are reported as ranges, and we say so rather than rounding a thin sample into a headline.

What we will never publish

A multiplier without N and a period · a comparison against an unnamed "industry average" · a number produced by the marketing side of the house · savings claims on total project cost beyond what the measurement supports.

Industries

Software and AI for every kind of business

From fintech and e-commerce to manufacturing and property development — the stack follows the domain, not an off-the-shelf product that fits none of them.

18 sectors

  • IT companies and digital agencies
  • Insurers and brokers
  • Lenders and microfinance
  • E-commerce
  • Wholesale and distribution
  • Manufacturing
  • Freight and forwarding
  • Clinics and medical centres
  • Dental
  • Car dealers and service centres
  • EdTech
  • Legal and consulting firms
  • Equipment rental
  • Real estate agencies
  • Property developers
  • Travel and events
  • Beauty and wellness
  • Construction and renovation
FAQ

Questions about the measurement

Ask the same questions of anyone else quoting an AI multiplier.

Why measure against your own baseline rather than the industry?

Because industry averages are unverifiable and self-serving. Our own historical hours are the only comparison we can actually evidence.

Does the discount apply to my whole invoice?

No. It applies to the code-writing share, which is roughly half a project. That is why we quote 30–50% on a full project rather than 2–4×.

Can we see the raw data?

We publish aggregates plus the method. Client-identifiable data stays private; if you need the methodology in more depth for procurement, ask and we will send it.

What if a quarter shows no improvement?

We publish it. A measurement programme that only reports good quarters is advertising with a spreadsheet attached.

How does this affect the price I pay?

Directly: our pricing takes the market average for the same scope and divides the programming share by 2.5. See how we price.

2–4×
faster on the code-writing share
in 3 weeks
from idea to MVP
230+
projects delivered
17 years
of engineering experience
Work

Work we can show you

Hold us to this

If a number on our site is not backed by the method on this page, tell us and we will fix or remove it.

The first full edition publishes once the current measurement window closes.