Article · 7 September 2026 · 5 min read

Nobody was overspending on purpose

Two platforms, two bills, no villain in either story. One had bought a Kafka cluster sized for a company it had not become yet. The other had copy-pasted a resource block until it covered three clusters. Neither shows up in a standup, and neither has a deadline.

The bill nobody re-read

Production Kafka on a managed platform: $58,000 a year, running at 7% cluster load. Priced at list against measured usage, the same workload landed near $9,200 a year on the cloud the company was already paying for. That is about $49,000 a year, from one line item, found by reading an invoice next to a metrics dashboard.

The method is dull and the dullness is the point. Measure first. Price every option at list against the profile you actually measured, using each vendor's own model rather than the one you are used to. Then name the single dependency that gates the move, because cheap is irrelevant if you cannot leave. In that case the blocker was never Kafka. It was two stream-processing pools tied to the incumbent's ecosystem.

The one that was not on any bill

The second finding was harder to see, because nothing was labelled with it. On GKE Autopilot the reservation is the invoice: you are billed for the CPU and memory your pods request, not for what they use. Nearly every service across three clusters carried the same resource block. It had been roughly right once, for one service, and then it was copied.

Thirty days of utilization dashboards, built from container metrics already present in the platform, and then sizing per workload against observed steady state. 31 vCPU of reservations came back, about a third off the Kubernetes line, with no application change, lowering requests where the measurement supported it and raising the two that were throttling. Put your own vCPU-hour rate against 31 vCPU held every hour of the year. The number is not small.

The measurement said some workloads needed more, not less. A right-sizing exercise that only ever goes down is not a measurement, it is a target.

The arithmetic, done honestly

A capable DevOps engineer in Europe is comfortably a six-figure package once salary, charges, tooling and the recruiter's cut are counted. One line item running at 7% of what it was sized for is not that whole package. It is a large enough slice of it that the second and third findings finish the job, and it recurs every year nobody looks.

So the honest claim is not "an audit pays for a hire." It is narrower and it holds up better: findings of this size are common enough that a few of them clear a year of headcount, and unlike a salary they compound in your favour. That is the sentence that survives a CFO reading it twice.

Why the team had not found it

Both findings were sitting in data the companies already had. Neither was found, for the same reason: it was nobody's job on a Tuesday. Cost work never has a deadline, so it loses every week to work that does. Product engineers can do it, at their salary, slowly, in the week they are not shipping. Mostly they do not, and they are right not to.

When there is nothing to find

Sometimes there is not. Some platforms are already sized correctly, and the honest output of a week spent looking is "your bill is about the right size, and here are the two things that will change that as you grow." That is a real result and it gets written down as one. Anyone who always finds a large saving is not finding, they are selling.

Which is why the entry point is scoped as a week rather than as a promise: five days, read-only, a risk register ranked by impact with the expected consequence and the effort to fix, and a 90-day plan that mostly routes to your own team. 4,900 EUR, credited against your first month if it turns into a retainer. If the week turns up one Kafka, the first year returns it more than ten times over. If it turns up nothing, you have that in writing, which is also worth having.

Figures are shifted from the engagements they come from so that no client is identifiable; the ratios, the pricing models and the method are unchanged. The Kafka figures are dollars and the fees are euros, so the comparisons here are orders of magnitude rather than a converted total.

Find out which line item it is

The platform audit is five days, read-only, 4,900 EUR: a risk register ranked by impact with the effort to fix, and a 90-day plan you can mostly run yourselves. Tell us what you run and we will tell you where to look first.

Get in touch