Azure and Cloud

Where Avoidable Cloud Spend Hides

HL
Harry Le
4 min read
Most cloud estates we review carry 20 to 30 per cent of avoidable spend, and it hides in the same places every time. A field guide, in descending order of embarrassment.

Most cloud estates we review carry 20 to 30 per cent of avoidable spend. That number stays consistent across industries and cloud vendors because the causes are structural: cloud makes spending easy and attributing spend hard, and the gap between those two compounds monthly. The waste hides in the same places every time, so it is worth naming them plainly.

The environments that never sleep

Non-production is the classic. Development and test environments sized like production, running around the clock, for teams that work business hours in one or two time zones. Scheduling non-production to switch off nights and weekends removes roughly two thirds of its compute hours, and the objection, "someone might need it at 2am", is answered with a self-service start button rather than a standing bill.

The grown-up version of this problem is the abandoned environment: spun up for a project, never torn down, orphaned when the project ended. Estates without expiry policies accumulate these like sediment, and each one costs real money to do nothing.

Paying on-demand prices for predictable workloads

Reservations and savings plans exist because vendors will discount heavily, often 40 to 60 per cent, for commitment. Estates that run steady baseline workloads on on-demand pricing are donating that discount back. The reason is rarely ignorance; it is that nobody owns the decision, because commitment feels like risk and the bill belongs to everyone and no one. A quarterly review of steady-state usage against commitment coverage is an hour of work with one of the best hourly rates in the industry.

Storage that nobody has tiered

Storage grows monotonically, because deleting requires a decision and keeping requires nothing. The result is hot-tier storage full of cold data: logs from 2022, backups of decommissioned systems, snapshots taken before changes nobody remembers. Lifecycle policies that age data through cool and archive tiers are boring, standard and frequently unconfigured. On large estates, storage tiering alone has paid for the entire cost review.

Oversized by inheritance

Instances get sized at migration time, generously, because nobody wants the go-live incident, and then never revisited. Utilisation data makes right-sizing a mechanical exercise: sustained single-digit CPU on a large instance is a resize, not a mystery. The same inheritance problem applies to premium disk tiers, redundant IP addresses, and load balancers fronting things that no longer exist.

Why the bill stays illegible

Underneath all of these is attribution. When resources are not tagged to workloads and owners, the invoice is one number that belongs to everyone, and no line of it is anyone's problem. Tagging discipline, enforced by policy at the platform level rather than by reminder emails, is what converts a bill into a set of accountable numbers. This is why cost work belongs to engineering rather than procurement: every lever above is an architectural or platform decision.

The monthly ritual that keeps an estate honest reports dollars per workload against last month, with a named owner per line. Estates reviewed that way stay lean. Estates reviewed as a single total drift back to the 20 to 30 per cent within a couple of years, and the cycle repeats.

Cost management is part of the operations line of our Azure and cloud engineering practice, alongside the monitoring and platform work it depends on.

Harry Le is a cloud engineer at Coder Trove.