The bill was visible. The reason was not.
Spend reports arrived by account and service, but product owners needed cost per payment, per merchant and per risk check. Idle capacity protected peak events, so indiscriminate right-sizing would have damaged the reliability objective.
Cloud cost could not be tied to product traffic or ownership.
Static fleets were sized for rare settlement windows.
Snapshots, IPs and test clusters survived because deletion had no accountable workflow.
Cost signals joined the same feedback loop as latency and errors.
Mandatory service, team and product tags
CUR joined with transaction telemetry
Settlement-aware demand profiles
Rightsize, schedule and tier safely
Freeze changes when error budget burns
Savings came from behavior, not one cleanup sprint.
Teams saw rupees per transaction next to p95 latency, making architectural trade-offs explicit.
Every exception to resource policy required an owner, reason and automatic expiration.
FinOps automation never overruled an active SLO protection state.
A leaner estate with a safer operating model.
| Measure | Before | After |
|---|---|---|
| Annualized cloud run-rate | Baseline | $240K lower |
| Cost attribution | Account level | Product transaction |
| Idle non-prod compute | Always on | Schedule + scale to zero |
| Availability | 99.91% | 99.99% |
Make cost an engineering signal.
We can trace spend to workloads and identify the changes that will not compromise your SLOs.