FinTechFinOpsResilience

Cut cloud spend by $240K—without cutting reliability.

The brief was not “make AWS cheaper.” It was to make cost an observable property of every transaction while preserving banking-grade availability and auditable controls.

$240KAnnual savings
99.99%Availability
31%Lower unit cost
0Unplanned outages
01 · Constraint

The bill was visible. The reason was not.

Spend reports arrived by account and service, but product owners needed cost per payment, per merchant and per risk check. Idle capacity protected peak events, so indiscriminate right-sizing would have damaged the reliability objective.

No unit economics

Cloud cost could not be tied to product traffic or ownership.

Peak-shaped capacity

Static fleets were sized for rare settlement windows.

Unowned waste

Snapshots, IPs and test clusters survived because deletion had no accountable workflow.

02 · Control plane

Cost signals joined the same feedback loop as latency and errors.

AttributeOwnership graph

Mandatory service, team and product tags

MeasureCost per event

CUR joined with transaction telemetry

ForecastWorkload envelopes

Settlement-aware demand profiles

ActPolicy automation

Rightsize, schedule and tier safely

ProtectSLO guard

Freeze changes when error budget burns

Failure path: scaling policies were bounded by tested minimum capacity; any rising error budget suspended cost actions and returned control to reliability automation.
03 · Decisions

Savings came from behavior, not one cleanup sprint.

D-01Optimize unit cost

Teams saw rupees per transaction next to p95 latency, making architectural trade-offs explicit.

D-02Policy with expiry

Every exception to resource policy required an owner, reason and automatic expiration.

D-03Reliability veto

FinOps automation never overruled an active SLO protection state.

04 · Outcome

A leaner estate with a safer operating model.

MeasureBeforeAfter
Annualized cloud run-rateBaseline$240K lower
Cost attributionAccount levelProduct transaction
Idle non-prod computeAlways onSchedule + scale to zero
Availability99.91%99.99%
AWS CURTerraformOPAKarpenterGrafanaOpenTelemetryS3 Intelligent-Tiering
Operating principle: a recommendation was not counted as savings until it survived a full settlement cycle and appeared in the finance baseline.

Make cost an engineering signal.

We can trace spend to workloads and identify the changes that will not compromise your SLOs.

Audit my cloud economics →