The cost of failure never lands on the model that failed. Retries inflate call volume. Escalations show under the premium model's line. Rework shows up in salaries.
Built a calculator.
The cost of failure never lands on the model that failed. Retries inflate call volume. Escalations show under the premium model's line. Rework shows up in salaries.
Built a calculator.
Not auditor access. Not admin.
They were given “tenant owner” privileges in Azure — full control over the NLRB’s cloud, above the CIO himself.
This is never supposed to happen.
Not auditor access. Not admin.
They were given “tenant owner” privileges in Azure — full control over the NLRB’s cloud, above the CIO himself.
This is never supposed to happen.
Almost none of that is architecture.
Here is the AWS waste sweep, in the order I run it.
Almost none of that is architecture.
Here is the AWS waste sweep, in the order I run it.
Cost Explorer, Budgets, Cost Anomaly Detection, Compute Optimizer, Cost Optimization Hub. All native, all free, genuinely capable.
We put our own tool on the list too, and said where it does not belong.
Cost Explorer, Budgets, Cost Anomaly Detection, Compute Optimizer, Cost Optimization Hub. All native, all free, genuinely capable.
We put our own tool on the list too, and said where it does not belong.
No instance. No resource to delete. The unit of cost is the token, and the bill is driven by HOW you call the model - not what you provisioned.
Per-model token visibility comes first.
No instance. No resource to delete. The unit of cost is the token, and the bill is driven by HOW you call the model - not what you provisioned.
Per-model token visibility comes first.
The same spike found on next month's invoice costs you thirty.
AWS Cost Anomaly Detection is free, ML-based, and learns your pattern per service. Five minutes to set up.
Most teams still have not switched it on.
The same spike found on next month's invoice costs you thirty.
AWS Cost Anomaly Detection is free, ML-based, and learns your pattern per service. Five minutes to set up.
Most teams still have not switched it on.
Compute Savings Plan ~66% - flexes across family, region, even Graviton
EC2 Instance SP ~72% - locked to family + region
Standard RI ~72% - no family change, but sellable
Convertible RI ~54% - exchangeable
Depth costs flexibility.
Compute Savings Plan ~66% - flexes across family, region, even Graviton
EC2 Instance SP ~72% - locked to family + region
Standard RI ~72% - no family change, but sellable
Convertible RI ~54% - exchangeable
Depth costs flexibility.
Which quietly turned "we have a few spare Elastic IPs" into a line item.
Releasing idle ones is the cheapest cleanup in AWS. No performance risk, no sign-off, no downside.
Which quietly turned "we have a few spare Elastic IPs" into a line item.
Releasing idle ones is the cheapest cleanup in AWS. No performance risk, no sign-off, no downside.
As CSV the headers are declared once. Models read delimited data fine.
Same answer, a fraction of the input tokens.
Token count is a property of your code, not the model.
As CSV the headers are declared once. Models read delimited data fine.
Same answer, a fraction of the input tokens.
Token count is a property of your code, not the model.
Models are stateless, so the orchestrator re-sends the whole history every step. Step 10 pays for steps 1-9 again.
Cost grows with the SQUARE of the step count.
max_tokens does not bound a loop.
Models are stateless, so the orchestrator re-sends the whole history every step. Step 10 pays for steps 1-9 again.
Cost grows with the SQUARE of the step count.
max_tokens does not bound a loop.
Someone ran a POC. The project moved on. The deployment did not.
No quality trade-off, no code change, no A/B test. Just delete it.
Cleanest money on any AI bill.
Someone ran a POC. The project moved on. The deployment did not.
No quality trade-off, no code change, no A/B test. Just delete it.
Cleanest money on any AI bill.
Someone ran a POC. The project moved on. The deployment did not.
No quality trade-off, no code change, no A/B test. Just delete it.
Cleanest money on any AI bill.
Someone ran a POC. The project moved on. The deployment did not.
No quality trade-off, no code change, no A/B test. Just delete it.
Cleanest money on any AI bill.
It drops to "available", attached to nothing, billed in full - EBS bills provisioned capacity, not usage.
A forgotten 500GB io2 with 5k IOPS is about $4,700 a year.
It drops to "available", attached to nothing, billed in full - EBS bills provisioned capacity, not usage.
A forgotten 500GB io2 with 5k IOPS is about $4,700 a year.
FOCUS ends that. One billing schema, and every major cloud now emits it natively - Microsoft, AWS, Google, Oracle.
One query, any provider.
FOCUS ends that. One billing schema, and every major cloud now emits it natively - Microsoft, AWS, Google, Oracle.
One query, any provider.
The most recent 1-2 days of cost data are ESTIMATES, and they often revise down.
A good share of "our bill exploded" turns out to be unfinalized days.
Rule out the false alarm before you start the hunt.
The most recent 1-2 days of cost data are ESTIMATES, and they often revise down.
A good share of "our bill exploded" turns out to be unfinalized days.
Rule out the false alarm before you start the hunt.
You do not pay for pods. You pay for nodes.
The scheduler packs by what pods REQUEST, not what they use - so most clusters run at 30-40% real CPU while paying for 100% of the nodes.
That gap is the bill.
You do not pay for pods. You pay for nodes.
The scheduler packs by what pods REQUEST, not what they use - so most clusters run at 30-40% real CPU while paying for 100% of the nodes.
That gap is the bill.
It does not tell you what you WASTED - and that is not a flaw, it is not its job.
It will never surface an orphaned disk, an idle gateway or missing Hybrid Benefit.
Budgets and alerts there on day one. Then go looking.
It does not tell you what you WASTED - and that is not a flaw, it is not its job.
It will never surface an orphaned disk, an idle gateway or missing Hybrid Benefit.
Budgets and alerts there on day one. Then go looking.
Stopping a VM stops the compute bill. It does not stop the disk bill.
Every disk on every deallocated VM is still billing you at full rate, for capacity nobody has read since.
Deallocated is not deleted.
Stopping a VM stops the compute bill. It does not stop the disk bill.
Every disk on every deallocated VM is still billing you at full rate, for capacity nobody has read since.
Deallocated is not deleted.
A cheaper model that fails more often can cost more overall.
Self-hosting shifts cost from marginal to fixed.
A stricter data policy may rule out the best-performing provider.
A cheaper model that fails more often can cost more overall.
Self-hosting shifts cost from marginal to fixed.
A stricter data policy may rule out the best-performing provider.
Nobody argues about whether the bill went down.
Security and reliability wins are real and hard to prove. Cost wins buy you the credibility to go after them.
Nobody argues about whether the bill went down.
Security and reliability wins are real and hard to prove. Cost wins buy you the credibility to go after them.
FinOps owns one pillar in depth - cost, and the business value of spend.
A CCoE owns the operating model, and is where cost trades off against reliability, security and speed.
Small org: same people. Large org: FinOps sits inside it.
FinOps owns one pillar in depth - cost, and the business value of spend.
A CCoE owns the operating model, and is where cost trades off against reliability, security and speed.
Small org: same people. Large org: FinOps sits inside it.
It's whether scaling something DOWN is routine rather than exceptional.
Most orgs have a well-worn path for growing a resource and no path at all for shrinking one.
It's whether scaling something DOWN is routine rather than exceptional.
Most orgs have a well-worn path for growing a resource and no path at all for shrinking one.
Performance asks: are resources matched to demand?
Cost asks: are we paying the best rate for what we use?
A workload can be perfectly sized and still overpriced, because nobody bought a commitment.
Performance asks: are resources matched to demand?
Cost asks: are we paying the best rate for what we use?
A workload can be perfectly sized and still overpriced, because nobody bought a commitment.
All five will say tier one.
That's why workload tiering can't be self-assessed. Without someone owning the classification, redundancy gets applied uniformly and expensively, or inconsistently and dangerously.
Usually both.
All five will say tier one.
That's why workload tiering can't be self-assessed. Without someone owning the classification, redundancy gets applied uniformly and expensively, or inconsistently and dangerously.
Usually both.
Publish standards without paving the road -> documents nobody reads.
Become an approval gate -> a queue teams route around.
The working version makes the compliant path the easiest path. A landing zone nobody uses is a project, not an operating model.
Publish standards without paving the road -> documents nobody reads.
Become an approval gate -> a queue teams route around.
The working version makes the compliant path the easiest path. A landing zone nobody uses is a project, not an operating model.