Cloud spending rebounds as AI workloads move to production

After two years of optimization, cloud budgets are growing again — and AI inference is the reason.

ThemeAnax 6 min read
A detailed view of a blue lit computer server rack in a data center showcasing technology and hardware.
Share

There is no shortage of advice about Industry insights. There is a shortage of advice that survives contact with a real week.

The expensive mistake

The compounding effects matter far more than the individual wins. Teams that pick both end up with neither, and usually discover this at the point where reversing would have mattered. Set a date at which you will stop, and write down in advance what would make you stop earlier.

graphical user interface
Photo by Deng Xiang on Unsplash

Measurement is usually where this falls apart. Choosing infrastructure before agreeing what it is for is how organisations end up maintaining a system nobody wanted. Try writing the constraint on one line before opening a vendor comparison; the line is usually harder than the comparison.

The advice worth ignoring

The expensive mistakes here are rarely the technical ones. A small improvement applied consistently beats a dramatic one applied once, which is unsatisfying advice precisely because it is correct. The evidence here is thinner than anyone quoting it tends to admit.

It helps to separate the decision from the execution. A team that changes approach every quarter pays a coordination tax that routinely exceeds whatever the change was meant to fix. Ask what would have to be true for the opposite approach to be correct, and see whether anyone can answer.

Most of the difficulty lives at the boundaries, not in the middle. Success has many causes and teaches very little; failure tends to have one, and it is usually obvious in hindsight. We ran both approaches in parallel for six weeks. The difference was smaller than the cost of the debate about it.

The one that only matters at scale

What looks like a process problem is frequently an ownership problem. Cutting scope early is cheap and slightly embarrassing; cutting it late is expensive and deeply embarrassing. The counter-argument deserves a hearing, and it is stronger than its usual proponents make it sound.

Feedback loops shorter than the planning cycle change everything. Handoffs between people who each hold a coherent local picture and no shared one produce most of the pain later attributed to tooling. That said, none of this generalises cleanly across team sizes.

What to do first

Nobody gets credit for the work that did not need doing. If you learn on Friday what you assumed on Monday, the assumption never has time to become an architecture. The clearest signal was that people stopped asking where things were.

The interesting constraint is almost never the one in the brief. It is comfortable, it is legible to management, and it is close to worthless once you measure what it actually changes. The caveat is that all of this assumes the underlying goal is settled, which is frequently the actual problem.

The quiet win

A shared definition of "done" removes more friction than any tool. Being right sixty per cent of the time builds exactly the kind of confidence that makes the other forty per cent expensive. It is worth saying that we have not run this long enough to be confident.

Consider the failure mode rather than the success case. Industry insights rewards clarity here more than almost anywhere else, because the wrong target produces work that looks productive and moves nothing. A useful test: if this disappeared tomorrow, how long before anyone noticed?

The cost of a bad decision is rarely the decision. It is the six months of building on top of it.

— Overheard in a retrospective

Begin with the obvious one

Speed and reversibility are the trade-off worth naming out loud. The decision is usually cheap and reversible; the execution is where the cost lives, and that is where the argument should have happened.

Documentation is a symptom: you write it where the design is unclear. When responsibility is spread across a group, the work that falls between the named parts is the work that does not happen. Reasonable people land elsewhere on this, usually because their constraints differ more than the vocabulary suggests.

The one people skip

The default answer is right often enough to be dangerous. The stated constraint is usually a proxy for a real one nobody wants to say aloud, and optimising the proxy is wasted effort. There are organisations where the opposite is true, and they are not obviously worse off.

The first thing to establish is what you are actually optimising for. The first quarter shows the intended effect; the second shows what the intended effect displaced. The version of this that works fits on an index card. The version that fails needs an onboarding session.

Where to start on Monday

Scope is the variable everyone adjusts last and should adjust first. The things that are easy to count are rarely the things that matter, and once a number reaches a dashboard it starts shaping behaviour whether or not it deserves to. One team we spoke to cut their review stage entirely and found throughput unchanged, which told them something the metrics had not.

The tooling question is downstream of the constraint question. They are decisions made quickly, defended slowly, and built upon for six months before anyone recalculates.

The habit that compounds

The second-order effects arrive about a quarter after the first-order ones. Subtraction is structurally underrated: the meeting that stopped happening leaves no artefact to point at in a review. In practice the answer showed up in the calendar before it showed up in the dashboard.

Consistency is worth more than any individual improvement to Industry insights. Most disagreements that present as strategic turn out, on inspection, to be two people using one word for two things. This is easier to write than to hold to when a deadline appears.

There is a version of Industry insights that is mostly ritual. Where a design is obvious the prose is short, so the length of an explanation is a reasonable proxy for where to look next.

The checklist we ended up with:

  1. Decide in advance what would make you stop
  2. Review the numbers monthly; change the targets rarely
  3. Prefer the reversible option when the evidence is thin
  4. Keep the feedback loop shorter than the planning cycle
  5. Write the constraint down before choosing a tool

The expensive mistake

The compounding effects matter far more than the individual wins. A small improvement applied consistently beats a dramatic one applied once, which is unsatisfying advice precisely because it is correct. When we mapped it out, four of the seven steps existed only to compensate for the second one.

Speed and reversibility are the trade-off worth naming out loud. Being right sixty per cent of the time builds exactly the kind of confidence that makes the other forty per cent expensive. In practice the answer showed up in the calendar before it showed up in the dashboard.

The tooling question is downstream of the constraint question. It is comfortable, it is legible to management, and it is close to worthless once you measure what it actually changes. The version of this that works fits on an index card. The version that fails needs an onboarding session.

The advice worth ignoring

What looks like a process problem is frequently an ownership problem. Cutting scope early is cheap and slightly embarrassing; cutting it late is expensive and deeply embarrassing. A useful test: if this disappeared tomorrow, how long before anyone noticed?

The interesting constraint is almost never the one in the brief. Teams that pick both end up with neither, and usually discover this at the point where reversing would have mattered. The caveat is that all of this assumes the underlying goal is settled, which is frequently the actual problem.

A shared definition of "done" removes more friction than any tool. A team that changes approach every quarter pays a coordination tax that routinely exceeds whatever the change was meant to fix.

The one that only matters at scale

Consider the failure mode rather than the success case. If you learn on Friday what you assumed on Monday, the assumption never has time to become an architecture.

The first thing to establish is what you are actually optimising for. The decision is usually cheap and reversible; the execution is where the cost lives, and that is where the argument should have happened. This is easier to write than to hold to when a deadline appears.

The default answer is right often enough to be dangerous. Subtraction is structurally underrated: the meeting that stopped happening leaves no artefact to point at in a review. The clearest signal was that people stopped asking where things were.

None of this generalises perfectly. Take the parts that map onto your constraints and discard the rest — that is what the framing is for.

baseline

Thoughts, stories and ideas.

Check your inbox for the link.

Free. Unsubscribe from the foot of any issue.