5 min read
Autoscale is a lagging indicator
What a storm taught me about scaling a customer platform: the metric you scale on is already late, and the cheapest fix is the read you stop serving.
Your platform is fine. It has been fine for two years. Then the day arrives that it exists for — for a utility’s customer site that is a storm, for a retailer it is a sale, for a payroll system it is the last working day of the month — and the instances max out, the gateway starts throttling, and the dashboard you never wired to an alert shows you, in real time, exactly how fine it is not.
I was a junior developer on a platform like that when its storm came. Twenty-plus people on the account, a design I did not own, and two of the fixes that followed were mine. This is what the day taught me, written for the version of you who has autoscale switched on and has never watched it lose.
What actually gave way
Not the database. Not the slow downstream everyone worried about. The tier that fell over first was the one with the most elastic-sounding name: the App Service plan, whose instances hit their ceiling, and the API gateway in front of it, whose throttling then did precisely what it was configured to do.
That ordering matters, because it tells you where the levers are. When the front of the path gives way first, capacity and traffic shape are the problem, and everything behind them is a second-order concern for the day.
Scroll the diagram sideways
- 1 Callers — the website and the contact centre’s tools. On a storm day, several times the usual number at once.
- 2 The API gateway. Where throttling tripped, and where the reference reads are now answered from cache.
- 3 The service tier, one service per downstream. Instances maxed out here first; autoscale now adds them sooner.
- 4 The database, the system of record — kept to the size of the work by scheduled purge jobs.
- 5 A client-owned SOAP layer over a mainframe: fixed capacity, and where a colleague’s circuit breaker went.
- 6 Telemetry that recorded every hop and alerted nobody.
The metric you scale on is already late
Autoscale on CPU percentage is the default, and it is what we had. It works, in the sense that instances do arrive. But CPU is a lagging indicator of load: it climbs after requests have started queueing, and by the time the rule fires and the new instance warms up, the queue behind it has been growing for minutes. During a surge that is the whole event.
Three things follow, in the order I would do them today:
- Scale on a leading signal. HTTP queue length on the plan — requests waiting for a worker — moves before CPU does, and it moves in proportion to what callers are experiencing. If the platform reports it reliably, it is the better rule.
- Make the rule asymmetric. Scale out fast and on a low threshold; scale in slowly and on a long cool-down. The cost of an unneeded instance for twenty minutes is small. The cost of not having it is the outage page being an outage.
- Accept the lag you cannot remove, and plan for it. Every rule fires late by some margin. Know the margin, and make sure the platform survives it — which is what the next section is for.
We scaled on CPU because it was the signal the platform reported and the team trusted, and the instances maxing out was the first failure, so instance count was the first lever. It was the right first move and the wrong final state.
The cheapest fix is the read you stop serving
Here is what the storm traffic was mostly made of: the same handful of reference and lookup reads behind every page — the kind of data nothing in the storm changes — each one travelling through the gateway to a service to a database to return the answer it returned an hour earlier.
Caching those at the gateway took them off the service tier and the database entirely. It was cheaper than making either faster, and it cost nothing to build: a response-cache policy on operations that already existed.
Two rules for doing this without regret:
- Cache what the event does not change. Reference data, lookups, static lists. Not outage status, which changes by the minute and is the one thing everyone has come for.
- State the staleness window and defend it. A cache is a decision that some reader will see an answer up to N minutes old. Say N. If you cannot defend N for that data, do not cache it.
The fixes that were not mine
The queue-based decoupling of writes — accepting an outage report onto a queue and processing it asynchronously, so a burst of them never waited on the database — and the timeouts and circuit breaking in front of the fixed downstream were colleagues’ work. They mattered at least as much as anything above. I mention them for two reasons.
First, because a write-up that quietly absorbs them would be a better story and a worse record.
Second, because they are the part of the response that depended on the design. The service boundaries had been drawn by dependency, so there was a place to put a circuit breaker in front of the SOAP layer and nowhere else; the writes could be queued because the outage-report service owned its own path. If you are the junior in the room, the boundaries someone else drew decide what you can fix. Learn to read them before the day comes.
What I would check on your platform tomorrow
- Which tier gives way first under a surge, and whether the autoscale rule is on a signal that leads or lags that failure.
- Whether the scale-out rule is as aggressive as the scale-in rule is cautious. It usually is not.
- What fraction of your traffic is reads that the event does not change, and whether the gateway is answering them or merely forwarding them.
- Whether anything alerts. Telemetry that records everything and pages nobody is how a surge gets discovered by a person watching a chart — which is how we discovered ours.
The platform survived its next storm. The case study, with the decisions and what each one cost, is on the portfolio: The day the platform was built for.