In short
- A model request lasts seconds, not milliseconds, so the same traffic puts far more work in flight than a web page does.
- Serving systems get their throughput from the scheduler and from memory management: continuous batching and paged key-value caches.
- Startups usually break earlier, in their own code: handlers that wait on the model, queues with no limit and retries that arrive together.
- The fixes are plain engineering. Answer in the background and stream, cap the line, meter tokens per user, back off with jitter and measure time to first token.
A model request is not a web request
Most of what backend engineers know about load comes from web requests. A page is built in tens of milliseconds, a worker finishes and takes the next one, and a single server can handle hundreds a second. Pool sizes, timeouts, autoscaling rules and load tests all quietly assume that shape.
A model request breaks those assumptions. It can run for ten seconds or more, it keeps expensive hardware busy the whole time, and the answer arrives a few words at a time instead of in one piece. The first thing to get right is how much that changes the arithmetic.
The rule is Little's law: the number of requests in flight equals the arrival rate times the time each one takes. Four requests a second at 50 milliseconds each means 0.2 requests in flight, which is nothing. Four requests a second at 12 seconds each means 48 in flight. A pool of 32 workers that felt generous for a website is full after eight seconds, and everything behind it waits.
Swipe sideways to see the whole figure
Switch the figure between the two request types and drag the rate. The traffic is the same size in both cases. Only the duration changed, and that alone fills the pool.
What happens between the question and the first word
Inside an AI product, one request passes more stations than most diagrams show. The gateway checks the key and applies limits. A queue holds the request until there is room. A scheduler decides which GPU and which batch it joins. Only then does the model run, and it runs in two very different phases.
Swipe sideways to see the whole figure
Prefill reads the whole prompt in one parallel pass. It is heavy on arithmetic and it sets how long you wait for the first token. Decode then produces the answer one token at a time. Each step has to read the model's weights and the growing memory of the conversation, so it is limited by memory bandwidth more than by compute.
Two phases with two different bottlenecks mean two numbers to watch. Time to first token tells you about the queue and prefill. Tokens per second tells you about decode. An average latency figure hides both.
Why serving systems batch, and why the batching changed
A GPU is most efficient when it does many things at once, so serving systems group requests into batches. The first approach, static batching, collects a few requests, runs them together and returns when the batch is done. It has a flaw you can see in a picture: answers differ in length, so every request that finishes early sits idle while the longest one keeps the batch alive.
The Orca paper at OSDI 2022 fixed this with iteration-level scheduling, now usually called continuous batching. The scheduler decides again after every single step instead of after every batch, so a slot that frees up goes to a waiting request straight away.
Swipe sideways to see the whole figure
The Orca authors reported 36.9 times the throughput of NVIDIA FasterTransformer at the same latency, measured on a GPT-3 175B model. That is a best case from one benchmark, so do not plan capacity around it. What carries over is the lesson that the scheduler decides how much of an expensive machine is doing useful work.
Memory is the other limit
Every running request keeps a key-value cache, which holds the attention state for all the tokens so far. It grows with every token, and nobody knows in advance where an answer will end. Early systems reserved the longest possible length for each request in one contiguous block. That wastes memory on answers that end early and leaves gaps that nothing can use.
The vLLM paper, presented at SOSP 2023, borrowed an idea from operating systems. PagedAttention cuts the cache into small fixed blocks and hands them out on demand, with a table that tracks where each request's blocks live. The authors report 2 to 4 times the throughput of FasterTransformer and Orca at the same latency. More requests fit in the same memory, and bigger batches are where the throughput comes from.
Swipe sideways to see the whole figure
Where a startup's AI backend breaks
Most startups never touch a GPU. They call a hosted model, and the failures happen in their own code, in places that look fine in a demo.
- A request handler waits for the model. It holds a thread and often a database connection for the whole answer, so twenty slow answers can starve the rest of the site.
- The waiting line has no limit. A spike turns into a delay that never clears, people refresh, and the refreshes join the line.
- Retries have no backoff. The provider returns a rate-limit error, every client retries at once, and the retries become the incident.
- Everyone shares one quota. One user's agent loop spends the tokens everyone else needs.
- Cost gets looked at last. A runaway loop is an outage and a bill at the same moment.
Swipe sideways to see the whole figure
The figure uses a server that finishes ten requests a second. Below that rate the line comes and goes. Above it the line grows for as long as the traffic lasts, and each person in it waits longer than the one before. Capping the line turns the overload into quick, honest failures for the extra requests, which a client can handle, in place of slow failures for everybody.
Swipe sideways to see the whole figure
Retries need their own warning. In the figure, 300 clients lose their connection at the same instant. Retrying every second keeps them marching in step, so the service takes a wave of requests, serves a fraction and fails the rest, then gets the same wave again a second later. Marc Brooker's 2015 write-up on the AWS blog measured the alternatives and found that full jitter did the least work. Google's incident report for 12 June 2025 names the same missing piece: tasks that restarted overloaded the database they depend on, and the service had no randomized exponential backoff in place.
Patterns that hold up under load
Answer in the background and stream the result
Accept the request, run the model call in a worker and push tokens to the browser over server-sent events or a WebSocket. The web tier stays fast whatever the model does, and a slow answer costs one open stream instead of a blocked thread.
Limit the line and say no early
Give the queue a maximum length and answer extra requests with a 429 and a Retry-After header. A fast no is cheaper for you and kinder to the user than a slow yes that times out.
Meter tokens as well as requests
Providers limit both requests and tokens per minute, and your own limits should do the same. A token bucket per user means a runaway loop drains that user's bucket and nobody else's.
Retry with backoff, jitter and a budget
Wait a random time up to a ceiling that doubles with each attempt, honour Retry-After when the provider sends it, and cap the total number of attempts. Retry only calls that are safe to repeat, or attach an idempotency key.
async function callWithRetry(fn, { tries = 5, base = 500, cap = 20000 } = {}) {
for (let attempt = 0; ; attempt++) {
try {
return await fn();
} catch (err) {
if (attempt + 1 >= tries || !isRetryable(err)) throw err;
const ceiling = Math.min(cap, base * 2 ** attempt);
const wait = err.retryAfterMs ?? Math.random() * ceiling; // full jitter
await sleep(wait);
}
}
}Cache what repeats
Providers offer prompt caching for long prompts that share a prefix, which cuts both cost and time to first token. A cache of your own for repeated questions, embeddings and tool results saves calls you would otherwise pay for.
Plan for the provider being down
Set one timeout for the first token and another for the whole answer. Put a circuit breaker in front of the provider, and keep a second provider or a smaller model ready as a fallback for the questions that can use it.
Measure the right things
Chart time to first token, tokens per second, queue depth, cost per request and error rate by cause. An average response time tells you almost nothing here.
How this looked in a real product
Oh Crap! Chat is a paid AI coach we built with Prologue Partnerships on Jamie Glowacki's potty-training method. More than 700 parents have paid for access and over 400 are active; the case study has the full picture. Several of the patterns above are in it. Answers stream to the browser as Gemini writes them. The API has one time limit for the first words and another for the whole answer, and it retries with a growing delay when the model is busy, so a slow moment at Google becomes a short pause on screen. Access is checked before any model call, which means an expired pass never spends tokens, and a dashboard shows AI usage and estimated daily cost next to revenue.
Questions to answer before the first spike
- What happens to a user's request when the model takes 60 seconds?
- How long can the waiting line get, and what does the next request receive when it is full?
- What stops one user from using the whole token budget?
- Do retries wait a random time, and is there a cap on attempts?
- Which model or provider takes over when the main one is down?
- Can you see time to first token and cost per request today, for each customer?
Sources
- Orca: A Distributed Serving System for Transformer-Based Generative Models (OSDI 2022)
- Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM, SOSP 2023)
- Marc Brooker, Exponential Backoff And Jitter
- Google Cloud incident report, 12 June 2025
The animated figures are simplified illustrations that use the numbers stated in the text. Where a figure is modelled on a real incident, the sources above are the full accounts.
