AI systemsTechnical article

How AI backends handle millions of requests without falling over

A chat answer keeps a GPU, a block of memory and a connection busy for many seconds. Here is what that does to a backend, how the big serving systems cope, and which parts a startup should copy.

In short

  • A model request lasts seconds, not milliseconds, so the same traffic puts far more work in flight than a web page does.
  • Serving systems get their throughput from the scheduler and from memory management: continuous batching and paged key-value caches.
  • Startups usually break earlier, in their own code: handlers that wait on the model, queues with no limit and retries that arrive together.
  • The fixes are plain engineering. Answer in the background and stream, cap the line, meter tokens per user, back off with jitter and measure time to first token.

A model request is not a web request

Most of what backend engineers know about load comes from web requests. A page is built in tens of milliseconds, a worker finishes and takes the next one, and a single server can handle hundreds a second. Pool sizes, timeouts, autoscaling rules and load tests all quietly assume that shape.

A model request breaks those assumptions. It can run for ten seconds or more, it keeps expensive hardware busy the whole time, and the answer arrives a few words at a time instead of in one piece. The first thing to get right is how much that changes the arithmetic.

The rule is Little's law: the number of requests in flight equals the arrival rate times the time each one takes. Four requests a second at 50 milliseconds each means 0.2 requests in flight, which is nothing. Four requests a second at 12 seconds each means 48 in flight. A pool of 32 workers that felt generous for a website is full after eight seconds, and everything behind it waits.

Figure 1Same traffic, a different kind of request
4 / s × 12 s = 48 in flightMore than the 32 workers can hold, so a line forms.WaitingWorkersIn flight now32 / 32Waiting in line27Average wait4.6 s

Swipe sideways to see the whole figure

Request type
One square is one worker, and its fill shows how far along its request is. The AI case runs three times faster than real time so you do not have to wait twelve seconds per request.

Switch the figure between the two request types and drag the rate. The traffic is the same size in both cases. Only the duration changed, and that alone fills the pool.

What happens between the question and the first word

Inside an AI product, one request passes more stations than most diagrams show. The gateway checks the key and applies limits. A queue holds the request until there is room. A scheduler decides which GPU and which batch it joins. Only then does the model run, and it runs in two very different phases.

Figure 2Life of one request
Decode: one token per step, streamed to the user as it is madeYour appsends the promptGatewaykey, rate limitQueuewaits for a slotSchedulerpicks the batchGPU workerruns the modelINSIDE THE GPU WORKER1 Prefillwhole prompt at once2 Decodeone token per stepreads the weights and chat memory each stepWHAT THE USER SEESThe first thing to▍first token

Swipe sideways to see the whole figure

Simplified and slowed down: in a real system the whole trip takes a few seconds. Time to first token is set by the queue and by prefill. The speed of the words after that is set by decode.

Prefill reads the whole prompt in one parallel pass. It is heavy on arithmetic and it sets how long you wait for the first token. Decode then produces the answer one token at a time. Each step has to read the model's weights and the growing memory of the conversation, so it is limited by memory bandwidth more than by compute.

Two phases with two different bottlenecks mean two numbers to watch. Time to first token tells you about the queue and prefill. Tokens per second tells you about decode. An average latency figure hides both.

Why serving systems batch, and why the batching changed

A GPU is most efficient when it does many things at once, so serving systems group requests into batches. The first approach, static batching, collects a few requests, runs them together and returns when the batch is done. It has a flaw you can see in a picture: answers differ in length, so every request that finishes early sits idle while the longest one keeps the batch alive.

The Orca paper at OSDI 2022 fixed this with iteration-level scheduling, now usually called continuous batching. The scheduler decides again after every single step instead of after every batch, so a slot that frees up goes to a waiting request straight away.

Figure 3Static batching against continuous batching
Static batchingdone 12/12 · GPU busy 61%slot 1slot 2slot 3slot 4123456789101112Continuous batchingdone 12/12 · GPU busy 90%slot 1slot 2slot 3slot 4123456789101112all done at step 15static finished at step 22time, in decode steps →54 slot-steps of real work in both cases

Swipe sideways to see the whole figure

Same twelve requests, same four slots. The lengths are made up for the picture and the effect is real: hatched blocks are slots that finished and sit idle until the longest request in their group is done.

The Orca authors reported 36.9 times the throughput of NVIDIA FasterTransformer at the same latency, measured on a GPT-3 175B model. That is a best case from one benchmark, so do not plan capacity around it. What carries over is the lesson that the scheduler decides how much of an expensive machine is doing useful work.

Memory is the other limit

Every running request keeps a key-value cache, which holds the attention state for all the tokens so far. It grows with every token, and nobody knows in advance where an answer will end. Early systems reserved the longest possible length for each request in one contiguous block. That wastes memory on answers that end early and leaves gaps that nothing can use.

The vLLM paper, presented at SOSP 2023, borrowed an idea from operating systems. PagedAttention cuts the cache into small fixed blocks and hands them out on demand, with a table that tracks where each request's blocks live. The authors report 2 to 4 times the throughput of FasterTransformer and Orca at the same latency. More requests fit in the same memory, and bigger batches are where the throughput comes from.

Figure 4Memory for the conversation: reserved, or paged
Reserve the longest answereach request books 24 blocks up frontHand out small blocks as neededa request takes blocks one by one as it growsrunning 3 · holds real data 31%memory booked 100%Request D is waiting for memoryrunning 4 · holds real data 38%memory booked 38%Request D started straight away

Swipe sideways to see the whole figure

An illustration, not a measurement. Each square stands for a block of the key-value cache. Solid squares hold real data, outlined squares are reserved and empty.

Where a startup's AI backend breaks

Most startups never touch a GPU. They call a hosted model, and the failures happen in their own code, in places that look fine in a demo.

  • A request handler waits for the model. It holds a thread and often a database connection for the whole answer, so twenty slow answers can starve the rest of the site.
  • The waiting line has no limit. A spike turns into a delay that never clears, people refresh, and the refreshes join the line.
  • Retries have no backoff. The provider returns a rate-limit error, every client retries at once, and the retries become the incident.
  • Everyone shares one quota. One user's agent loop spends the tokens everyone else needs.
  • Cost gets looked at last. A runaway loop is an outage and a bill at the same moment.
Figure 5The same server, with and without a limit on the line
This server finishes about 10 requests a second12 arrive each second, so the line cannot clear.Waiting now106Average wait9.1 sTurned away (429)01209060300last 30 seconds →requests waiting in line

Swipe sideways to see the whole figure

Waiting line
A single server that finishes about ten requests a second on average. Arrivals are random, which is why the line moves even below the limit. Move the slider past ten and watch the difference.

The figure uses a server that finishes ten requests a second. Below that rate the line comes and goes. Above it the line grows for as long as the traffic lasts, and each person in it waits longer than the one before. Capping the line turns the overload into quick, honest failures for the extra requests, which a client can handle, in place of slow failures for everybody.

Figure 6A retry storm, and the same outage with jittered backoff
Retry every secondall in step · failed attempts 2,891100Backoff with full jitterspread out · failed attempts 9911000s5s10s15s20s25sservice downservice downcaught up after 18 scaught up after 7 s

Swipe sideways to see the whole figure

Attempts per second, green when the service accepted them and red when it turned them away. The service can handle 100 a second. The red band is the outage, and the same 300 clients and the same outage are used in both charts.

Retries need their own warning. In the figure, 300 clients lose their connection at the same instant. Retrying every second keeps them marching in step, so the service takes a wave of requests, serves a fraction and fails the rest, then gets the same wave again a second later. Marc Brooker's 2015 write-up on the AWS blog measured the alternatives and found that full jitter did the least work. Google's incident report for 12 June 2025 names the same missing piece: tasks that restarted overloaded the database they depend on, and the service had no randomized exponential backoff in place.

Patterns that hold up under load

Answer in the background and stream the result

Accept the request, run the model call in a worker and push tokens to the browser over server-sent events or a WebSocket. The web tier stays fast whatever the model does, and a slow answer costs one open stream instead of a blocked thread.

Limit the line and say no early

Give the queue a maximum length and answer extra requests with a 429 and a Retry-After header. A fast no is cheaper for you and kinder to the user than a slow yes that times out.

Meter tokens as well as requests

Providers limit both requests and tokens per minute, and your own limits should do the same. A token bucket per user means a runaway loop drains that user's bucket and nobody else's.

Retry with backoff, jitter and a budget

Wait a random time up to a ceiling that doubles with each attempt, honour Retry-After when the provider sends it, and cap the total number of attempts. Retry only calls that are safe to repeat, or attach an idempotency key.

async function callWithRetry(fn, { tries = 5, base = 500, cap = 20000 } = {}) {
  for (let attempt = 0; ; attempt++) {
    try {
      return await fn();
    } catch (err) {
      if (attempt + 1 >= tries || !isRetryable(err)) throw err;
      const ceiling = Math.min(cap, base * 2 ** attempt);
      const wait = err.retryAfterMs ?? Math.random() * ceiling; // full jitter
      await sleep(wait);
    }
  }
}

Cache what repeats

Providers offer prompt caching for long prompts that share a prefix, which cuts both cost and time to first token. A cache of your own for repeated questions, embeddings and tool results saves calls you would otherwise pay for.

Plan for the provider being down

Set one timeout for the first token and another for the whole answer. Put a circuit breaker in front of the provider, and keep a second provider or a smaller model ready as a fallback for the questions that can use it.

Measure the right things

Chart time to first token, tokens per second, queue depth, cost per request and error rate by cause. An average response time tells you almost nothing here.

How this looked in a real product

Oh Crap! Chat is a paid AI coach we built with Prologue Partnerships on Jamie Glowacki's potty-training method. More than 700 parents have paid for access and over 400 are active; the case study has the full picture. Several of the patterns above are in it. Answers stream to the browser as Gemini writes them. The API has one time limit for the first words and another for the whole answer, and it retries with a growing delay when the model is busy, so a slow moment at Google becomes a short pause on screen. Access is checked before any model call, which means an expired pass never spends tokens, and a dashboard shows AI usage and estimated daily cost next to revenue.

Questions to answer before the first spike

  1. What happens to a user's request when the model takes 60 seconds?
  2. How long can the waiting line get, and what does the next request receive when it is full?
  3. What stops one user from using the whole token budget?
  4. Do retries wait a random time, and is there a cap on attempts?
  5. Which model or provider takes over when the main one is down?
  6. Can you see time to first token and cost per request today, for each customer?

Sources

  1. Orca: A Distributed Serving System for Transformer-Based Generative Models (OSDI 2022)
  2. Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM, SOSP 2023)
  3. Marc Brooker, Exponential Backoff And Jitter
  4. Google Cloud incident report, 12 June 2025

The animated figures are simplified illustrations that use the numbers stated in the text. Where a figure is modelled on a real incident, the sources above are the full accounts.

The service behind this articleAI Systems, Agents, and Workflow AutomationRelated case studyPrologue Partnerships: an author's book, turned into a paid AI coach
Aslam Sarfraz profile photo

Aslam Sarfraz

AI, DevOps and full-stack engineer

More articles

Full-stack engineering · 8 min read

Why startup backends break at the first big traffic spike

Ticketmaster took 3.5 billion requests in one sale and Segment merged more than 140 services into one. The failures behind spikes are rarely exotic, and each one has arithmetic you can do before launch.

DevOps and reliability · 8 min read

How one config change takes down production, and how to contain yours

CrowdStrike, Google Cloud, AWS and Cloudflare each had a major outage between July 2024 and November 2025, and none began with an attack. What happened, what it teaches and what to put in place.

LinkedIn outreach · 7 min read

Designing a LinkedIn outreach pipeline: a fixed capacity and an unseen limit

Outreach is a pipeline with a fixed capacity, a limit nobody publishes and very small samples. The planning numbers I use and the engineering behind them.