Design a Streaming Chat Completions API
We are opening our language models to developers through a public API: they send a conversation and get the model's reply. We want to design that API and the service that sits between developers and our GPUs.
This brief is incomplete on purpose, as it would be in a real interview. Ask the interviewer about the users, the features, the targets and the traffic. Whatever you uncover is added below.
- A developer sends a conversation and receives the model's reply streamed token by token as it is generated.
- Not uncovered yet
- Not uncovered yet
- Not uncovered yet
- Not uncovered yet
- Not uncovered yet
- Not uncovered yet
- Not uncovered yet
- Not uncovered yet
- Not uncovered yet
- Not uncovered yet
- Not uncovered yet
- Not uncovered yet
- The API: the request, the streamed response and its errors
- The request path from a client to a GPU and back
- How usage is metered and billed
- The capacity estimate behind your choices
Replies stream over server-sent events (or chunked HTTP) with typed events (content deltas, tool calls, usage, done, errors) and event ids; a dropped client resumes from its last event id, and a cancellation propagates all the way to the worker that is generating.
Stateless gateways hold the long-lived connections and do authentication and rate limiting, a scheduler places requests on inference workers with free capacity (in another cluster when one fails), and overload is signalled with 429 or queueing rather than timeouts.
Per-key request and token budgets live in a shared fast store (token buckets), tokens are reserved from the prompt size up front and corrected when the reply ends, and the hottest keys are not a single hot shard.
Clients send an idempotency key; usage is emitted once per request (on completion or cancellation) as an event with the request id into a durable log, and aggregation deduplicates on that id.
3 billion requests a month is about 1,200 a second on average and 4,600 at peak; at 8 seconds per reply that is roughly 37,000 concurrent streams at peak and almost 2 million output tokens a second, which sizes the gateways and the GPU fleet.
Every functional requirement in the brief is visibly served by something on the board, and the non-functional targets are addressed rather than ignored.
Components are labelled, data flows are drawn as connections between them, and the direction of each flow is unambiguous.
Concentrate on the public API, the request path to the inference fleet, rate limiting and usage metering. How the model runs on a GPU, training and the developer dashboard are out of scope: assume inference workers that accept a prompt and emit tokens one at a time.
- Views
- 2