Design an Inference Router for a Model API
Our model API runs on GPU clusters in three regions. Since we launched a cheaper batch tier and long-context models, some clusters sit at 30% while others turn requests away, and our enterprise customers' latency has doubled. We need to rethink how requests get to GPUs.
This brief is incomplete on purpose, as it would be in a real interview. Ask the interviewer about the users, the features, the targets and the traffic. Whatever you uncover is added below.
Nothing uncovered yet.
Nothing uncovered yet.
Nothing uncovered yet.
- The request path from the API to a replica
- How the router decides where a request goes
- How the tiers share the fleet
- The capacity estimate behind your choices
The router places requests using fresh replica telemetry (models loaded, running requests, free KV-cache memory, queue length), with a choice rule such as power-of-two-choices or least-loaded that avoids sending every request to the same idle replica at once.
Requests are hashed by customer and prefix onto a small set of replicas (consistent hashing with bounded load), so cached prefixes are reused, but a popular prefix spreads to more replicas instead of overloading one.
Interactive, standard and batch have separate queues; batch fills idle capacity and is preempted or paused when interactive demand rises; reserved capacity is lent out and reclaimed within seconds; utilisation stays high without hurting interactive latency.
A global layer chooses the region and spills interactive traffic to another region when the local one is saturated, within the latency budget; replicas are health-checked and drained, and interactive requests on a failed replica are retried elsewhere.
4 million developers, 20% active, 1,000 requests each is 800 million requests a day: about 9,300 a second on average and 28,000 in the busiest hour; at 10 seconds each that is roughly 280,000 requests in flight, about 4,600 large-model replicas at 60 each, before batch is shifted to quiet hours.
Every functional requirement in the brief is visibly served by something on the board, and the non-functional targets are addressed rather than ignored.
Components are labelled, data flows are drawn as connections between them, and the direction of each flow is unambiguous.
Concentrate on deciding which GPU replica serves each request, and on how different kinds of traffic share the fleet. The public API, billing, and how a model runs on a GPU are out of scope: assume replicas that report their state and run any request they are given.
- Views
- 2