Ten questions on the engineering around large language models: tokens and sampling, how inference servers turn GPUs into throughput, retrieval, rate limits, caching, evals and tool use. Each question is a decision a team running a model platform makes. Pick the answer you would defend in a design review.