The Engine Is Allowed to Forget
An inference server stays fast by treating every request as its first. Agents need the opposite. vLLM's agentic layer puts the memory beside the engine, and most of its design follows from one rule about who owns the history.