From one little question
to 100 million people.
26 little lessons, A–Z. From tokens and GPUs to retrieval, tools, and trustworthy answers. No big words before their time.
A LITTLE IDEA
THE REAL NAME
This is a learning model of a large chat service, not a description of OpenAI’s private infrastructure. Animations are simplified. GPU throughput is a made-up, editable benchmark input, never a hardware performance claim. Training and serving are separate capacity budgets.
- NVIDIA DGX B200: eight GPUs, 1,440 GB GPU memory, 14.3 kW maximum system power
- NVIDIA DGX H100 / H200: GPU memory specifications
- NVIDIA GB200 NVL72: 72 GPUs, 18 compute trays, nine NVLink switch trays
- NVIDIA: the NVLink fabric and the separate cluster network
- vLLM: paged KV memory
- vLLM: measure serving throughput and latency
- Retrieval-augmented generation
- LoRA: adapting models with low-rank updates
- Hugging Face: temperature and top-p
- Anthropic: tools, workflows, and agents
Specifications checked September 2026. Product variants differ. Network calculations use an idealized, fully provisioned two-tier fabric, not a deployment-ready network design.