GPT-2 (355M)
The same architecture as the other demos, scaled up to OpenAI's 355M-parameter GPT-2 (24 layers, 1024-dim embeddings) and loaded from the released open weights — book chapter 5. Larger than the 124M model, so its completions tend to hang together better. Start a phrase and watch it continue.