← All posts

Blog

Inference: Prefill and Decode

Views —

Why LLM inference separates prompt processing from token-by-token generation.

The full article is temporarily unavailable in this reader.

Read next

Nov 20, 2025Learnings from Mankind’s Search for Meaning