A TINY NANOGPT LESSON · 01

What if we just… skipped some work?

The next word should be cat. Do we need to score every other word to learn that?

Switch off a word.
Watch what changes.

Keep cat — it’s the correct answer.

Score all fourScore only checked words

output columns kept / probability mass left out

Toy counts, not measured speed. The base scores stay fixed.

The math, the real recipe & sources

More confident ≠ more correct.

The gray bars start at 60%, 30%, 6%, and 4%. These are invented probabilities, not predictions from a trained language model. Removing choices shrinks the denominator; it doesn’t improve the model.

Each green probability is its original probability divided by the total probability of the kept words. A skipped word is shown as zero in this restricted distribution, not zero in the original model.

What are the little arrows?

They show the direction a gradient-descent step would push each output score: ↑ raise, ↓ lower, — no push. The numbers are the logit gradients, p − y, before → after. Here y = 1 for cat and 0 otherwise. A negative gradient raises a score; a positive one lowers it. Omitted output columns receive no direct gradient in this toy.

How this connects to the post

This is a toy restricted softmax, not the actual sampled-softmax implementation. The linked NanoGPT recipe keeps every local-batch target and adds deduplicated sampled non-targets. It increases the candidate set over training and finishes with the full vocabulary; validation uses the full vocabulary too.

We know all four probabilities here so you can see the trade-off. Selecting useful candidates isn’t magically free, and real training doesn’t get the full probabilities for free. Skipping low-probability words here is a teaching comparison, not a claim about the recipe’s selection rule.

The wider record also combines sparse embedding updates, sparse communication and optimizer state, custom attention, EMA, Anvil2 and network changes. Its huge parameter count comes from sparse embeddings, not a dense 65B transformer.

No training, GPU timing, speedup measurement, or NanoGPT benchmark reproduction happens here. Fewer columns and lower restricted loss do not establish better language modeling or a proportional speedup.

Runs in your browser. No uploads, APIs, analytics, or saved data. The hosting provider may keep ordinary request logs.