lenatriestounderstand

Chapter 32 of 32

Dark Knowledge: What a Classifier Says After Its Answer

Created Sep 19, 2026 Updated Sep 19, 2026

Show a trained classifier a photo of a cat and it answers cat. That is the part anyone looks at. What it actually produced is one number per class: something like 80% cat, 12% dog, 5% deer, a sliver for bird, almost nothing for truck. The ranking of the classes it did not choose is real information about the image, it reflects structure the model has learned, and it cannot be written in a label. That is what dark knowledge means — knowledge the model has and its top-1 answer never shows.

Why the leftovers are not noise. A confident network's non-top probabilities are tiny — 0.3%, 0.02% — and look like rounding. They are not. They come from perfectly ordinary logits, and consistent class-level structure appears in them: averaged over a whole test set, a network trained on CIFAR-10 puts its leftover probability for automobile on truck and ship, for cat on dog and bird, for horse on dog and deer. Nobody explicitly supplied those relationships — the labels supplied class identities and nothing about how the classes relate. They were learned from pixels.

Why a label cannot carry it. A label is one class: cat. It has nowhere to put "and this one is a bit dog-like, and definitely not a truck." Two very different cat photos — a clear one and one that is half-hidden — get the same target. A teacher's full distribution gives them different targets, and the difference is exactly the part the label threw away.

How you make it visible. Divide the logits by a temperature before the softmax:

p_k = exp(z_k / T) / Σ_j exp(z_j / T)

At T = 1 a well-trained network is nearly one-hot. At T = 4 the ordering of the wrong classes becomes legible — same model, same logits, nothing added.

This is mathematically the same logit scaling as the temperature in LLM decoding. What differs is what the softened distribution is for: in decoding, the next token is sampled from it; in distillation, it becomes a training target for the student. One operation, two jobs.

How much of it actually matters. This is where the popular version overshoots. Train a small network on a teacher's distributions instead of the labels and it improves — 77.7% to 79.9% in one controlled setup. But taking that gain apart shows the class ranking is not all of it. Replace the teacher's distribution with a uniformly softened label — same amount of probability moved off the top class, spread evenly, no structure at all — and most of the gain survives. Keep the teacher's exact entropy but scramble which wrong class gets which probability, and the student still lands close. On plenty of data, softness as such carries more of the benefit than the class geometry does.

The picture flips when data is scarce. With a tenth of the images, uniform softening hurts, and the real ranking of the wrong classes is worth several points over a scrambled one. One plausible reading: with enough examples a student can recover much of the class structure by itself, so the teacher's contribution is mostly regularization, while with few examples the teacher supplies structure the data no longer contains. The measurements are consistent with that; they do not establish it.

So dark knowledge is real, and it is not one thing. It carries at least three kinds of signal — overall target softness, example-specific confidence, and the relative geometry of the wrong classes — and they are not independent: confidence is itself expressed as softness on that example.

Dark knowledge is not a single hidden signal. It is a bundle of signals whose relative value depends on how much data the student has — which is why distillation can look mostly like regularization in one regime and like transfer of learned class structure in another. Measured in full, one ingredient at a time, in Knowledge Distillation: What a Teacher Actually Teaches.