
The Information Domain: Entropy as a Codelength, and Which of Your Objectives Are One by Theorem
A standalone tutorial, not a survey: entropy as a coding quantity with its asymptote attached, the cross-entropy floor under both label-smoothing conventions and the table of which held-out loss owns which floor, the plug-in mutual-information estimator as an uncertain measurement with a report format, discrete versus differential entropy in d dimensions before any VAE or kernel claim, the information plane measured three ways (exact count at float64 with a readout cast to float32 and float16, binned, noisy readout at absolute and scale-normalised noise), the ELBO as an exact split into a distortion and an upper bound on rate, with the three failures called posterior collapse and the gap between a β-VAE's curve and Shannon's R(D), the I-MMSE floor under a diffusion loss with both sides integrated, Kolmogorov–Szegő with its convention and hypotheses displayed and its circulant check kept apart from the Toeplitz limit and from the jitter it rides on, hypothesis-testing exponents (Stein, Chernoff, Sanov kept apart) checked against exact finite-sample errors, the Cover & Thomas and MacKay results with proof sketches, modern consumers from InfoNCE to speculative decoding, six checkpoint exercises, and an appendix that grades every connection to our earlier posts by how much it actually proves.
