Appearance
The 250GB training step
Kimi-K3 has 2.8 trillion parameters. Loading it takes roughly 3TB of VRAM just to hold the weights. That's why distillation became standard practice: train a smaller student to recover the teacher's capabilities, then deploy the student. Nvidia's Nemotron 3 Puzzle 75B and Multiverse Computing's Hypernova 60B are both products of this pipeline.
The distillation step decides most of the final quality, and it's the most expensive part of the pipeline. Online KL distillation keeps teacher and student in memory at the same time. Every training step runs a full teacher forward pass, even though the teacher's behavior never changes across the run.
The numbers get absurd fast. gpt-oss-120b has a vocabulary of 201,088 tokens. At sequence length 32K and batch size 4, the teacher probability tensor alone is 4 × 201,088 × 32,768. In bf16, that's about 50GB for one tensor. Add gradients, activations, weights, and optimizer states, and a single distillation iteration peaks near 250GB of VRAM. That exceeds what an H200 (141GB) or a B200 can provide. You need multiple GPUs and careful tensor parallelism just to run one step.
Four recent lines of work attack this wall from different angles. One caches the teacher. One chunks the loss. One decomposes pretraining itself. One streams layers out of VRAM. They share a theme: the expensive thing is materializing data that doesn't need to exist.
Stop recomputing the teacher
The first fix, from Multiverse Computing's distillation paper, is almost insultingly simple. Run the teacher once, cache the top-100 most likely tokens per position, and train the student against that cache. The teacher never sits in memory during training. The cache is reusable across ablations, so a whole hyperparameter sweep costs one teacher pass.
Is top-100 enough? In their benchmarks, yes. The loss curves for online distillation and offline distillation with cached top-100 logits overlap almost exactly. Offline training against a cache is lossless relative to online training at this scale.
Don't build the grid
The second fix targets the KL loss itself. To compute KL divergence, you need a number for every token position and every vocabulary entry: a grid of vocab × sequence length. With a 100K+ token vocabulary and long sequences, that grid is enormous. Default implementations build the whole thing before producing a single scalar.
The paper compares three mathematically equivalent ways to compute the same loss:
- Dense KL rebuilds a full teacher-probability grid from the cached logits and compares it against the student's dense grid. It's the correctness baseline, but it holds the vocab × sequence grid in memory twice.
- Forward-chunked KL keeps the teacher sparse and computes the loss slice by slice. It's the fastest of the three at 8K context. But the student's own logits are still computed in full and kept for the backward pass.
- Fused chunked KL fuses the output projection into the loss. It never produces the student's full logits grid. One chunk of the sequence goes end to end: project hidden states to logits, fold into the running loss, discard, move on. The backward pass recomputes each chunk on the fly.
The cost of fusing is doing the output projection twice, once forward and once in backward. In exchange, peak memory grows linearly with sequence length instead of spiking with vocab × sequence.
Quick Take: The cheapest distillation win is structural. Cache the teacher once, chunk the loss, and never build the full vocab × sequence grid.
What the memory numbers actually mean
Here's the head-to-head on a single H200 with Llama 3.1 8B Instruct as teacher and a 3.2B Llama student at 8K context:
| Method (8K context, single H200) | Peak memory | Iteration time | Throughput |
|---|---|---|---|
| Online distillation | 102.8 GB | 25.9 s | 237 TFLOP/s |
| Offline, dense KL | 78.3 GB | 18.5 s | 331 TFLOP/s |
| Offline, forward-chunked KL | 61.8 GB | 18.4 s | 335 TFLOP/s |
| Offline, fused chunked KL | 58.3 GB | 20.2 s | 304 TFLOP/s |
At 8K, the fused chunked loss isn't the fastest option. The extra backward projection costs about 10% throughput versus forward-chunked. Its advantage only shows up as context grows.
At 32K tokens, peak memory falls from 85.2 GiB with the dense loss to 5.45 GiB with the fully chunked version. That's a 15.6× reduction, and the dense loss fails outright from 64K onward. At 256K, the chunked loss uses 11.6 GiB against 134.2 GiB for the next-best variant, and runs about 3.3× faster per iteration.
The practical payoff: distilling a GPT-OSS 20B model at 32,768-token context shrank from four GPU nodes to one. Step time fell from 57.0 to 12.23 seconds, about 5× faster. Throughput per GPU went from 74.2 to 345.7 TFLOP/s.
Key numbers: online distillation peaks around 250GB VRAM per step on a 32K run. Cached top-100 logits drop that to 78GB at 8K. The fused chunked loss lands at 58GB, and at 32K it cuts memory 15.6×. The resulting 3.2B student keeps most of an 8B teacher's accuracy at less than half the parameters.
Pretraining gets modular
The same "don't materialize what you don't need" logic applies to pretraining itself. Mixture of Training (MoT) asks whether pretraining can be decomposed into smaller, independently trainable jobs that later recompose into a coherent model. The procedure partitions a target Transformer into contiguous layer blocks, trains each block inside a frozen pretrained aligner scaffold, then recomposes the blocks with an optional short end-to-end adaptation pass.
On a 1.3B Gemma-style model trained on C4, this is a proof of mechanism: independently trained depth slices recompose into a usable language model, and a quality-parity schedule reaches the same perplexity as the monolithic baseline.
The authors are careful about what this does and doesn't buy. The parity setting processes more aggregate tokens, and the critical path advantage depends on reusing the aligner across runs. If the aligner is single-use, the economics don't work. MoT is presented as a small-scale framework for studying whether scaffolded sub-runs can act as reusable training units, not as a general replacement for monolithic pretraining.
Normalization placement is a curriculum decision
A second pretraining paper, Rethinking Normalization Placement for LLMs, finds that a design choice everyone stopped thinking about matters again once you grow depth gradually. Pre-norm is the standard placement because it makes joint optimization of full-depth models work. But in curriculum depth growth, each appended block receives the boundary representation produced by a trained prefix, so placement affects forward conditioning.
In a controlled distillation study with a Qwen3-8B teacher and a nine-layer student, pre-norm and post-norm are indistinguishable under joint training, differing by 0.0004 validation CE. That's noise. Under curriculum growth, post-norm improves over pre-norm by 0.0328 validation CE, an order of magnitude larger. The ranking crosses over during the curriculum: post-norm takes the lead once blocks are appended.
Boundary diagnostics explain why. Post-norm keeps residual scales stable. Pre-norm shows structural-token scale drift, and on a fixed batch, the final pre-grow block is nearly identity-mapped. The takeaway: normalization placement and training curriculum are coupled design choices, at least in distillation settings. If you grow depth, test post-norm.
The tooling layer: fine-tuning on a 4GB laptop
The third attack is tooling. Soup is a one-command fine-tuning and post-training CLI whose headline feature is layer streaming: the frozen base model stays out of VRAM and feeds the GPU one decoder layer at a time, so only the adapter trains in memory.
Measured on an RTX 3050 Laptop with 4GB VRAM: Llama-3.1-8B-Instruct with NF4 quantization runs at 119.6 tok/s with 3.32GB peak, bit-exact against a normal resident run. The same config was reproduced independently on an H100 at 113.00 tok/s in the same 3.32GB.
Preference tuning works over layer streaming too. DPO normally needs a reference model, and a second copy would double memory. Soup uses the same streamed base with adapters switched off, measured at 0.914× the SFT peak, where forcing a real second instance cost +730MB, exactly one copy of the weights. The honest cost is time: DPO reads the layer stack 1.52× as often per step.
The practical envelope, with QLoRA 4-bit:
| VRAM | Max model | Example |
|---|---|---|
| 8 GB | ~7B | Llama-3.1-8B, Mistral-7B |
| 16 GB | ~14B | Phi-4-14B, Qwen2.5-14B |
| 24 GB | ~34B | CodeLlama-34B, Yi-1.5-34B |
| 48 GB | ~70B | Llama-3.3-70B |
| 80 GB+ | 70B+ (full) or MoE | Mixtral-8x22B, DeepSeek-V3 |
What the community is running into
The changelogs around these tools read like a catalog of the ways training pipelines quietly lie to you. I've hit several of these myself.
Greedy decoding is not deterministic on GPU. I've watched five runs of the same model, no adapter, spread 0.015 to 0.020 against a 0.05 significance threshold, with four of six paired deltas sitting inside the noise floor. If you're comparing two fine-tunes and one wins by 0.02, you haven't measured anything. The fix in the tooling is to re-run the base model N times and refuse to call any delta smaller than the measured spread significant.
Eval gates can be wrong in ways that look like findings. I ran a suite that scored a model 0.225 when it got 40/40 right. The scorer was ranking brace hygiene: the model emitted one closing brace short, the parse fell back to the inner object, and the scorer rejected it for lacking the outer key. Another suite scored Llama-3.1-8B at 0.423, below a 0.5B model, because the extractor didn't know \boxed{C} and the prompt never asked for a letter. After the fix, the same model scored 0.731.
Version bugs bite silently. If you trained an adapter with layer streaming on v0.72.0, that adapter is inert: its tensors were saved under keys with an extra .inner. segment, so every loader returned the untuned base. Nothing complains. You just get the base model's outputs.
The distribution side is moving too. A MiniMax-H3-Turbo-Lora space sits at 139 likes, and people are assembling multi-teacher distillation datasets with quality filtering baked in. The qwen3.8-max-glm5.2-kimi-k3-distillation dataset mixes traces from three frontier teachers with verifier_passed, family_oracle_passed, and reference_agreement flags, plus a sampling weight per row. The filtering happens before training, not after. Same structural instinct: don't pay to learn from bad examples.
Common Pitfalls
Don't use the dense KL implementation for long-context distillation. It fails outright from 64K tokens, and at 32K it uses 15.6× more memory than the fused chunked loss. The chunked version is mathematically equivalent; the only cost is a second output projection pass.
Don't trust a single eval run when comparing fine-tunes. Greedy decoding on GPU spreads 0.015 to 0.020 run to run. Measure the noise floor first, then call a delta significant.
Don't assume pre-norm is always the right placement. Under curriculum depth growing, post-norm wins by 0.0328 validation CE, an order of magnitude larger than the joint-training gap. If you append layers over time, test post-norm.
Don't treat MoT as a free lunch for pretraining. The parity schedule processes more aggregate tokens, and the compute advantage only materializes when the aligner is reused across runs. It's a research framework for reusable training units, not a drop-in replacement for monolithic pretraining.
Don't skip adapter key checks after upgrading training tools. The .inner. segment bug meant every loader returned the untuned base model, and nothing complained.
One thing to remember: every win in this cluster came from not materializing something that didn't need to exist. Cache the teacher. Chunk the loss. Stream the layers. Reuse the aligner. The math was never the bottleneck. Memory and redundant recomputation were.
The Bottom Line
If you're distilling a model above ~20B parameters at long context, adopt the fused chunked KL loss with cached top-100 logits. It took a GPT-OSS 20B run from four GPU nodes to one and cut step time 5×, with loss curves indistinguishable from online distillation.
If you're constrained to a single consumer GPU, skip full fine-tuning and use layer streaming with QLoRA. An 8B model trains at 119.6 tok/s in 3.32GB on a 4GB laptop card, bit-exact against a resident run, and DPO costs only 0.914× the SFT peak.
Watch modular pretraining: it's at proof-of-mechanism scale, and its economics hinge on aligner reuse. If aligners become reusable training units, expect pretraining experiments to get dramatically cheaper within a year.