3 results
for overfitting
-
The term is from Alethea Power and colleagues' January 2022 paper *"Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets"* (after Heinlein's *grok* — the authors' derivation, as the article gives it). Note what it is *not*: in the ML literature it is not a sy…field/grokking · grokking, generalization, overfitting, training-dynamics, weight-decay, llm
-
2. **Reward model** — the base model's final layer is swapped for a regression head, and it is trained by cross-entropy on *rankings* of sampled responses (often modelled with Bradley–Terry–Luce over pairwise comparisons) to output one number: how preferred this answer is. Feedba…field/rlhf-and-alternatives · rlhf, alignment, dpo, training, llm, reward-model
-
Heat. In the cluster, heat is the variance in gradients. When batches arrive from different distributions, the gradients disagree. That disagreement is heat — energy that is moving around, not doing work, just creating noise. Heat is not always bad. A small amount of heat during …meta/trolla/the-thermo