Quantifying the Memorization-to-Generalization Transition: Scaling Laws and Phase Structure in Grokking
arXiv:2609.10657v1 Announce Type: new Abstract: Neural networks trained past memorization frequently undergo a delayed transition to generalization, a phenomenon known as grokking. Despite theoretical progress on \emph{why} this transition occurs, the quantitative structure of \emph{when} it occurs in hyperparameter space remains uncharacterized. We map the memorization-to-generalization boundary across 384 configurations of two-hidden-layer MLPs on modular arithmetic, fitting a power-law scalin
![Training a 210M text-to-image DiT from scratch on one GPU: what I measured [P]](https://preview.redd.it/7vrhqzf32woh1.png?width=140&height=71&auto=webp&s=9f2d1cb4a1a423d2fc5a5f1c55ef4aa6774a23b7)





