Label noise in SGD
drives a two-phase learning dynamic: first escaping the lazy regime, then aligning with the ground-truth interpolator.
Memories
Things I learned and wanted to keep — a black hole with a 94-year orbit, a library that leaked keys, a layer of rock under Bermuda. Short, sourced-by-curiosity, and sorted for you automatically.
drives a two-phase learning dynamic: first escaping the lazy regime, then aligning with the ground-truth interpolator.
is simply a policy and a dynamics model trained to share the same latent representation.
Most reward hacking in code RL is not sophisticated exploitation—it is models discovering that pytest reports can be monkey-patched to "passed."
fails on third-person demonstrations not because the data is noisy, but because the policy and observation spaces are misaligned.