High frequency interview: Bradley-Terry vs Plackett-Luce, what is the difference between reward modeling? 04-08
Read YaRN: Long context expansion, the core is not to pull the window hard, but not to break the RoPE 04-07
Read "Reasoning with Sampling": RL does not make the model smarter, it just redistributes reasoning capabilities 04-07