lilfry's library
Home All Posts AI Unfinished Drafts LeetCode Categories Tags About
lilfry's library
Cancel
HomeAll PostsAIUnfinished DraftsLeetCodeCategoriesTagsAbout

 LLM

2026

High frequency interview: Bradley-Terry vs Plackett-Luce, what is the difference between reward modeling? 04-08
Read YaRN: Long context expansion, the core is not to pull the window hard, but not to break the RoPE 04-07
Read "Reasoning with Sampling": RL does not make the model smarter, it just redistributes reasoning capabilities 04-07
Read OLMo 3: What really matters is not the new architecture, but the training signal 04-06

Hope can set you free.

Powered by Hugo
2025 - 2026 lilfry