<rss xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title>OLMo 3 - Tag - lilfry's library</title><link>https://lilfry09.github.io/en/tags/olmo-3/</link><description>OLMo 3 - Tag - lilfry's library</description><generator>Hugo -- gohugo.io</generator><language>en</language><managingEditor>lilfry@sjtu.edu.cn (lilfry)</managingEditor><webMaster>lilfry@sjtu.edu.cn (lilfry)</webMaster><copyright>All rights reserved.</copyright><lastBuildDate>Mon, 06 Apr 2026 19:00:00 +0800</lastBuildDate><atom:link href="https://lilfry09.github.io/en/tags/olmo-3/" rel="self" type="application/rss+xml"/><item><title>Read OLMo 3: What really matters is not the new architecture, but the training signal</title><link>https://lilfry09.github.io/en/ai/olmo3-training-signals/</link><pubDate>Mon, 06 Apr 2026 19:00:00 +0800</pubDate><author>fry</author><guid>https://lilfry09.github.io/en/ai/olmo3-training-signals/</guid><description><![CDATA[<p>When I read <code>OLMo 3</code> recently, my biggest feeling was not that &ldquo;the open source community has made a stronger model&rdquo;, but that it made something clear that is often said very vaguely:</p>
<p>**What a large model will eventually look like is determined not by a few architectural fine-tunings, but by what training signals it repeatedly receives at each stage. **</p>
<p>This statement may sound naive, but if you really think about it, many questions about LLM will become clearer. Why do some models have good bases but cannot be pulled up after training? Why do some model context windows become longer, but they still don&rsquo;t really take advantage of the long context? Why do some RL recipes seem advanced but have no obvious benefits in the end? The answer is often not &ldquo;what has changed in the structure&rdquo;, but &ldquo;where does the gradient come from?&rdquo;</p>]]></description></item></channel></rss>