<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
        <title>Scaling to Long Horizons — Ross Taylor &amp; Chengxi Taylor, General Reasoning</title>
        <link>https://video.ut0pia.org/videos/watch/f772606f-cfc4-4191-b922-94e531ea501f</link>
        <description>Ross Taylor opens with some history: back in 2022 he worked on Galactica, an early large model for science that briefly crossed the Rubicon on curated high quality data and intermediate reasoning tokens before the reaction overshadowed the work. That obsession, optimizing what happens between the question and the answer, is where this talk on long horizon reinforcement learning picks up. He and Chengxi Taylor of General Reasoning treat long horizon less as a benchmark and more as a mindset: if you want agents that stay coherent over hours, you have to be patient about signal and deliberate about how you spend tokens. The mechanics they walk through are the ones that make long rollouts trainable. Value models reduce variance and help with credit assignment, bootstrapping pulls signal out of sparse rewards, and the real constraint becomes the tradeoff between off policy staleness and GPU utilization as sequences get longer. They make it concrete with a task where frontier models were handed real money to trade football matches and did poorly, exposing how little the environment was actually simulated. The takeaway is that scaling to long horizons demands better environments and simulation, not just bigger context windows, and they point listeners to openreward.ai to go deeper. Speaker info: Ross Taylor (General Reasoning): https://x.com/rosstaylor90, https://www.linkedin.com/in/rosstaylor90/, https://rossjtaylor.com, Chengxi Taylor (General Reasoning): https://x.com/chengxitaylor, https://www.linkedin.com/in/chengxi-taylor/, https://www.chengxitaylor.com/, Timestamps: 0:00 - Introduction and a look back 1:57 - The Galactica story 5:15 - Curated data and thinking tokens 8:09 - What got RL cooking 9:12 - Long horizon as a mindset 10:16 - Why value models help 11:08 - Credit assignment and bootstrapping 12:38 - Trading football matches for real money 13:44 - Why models struggled 14:36 - Off policy staleness versus GPU use 16:18 - openreward.ai and what's next</description>
        <lastBuildDate>Sat, 01 Aug 2026 18:04:06 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>PeerTube - https://video.ut0pia.org</generator>
        <image>
            <title>Scaling to Long Horizons — Ross Taylor &amp; Chengxi Taylor, General Reasoning</title>
            <url>https://video.ut0pia.org/lazy-static/avatars/0287a09a-aae7-4840-9843-b416426e7046.webp</url>
            <link>https://video.ut0pia.org/videos/watch/f772606f-cfc4-4191-b922-94e531ea501f</link>
        </image>
        <copyright>All rights reserved, unless otherwise specified in the terms specified at https://video.ut0pia.org/about and potential licenses granted by each content's rightholder.</copyright>
        <atom:link href="https://video.ut0pia.org/feeds/video-comments.xml?videoId=f772606f-cfc4-4191-b922-94e531ea501f" rel="self" type="application/rss+xml"/>
    </channel>
</rss>