<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
        <title>Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software</title>
        <link>https://video.ut0pia.org/videos/watch/0ef375b9-360a-421b-ad64-1f1ea0c4c262</link>
        <description>Everyone wants agents that handle long horizon work, but Rayan Garg starts with the awkward question of what long horizon even means. One popular answer measures the time horizon as the task length at which an agent crosses a success threshold, like the sixteen hour mark, which is a useful endpoint but a noisy one, since human time estimates vary and the same wall clock hides very different amounts of real difficulty. How you choose to measure this has an outsized effect on what you conclude about a model. From there Theta Software's work is about designing the environments and verifiers that make those measurements honest. A task can be artificially stretched by forcing serial dependencies, or made genuinely hard when a bad early query cascades through everything after it, and as environments grow more complex, standardized evaluation gets harder and correctness is best verified from the final state rather than a judge's guess. Garg walks through collapsing a huge state space with sample trajectories, being careful that judges do not see information they should not, and reusing agents to sift artifacts like CI logs. The recurring principle is that long horizon progress lives or dies on environment and verifier design, not on the headline benchmark number. Speaker info: https://x.com/RayanGarg, https://www.linkedin.com/in/rayan-garg/, Timestamps: 0:00 - What does long horizon mean? 1:13 - Time horizon and the threshold metric 3:17 - Why the metric is noisy 4:20 - Measuring what actually matters 6:38 - Creating tasks and environments 7:42 - When a bad early step cascades 10:01 - Why standardized evaluation is hard 11:17 - Verifying from the final state 13:46 - Judges, tools, and reused agents 17:45 - Rubrics, QA, and careful grading</description>
        <lastBuildDate>Sat, 01 Aug 2026 17:53:50 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>PeerTube - https://video.ut0pia.org</generator>
        <image>
            <title>Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software</title>
            <url>https://video.ut0pia.org/lazy-static/avatars/0287a09a-aae7-4840-9843-b416426e7046.webp</url>
            <link>https://video.ut0pia.org/videos/watch/0ef375b9-360a-421b-ad64-1f1ea0c4c262</link>
        </image>
        <copyright>All rights reserved, unless otherwise specified in the terms specified at https://video.ut0pia.org/about and potential licenses granted by each content's rightholder.</copyright>
        <atom:link href="https://video.ut0pia.org/feeds/video-comments.xml?videoId=0ef375b9-360a-421b-ad64-1f1ea0c4c262" rel="self" type="application/rss+xml"/>
    </channel>
</rss>