<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
        <title>What's Next After RLHF? — Diogo Almeida, TypeSafe AI</title>
        <link>https://video.ut0pia.org/videos/watch/6d9eb540-fd1f-47f2-952b-713ea58ddb68</link>
        <description>RLHF made models that are extraordinary at pleasing the human in the loop, and Diogo Almeida, a GPT-4 co author, argues that is exactly the problem. Optimizing for human preference optimizes for engagement and for overpromising, the same pressure that makes a model confidently agree that a fart audio file is a symphony. That produces two camps: one where models act as assistants with a human catching mistakes, where RLHF shines, and one where they operate autonomously with real stakes, where the same instinct to please quietly becomes a liability. So what comes next is not the Claude Code era but a shift in what you optimize. Almeida frames it through Sutton's bitter lesson: the task matters more than the data, and reinforcement learning with verifiable rewards points the model at real automation instead of human approval. He is careful that pre trained models are already incredibly capable and that the trap is bolting preference optimization on top, which teaches confidence and drops modes. The through line is that assistance and automation pull in different directions in optimization space, and the field is only starting to say plainly which one it is building. Speaker info: https://x.com/CompleteSkeptic, https://www.linkedin.com/in/diogomda/, https://typesafe.ai/, Timestamps: 0:00 - Not the Claude Code era 1:40 - The state of the field 3:14 - Two camps: assistance and autonomy 4:31 - Why models please the human in the loop 6:37 - How RLHF actually works 7:31 - Preference versus what's true 8:10 - When the consequences get real 8:47 - So what's next 9:35 - Assistance is not automation 14:31 - Is pre-training the problem? 15:43 - RLVR and Sutton's bitter lesson</description>
        <lastBuildDate>Sat, 01 Aug 2026 17:44:57 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>PeerTube - https://video.ut0pia.org</generator>
        <image>
            <title>What's Next After RLHF? — Diogo Almeida, TypeSafe AI</title>
            <url>https://video.ut0pia.org/lazy-static/avatars/0287a09a-aae7-4840-9843-b416426e7046.webp</url>
            <link>https://video.ut0pia.org/videos/watch/6d9eb540-fd1f-47f2-952b-713ea58ddb68</link>
        </image>
        <copyright>All rights reserved, unless otherwise specified in the terms specified at https://video.ut0pia.org/about and potential licenses granted by each content's rightholder.</copyright>
        <atom:link href="https://video.ut0pia.org/feeds/video-comments.xml?videoId=6d9eb540-fd1f-47f2-952b-713ea58ddb68" rel="self" type="application/rss+xml"/>
    </channel>
</rss>