<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
        <title>Reinforcement Learning without Verifiable Rewards — Will Brown, Prime Intellect</title>
        <link>https://video.ut0pia.org/videos/watch/dc2ce06a-cd45-4fea-810f-b37b917e4878</link>
        <description>Reinforcement learning has been easy to sell where the answer is checkable, like math or code, and Will Brown's talk is about everything else. Most valuable tasks have no clean verifier, so Prime Intellect's work is on how you build reward signal when there is no ground truth waiting. He frames RL simply first, a model acting in a harness with tools and skills, getting a reward, and nudging its weights, then asks how you keep climbing once you leave the verifiable island behind. His answer leans on environments as the anchor. You can set up judges, generate question and answer pairs grounded in real documents and repos, and use a reverse direction trick where you hide something, like a bug or a backdoor, so the model can learn to find it again, which conveniently gives you a difficulty dial to keep tasks not too easy and not too hard. He is direct about the dangers: reward hacking will find you if you are not careful, so you inspect traces, run small experiments, and bring in expert understanding. The goal he keeps returning to is making this a real science, with open models and shared benchmarks, where environments turn into new tasks and higher levels of ability. Speaker info: https://x.com/willccbb, https://www.linkedin.com/in/willcb/, https://willcb.com, Timestamps: 0:00 - RL without verifiable rewards 1:17 - How RL works, simply 2:43 - The tooling that powers it 4:13 - Where verifiable rewards run out 6:24 - Being careful about reward design 8:04 - Making RL a science 9:20 - Judges and grounded question answer pairs 10:46 - The reverse direction trick 14:19 - Calibrating difficulty 15:08 - Hunting for reward hacks 18:27 - Environments as the anchor</description>
        <lastBuildDate>Sat, 01 Aug 2026 17:51:03 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>PeerTube - https://video.ut0pia.org</generator>
        <image>
            <title>Reinforcement Learning without Verifiable Rewards — Will Brown, Prime Intellect</title>
            <url>https://video.ut0pia.org/lazy-static/avatars/0287a09a-aae7-4840-9843-b416426e7046.webp</url>
            <link>https://video.ut0pia.org/videos/watch/dc2ce06a-cd45-4fea-810f-b37b917e4878</link>
        </image>
        <copyright>All rights reserved, unless otherwise specified in the terms specified at https://video.ut0pia.org/about and potential licenses granted by each content's rightholder.</copyright>
        <atom:link href="https://video.ut0pia.org/feeds/video-comments.xml?videoId=dc2ce06a-cd45-4fea-810f-b37b917e4878" rel="self" type="application/rss+xml"/>
    </channel>
</rss>