<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
        <title>Build Evals That Actually Matter - Nick Ung, Lyft</title>
        <link>https://video.ut0pia.org/videos/watch/91f39714-b72b-45ca-8200-cf818de84c53</link>
        <description>Your agent passes offline evals at 90%. You ship. Production immediately finds failure modes your eval never saw. Sound familiar? The culprit is almost always the same: the "customer" in your offline eval is an off-the-shelf LLM that sounds nothing like your real users, and your synthetic test set doesn't capture how messy, angry, or off-topic real conversations get. Your eval was too easy. At Lyft, our customer-care agents resolve roughly a third of all customer issues — millions of conversations a month. To trust them at that scale, we built an adversarial user simulator: a fine-tuned LLM trained on real Lyft rider and driver transcripts that can role-play frustrated, confused, and adversarial users with the same distribution as production. It found regressions our synthetic dataset missed for months. This talk walks through the full eval lifecycle that surrounds it: the harness primitives that let any engineer write a benchmark in 20 lines, how we calibrate LLM-judge rubrics against human labels until they match inter-rater agreement, how we route failed production traces back into the offline test set, and the continual-learning loop that feeds improvements into prompts, harness, and the model. Speakers: Nick Ung (Lyft): Nick Ung leads Data Science for Safety &amp; Customer Care at Lyft, where his team built and operates the multi-agent platform that powers AI agents resolving roughly a third of all Lyft customer issues. LinkedIn: https://www.linkedin.com/in/unglikteng</description>
        <lastBuildDate>Tue, 21 Jul 2026 10:06:41 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>PeerTube - https://video.ut0pia.org</generator>
        <image>
            <title>Build Evals That Actually Matter - Nick Ung, Lyft</title>
            <url>https://video.ut0pia.org/lazy-static/avatars/0287a09a-aae7-4840-9843-b416426e7046.webp</url>
            <link>https://video.ut0pia.org/videos/watch/91f39714-b72b-45ca-8200-cf818de84c53</link>
        </image>
        <copyright>All rights reserved, unless otherwise specified in the terms specified at https://video.ut0pia.org/about and potential licenses granted by each content's rightholder.</copyright>
        <atom:link href="https://video.ut0pia.org/feeds/video-comments.xml?videoId=91f39714-b72b-45ca-8200-cf818de84c53" rel="self" type="application/rss+xml"/>
    </channel>
</rss>