<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
        <title>The Inference Inflection from First Principles — swyx &amp; Rob Wachen, Etched</title>
        <link>https://video.ut0pia.org/videos/watch/6cf453b7-c56d-4b67-9225-0b2f5198f24b</link>
        <description>Recorded live at the AI Engineer World's Fair, swyx sits down with Rob Wachen, co founder and president of Etched, to trace the inference inflection from first principles. Etched's bet is narrow on purpose: rather than build another general GPU designed for the lowest common denominator, they burn transformers into custom silicon and design the whole system around running today's large mixture of experts models at very high throughput. Wachen frames the company around a chunk of the world's inference capacity, built in house, where the chip is only the starting point and everything from the board to the rack has to be reinvented to actually serve tokens at scale. The conversation goes deep on the engineering reality behind that pitch. Power, not raw flops, is the wall you hit, so the team obsesses over how many cores they can light up, clock speed, voltage, and advanced packaging, alongside yield, testing, and mean time between failures. Wachen talks through a culture where production is the product, a mini data center built in their own office with a chiller piped up to the roof, a Taiwan team assembling racks ahead of time, and a memory roadmap reaching toward wafer scale and ten thousand plus chips in a single scale up. He is candid about staying cagey on the parts that are still trade secrets, and closes with a straight call to action: they are hiring across production, supply, and infrastructure for people who want to build something that has never been shown before. Speaker info: swyx (host): https://x.com/swyx, https://www.latent.space, Rob Wachen (Etched): https://x.com/robertwachen, https://www.etched.com, Timestamps: 0:00 - A special live edition with Rob Wachen 2:30 - The hardware powering AI models 3:47 - Building inference capacity in house 5:52 - A codesigned silicon team 9:21 - The transformer north star 11:57 - Recruiting for an extraordinary mission 14:33 - High throughput inference and agents 18:34 - Why power is the real wall 22:45 - Packaging, yield, and reliability 23:45 - Production is the product 24:10 - A data center built in the office 26:31 - The Taiwan team and prebuilt racks 29:04 - Memory, wafer scale, and scaling up 33:41 - A four trillion dollar problem 38:17 - Previewing what's never been shown 42:09 - A call to action: they're hiring</description>
        <lastBuildDate>Fri, 31 Jul 2026 16:43:07 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>PeerTube - https://video.ut0pia.org</generator>
        <image>
            <title>The Inference Inflection from First Principles — swyx &amp; Rob Wachen, Etched</title>
            <url>https://video.ut0pia.org/lazy-static/avatars/0287a09a-aae7-4840-9843-b416426e7046.webp</url>
            <link>https://video.ut0pia.org/videos/watch/6cf453b7-c56d-4b67-9225-0b2f5198f24b</link>
        </image>
        <copyright>All rights reserved, unless otherwise specified in the terms specified at https://video.ut0pia.org/about and potential licenses granted by each content's rightholder.</copyright>
        <atom:link href="https://video.ut0pia.org/feeds/video-comments.xml?videoId=6cf453b7-c56d-4b67-9225-0b2f5198f24b" rel="self" type="application/rss+xml"/>
    </channel>
</rss>