<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
        <title>Why Large? Tiny LMs &amp; Agents on Edge/Robotics — Cormac Brick, Google</title>
        <link>https://video.ut0pia.org/videos/watch/68ec35d3-7724-41e1-99ef-c1148b3aab49</link>
        <description>The constraint on edge AI is not compute, it is RAM, and it is getting worse: phone makers are shipping less of it this year, and a 6GB Raspberry Pi costs 2.5 times what it did at launch. So Cormac Brick's team at Google AI Edge spends its effort making models small enough to fit. A 2 billion parameter Gemma, quantized to 2.9 bits per weight, runs on a Raspberry Pi at about 8 tokens per second and on a Qualcomm NPU fast enough for a few frames of vision a second. Below that sit tiny models, from 500 million parameters down to 50, that reach the older laptops and cheap devices where even a small model will not fit. They usually need fine tuning rather than prompting, but the payoff is real: a fine tuned Gemma turns free text into the right function call across ten actions at over 86% reliability, and putting a speech model in front gives you voice to function calling. One shipped example is an offline voice dictation app with no subscription, built on two sub billion Gemma models that also strip your ums and ahs. Speaker info: https://x.com/cormacb, https://www.linkedin.com/in/cbrick/, https://github.com/google-ai-edge/gallery, Timestamps: 0:00 - Why intelligence at scale needs tiny models 1:17 - The Google AI Edge team and its open source stack 2:35 - Why run on the edge at all 3:25 - The real constraint: DRAM cost 4:40 - Small models: 1 to 4 billion parameters 6:08 - Shrinking Gemma to 2.9 bits per weight 7:36 - Decode speeds across Raspberry Pi, Jetson, and NPUs 9:30 - Try it yourself: AI Edge Gallery and a hobby robot 12:07 - When small is still too big: tiny models 13:24 - Off the shelf tiny models: ASR, vision, embeddings 14:28 - Fine tuning for voice to function calling 17:50 - In production: offline voice dictation 19:30 - Takeaways and Q&amp;A</description>
        <lastBuildDate>Sun, 26 Jul 2026 23:12:51 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>PeerTube - https://video.ut0pia.org</generator>
        <image>
            <title>Why Large? Tiny LMs &amp; Agents on Edge/Robotics — Cormac Brick, Google</title>
            <url>https://video.ut0pia.org/lazy-static/avatars/0287a09a-aae7-4840-9843-b416426e7046.webp</url>
            <link>https://video.ut0pia.org/videos/watch/68ec35d3-7724-41e1-99ef-c1148b3aab49</link>
        </image>
        <copyright>All rights reserved, unless otherwise specified in the terms specified at https://video.ut0pia.org/about and potential licenses granted by each content's rightholder.</copyright>
        <atom:link href="https://video.ut0pia.org/feeds/video-comments.xml?videoId=68ec35d3-7724-41e1-99ef-c1148b3aab49" rel="self" type="application/rss+xml"/>
    </channel>
</rss>