Etched - Building AI Hardware to Make Inference Faster and Cheaper - [Invest Like the Best, EP.480]

Summary of Etched - Building AI Hardware to Make Inference Faster and Cheaper - [Invest Like the Best, EP.480]

by Colossus | Investing & Business Podcasts

1h 27mJune 30, 2026

Overview of Etched – Building AI Hardware to Make Inference Faster and Cheaper

This episode is a deep dive into Etched, the AI hardware startup building a full-stack inference system—chip, rack, power delivery, interconnects, and manufacturing—designed to make AI model serving dramatically faster and cheaper. Founders Gavin Uberti and Rob explain why they believe modern AI inference is a fundamentally new hardware problem, how they designed around the specific needs of prefill and decode, and why they think inference will become one of the largest markets in the world. The conversation also covers their unlikely founding story, intense execution culture, fundraising struggle, and the technical challenges they had to solve to get a working chip taped out on the first try.

Why Etched Exists

  • Etched’s core thesis is that AI inference is becoming the bottleneck in the AI economy.
  • The company is not just building a chip; it is building a complete inference rack-scale product.
  • Their belief is that the current hardware stack was built before ChatGPT, so it is retrofitted for modern workloads rather than optimized for them.
  • They see a future where:
    • inference is a massive share of global compute,
    • token production becomes an economy-of-scale business,
    • and the winners will be companies that can produce the most tokens at the lowest cost.

The Technical Bet: How Etched’s Architecture Works

Prefill vs. Decode

  • They break inference into two major stages:
    • Prefill: ingesting context and preparing the model’s memory state.
    • Decode: generating output tokens.
  • A major system design choice is PD disaggregation:
    • one cluster handles prefill,
    • another handles decode,
    • and the KV cache is transferred between them.

Low-Voltage Inference

  • One of Etched’s key insights is that GPU-style voltage assumptions are too high for inference-optimized hardware.
  • They developed a mechanism for running at much lower voltage than typical AI chips.
  • The goal is to:
    • reduce thermal throttling,
    • increase real-world flop utilization,
    • and improve performance per watt.

Cluster-Scale Memory

  • For decode, memory bandwidth and latency matter more than raw compute alone.
  • Etched is building a custom interconnect stack so a cluster of chips can behave more like a shared memory pool.
  • They argue that the important metric is not memory bandwidth on one chip, but memory bandwidth across the entire scale-up cluster.
  • Their custom networking/interconnect approach aims to reduce chip-to-chip latency by more than 5x versus current systems.

Why They Rejected Generic Compiler-First Approaches

  • Early on, they explored many architectures:
    • advanced packaging,
    • HBM-heavy designs,
    • 3D DRAM,
    • shared memory pools,
    • graph compilers.
  • They ultimately chose a kernels-first strategy:
    • optimize for a small number of important models,
    • don’t support every possible framework or graph type,
    • and let the most advanced customers work close to the metal.
  • Their view: the best AI hardware should be built for the models that matter, not for arbitrary generality.

Building the Company: Speed, Vertical Integration, and Talent

“Velocity, velocity, velocity”

  • Etched’s culture is centered on shipping fast.
  • They believe the best way to win is to:
    • move quickly,
    • make decisions fast,
    • and parallelize aggressively.

Extreme Vertical Integration

  • Their philosophy is essentially: “the best vendor is no vendor.”
  • They try to keep as much as possible in-house:
    • chips,
    • boards,
    • cold plates,
    • interconnects,
    • and even production coordination.
  • They are, unusually, building both the chip and the rack as a startup.

Hiring: “Legends + Naivete”

  • Their talent strategy is intentionally bimodal:
    • legends: world-class experts with deep experience,
    • young, first-principles talent: ambitious people willing to push hard on assumptions.
  • They built a project-based recruiting process to find the best person in the world for each technical problem.
  • A standout example is hiring NVIDIA rack-scale veteran Brian to lead systems work alongside very young builders.

Intense Execution

  • They routinely do things that compress timelines dramatically:
    • sent engineers to Bangalore for months to unblock vendor issues,
    • ran a 24/7 development cycle with U.S. and India handoffs,
    • built and tested racks before chips returned,
    • and used thermal proxy chips to validate cold plates early.
  • They emphasize that speed comes from parallelization and willingness to spend money when ROI is clear.

The Hardest Moments

Raising Money

  • Early fundraising was extremely difficult.
  • Investors often passed because:
    • the founders were very young,
    • they had no taped-out product yet,
    • and many thought inference was not yet the main market.
  • They wrote a dense, technical memo and were initially rejected by many major investors.
  • Eventually, they pieced together a large round from believers and strategic partners.

A Near-Shipwreck Hardware Bug

  • One of the most intense technical crises involved a clock-domain crossing / backpressure issue that caused incorrect results.
  • The fix required aligning two clock signals to within 50 picoseconds.
  • They solved it in about two weeks, after people had said it was impossible.

First Wafer Sort

  • During early wafer testing, every chip initially showed red instead of green.
  • The team stayed calm, treated it as a solvable puzzle, and found a fix quickly.

Why They Think Inference Will Matter So Much

  • They believe inference will shape:
    • software costs,
    • AI product economics,
    • productivity,
    • and eventually broader GDP.
  • Their argument is that cheaper inference enables:
    • longer-running agents,
    • bigger context windows,
    • more concurrent users,
    • and faster scientific and software breakthroughs.
  • They expect AI systems to support:
    • massive scale-up clusters,
    • distributed “brains” made of many experts,
    • and eventually data centers that look more like token factories than traditional compute farms.

Outlook on the Future of AI Hardware and Models

  • They think future models will likely use:
    • more compute,
    • more dynamic routing,
    • more experts,
    • and larger distributed memory systems.
  • They believe hardware will increasingly shape model design:
    • models will be built to exploit large-scale interconnects and memory pools,
    • and chips will need to support highly dynamic workloads.
  • Their long-term view is that:
    • more tokens will be produced than ever before,
    • and the economy will increasingly be organized around the production of intelligence.

Key Takeaways

  • Etched is building full-stack inference infrastructure, not just a chip.
  • Their technical edge comes from:
    • low-voltage inference,
    • cluster-scale memory,
    • and custom interconnects.
  • They reject generic “one-size-fits-all” chip design in favor of workload-specific optimization.
  • Their operating model is built on:
    • speed,
    • vertical integration,
    • and recruiting exceptional people who believe the mission is possible.
  • The founders believe inference is on track to become one of the largest markets in the world.

Notable Ideas

  • “The best vendor is no vendor.”
  • “Production is the product.”
  • “Speed wins.”
  • “Assume it is possible.”
  • “The biggest risk is not taking risk.”

Closing Note

The episode is part technical deep dive, part founder origin story, and part manifesto for a future where AI inference becomes a core industrial layer of the economy. Etched’s bet is that the next era of AI won’t just be about better models—it will be about building the hardware and systems that can serve those models at massive scale, cheaply and fast.