Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

Summary of Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

by Colossus | Investing & Business Podcasts

1h 18m•August 25, 2026

Overview of Neil Movva - Making AI 10x Cheaper

In this episode of Invest Like the Best, Patrick O’Shaughnessy speaks with Neil Movva, founder of SAIL Research, about building what he calls a “token factory”: an inference company designed to make AI dramatically cheaper, especially for long-running background agents rather than real-time chat. Movva’s core thesis is that the future of AI is less about fast responses and more about abundant, persistent intelligence running in the background for hours or days, doing deep research, cybersecurity, and proactive personal assistance. The conversation is a detailed tour through the full stack of AI economics — software, chips, memory, data centers, and power — and why he believes the biggest opportunity is not just better models, but cheaper tokens at massive scale.

Main Thesis: Intelligence Becomes a Commodity

Movva’s worldview is built around one idea: if you make intelligence 10x cheaper, you create a new product category.

What SAIL Research does

  • SAIL provides an API for serving open-source large language models at extremely low cost.
  • It also supports long-running agents through “SAIL boxes,” virtual machines built for tasks that last hours, days, or even weeks.
  • The company is optimized for background work, not interactive chat.

Why this matters

  • He believes the market is shifting from human-in-the-loop, low-latency inference to agentic, long-horizon inference.
  • In that world, cost matters more than latency.
  • His North Star is simple: lowest cost per token in the industry.

His broader vision

  • Abundant cheap tokens will let users and companies:
    • run deep research at massive scale,
    • proactively monitor systems and data,
    • automate cybersecurity discovery and patching,
    • and eventually create proactive assistants that understand your life in the background.

Why Background Agents Matter More Than Chatbots

Movva argues that the future of AI is not a human waiting for a chatbot reply, but a machine doing work before you ask.

The shift he sees

  • Today’s dominant AI use case is still interactive chat.
  • But he thinks the next wave is background and proactive agents.
  • He predicts the market will move from roughly 50/50 background vs. real-time workloads to 90/10 in favor of background.

Best use cases for long-running inference

  • Deep research
    • Synthesizing answers from 10,000+ sources.
    • Building authoritative indexes of knowledge.
  • Cybersecurity
    • Using agents to break software in many ways and then patch vulnerabilities.
    • He describes security as increasingly becoming a kind of “proof of work” spent on model-driven attacks.
  • Proactive personal assistance
    • A future Siri-like agent that watches emails, texts, calendars, and app behavior continuously.
    • The goal is an assistant with an encyclopedic view of your life.

Key insight

Movva’s favorite framing is:

“The best latency is no latency at all.”

The ideal agent does the work while you sleep, rather than while you wait.

The Full-Stack Strategy: Software, Chips, and Power

A major theme of the episode is Movva’s belief that cheap AI requires optimization across the whole stack.

Software: Squeezing More Out of Existing GPUs

Movva starts with software because it is the fastest place to find leverage.

Core software ideas

  • Optimize for peak GPU efficiency.
  • Build kernels and serving systems to maximize tokens per chip.
  • Fuse operations and reduce unnecessary memory round-trips.
  • Design for throughput, not just latency.

Why throughput matters

  • GPUs are naturally throughput machines.
  • Chatbots favor low latency, which keeps GPUs away from their ideal operating mode.
  • Background agents let you batch more work and keep the hardware busier for longer.

His takeaway

  • He is less interested in creating a single “best” model stack than in making any chip work well at the right price.

Hardware: A Heterogeneous, Arbitrage-Driven View

Movva is highly pragmatic about chips. His view is that there are no bad chips, only bad pricing.

NVIDIA

  • He respects NVIDIA deeply and learned much of his engineering ethos there.
  • He believes NVIDIA is still excellent for both:
    • low-latency inference
    • and high-throughput inference
  • But he does not think NVIDIA’s performance-per-watt improvements have been as dramatic as many assume.

His contrarian view

  • The industry’s obsession with NVIDIA can obscure the opportunity in other chips.
  • He thinks there’s meaningful alpha in using:
    • AMD,
    • TPUs,
    • Trainium,
    • Cerebras,
    • Groq,
    • and other emerging accelerators.
  • The key is finding each chip’s comparative advantage and fitting it into a heterogeneous serving system.

Memory hierarchy matters

He spends a lot of time on the trade-offs between:

  • SRAM: very fast, but low capacity and expensive in die area.
  • DRAM/HBM: much denser, but slower.

This leads to his view that:

  • Cerebras/Groq-style chips can be excellent for certain workloads,
  • but they are limited by the KV cache, which grows dynamically during inference.
  • The future is likely hybrid systems:
    • fast on-chip memory for some components,
    • larger external memory for others,
    • and different chips for different model subsystems.

Transformers, KV Cache, and the Limits of Today’s Architecture

Movva gives a useful technical framing of why current AI systems still have bottlenecks.

Transformers

  • He sees transformers as incredibly powerful because they:
    • learn sequences well,
    • scale cleanly with more compute,
    • and can model any pairwise relationship in the input.
  • He thinks transformers remain a dominant architecture because they are such strong general learners.

But there are limits

  • The attention layer is memory-bound.
  • The MLP layer is compute-bound.
  • At extreme context lengths, those two components stress different parts of the hardware stack.

KV cache

  • Movva highlights the KV cache as one of the biggest inefficiencies in modern inference.
  • It stores the working memory of the conversation and can become larger than the model weights themselves.
  • He believes there is still 1–2 orders of magnitude of potential improvement in KV cache compression and management.

Data, RL Environments, and the Next Phase of Model Improvement

Movva argues that the internet was effectively a one-time subsidy on training data.

His view on data

  • Most high-quality internet text has already been consumed many times over.
  • The next frontier is not more scraped web data.
  • Instead, it is model self-improvement through verifiable environments.

What that means

  • Give the model a hard, measurable task.
  • Let it act in an environment.
  • Score its progress.
  • Use that feedback as training data.

Examples

  • Coding tasks
  • Math problems
  • Research tasks
  • Security challenges

He sees this as a path toward recursive improvement on verifiable problems.

Data Centers and Power: The “Scavenger Strategy”

A big part of SAIL’s edge, in Movva’s view, is not just software or chips but where and how compute is sourced.

His “scavenger strategy”

  • Buy any chip, anywhere, for any duration.
  • Use power and facilities others overlook.
  • Avoid head-to-head competition with OpenAI or Anthropic for the highest-end, most obvious compute.

Data center philosophy

  • He prefers many small, distributed data centers over huge centralized ones.
  • He thinks inference can tolerate:
    • lower uptime,
    • less redundancy,
    • and simpler networking than training can.

Why this works for him

  • Background agents can tolerate occasional failure or delay.
  • If a request needs to move to another GPU, that’s acceptable if the job is running for hours anyway.
  • He is willing to trade P99 latency for dramatically better economics.

Power sourcing

  • He is enthusiastic about solar and wind, especially if paired with flexibility in workload scheduling.
  • Because inference is more tolerant of interruptions, he thinks variable power sources can become more viable.

Market Views: Semis, NVIDIA, and the AI Boom

Movva is bullish on AI demand, but he is skeptical of some of the market’s long-term assumptions.

Why this cycle is different from past bubbles

  • In his view, today’s inference spend is not speculative.
  • Tokens are consumed because they are immediately valuable.
  • That differs from past cycles where infrastructure got built for demand that never materialized.

NVIDIA and the semiconductor market

  • He remains bullish on NVIDIA in the short term.
  • But he argues that the market overstates the performance-per-watt gap between leading chips and alternatives.
  • He also thinks geopolitical concerns about fabs are often exaggerated.

Memory and bottlenecks

  • He sees HBM and memory capacity as one of the biggest bottlenecks in the AI supply chain.
  • He expects memory shortages to persist and influence pricing across the ecosystem.

Hiring and Culture: What He Looks For

Movva’s team philosophy matches the rest of his worldview: performance, curiosity, and systems thinking.

Traits he values

  • Curiosity
  • Love of performance engineering
  • A desire to understand the machine at the microsecond level
  • People who are both good students and good teachers

Less important than people assume

  • Deep CUDA experience
  • Conventional “AI experience”
  • A long résumé in one narrow subfield

He wants people who care about the full stack and are willing to keep learning as the hardware and models change.

Notable Takeaways

  • Cheap tokens are the unlock for abundance in AI.
  • The future is likely background agents, not just chatbots.
  • Latency and throughput are in tension; background workloads let you optimize for throughput.
  • KV cache is one of the biggest technical inefficiencies in modern inference.
  • The most interesting systems will be heterogeneous, combining different chips and memory types.
  • The best data may come less from the internet and more from verifiable environments.
  • Movva’s long-term goal is not just to serve AI, but to make intelligence cheap enough that everyone can own it.

Bottom Line

Neil Movva’s vision for SAIL Research is to build the infrastructure for a world where intelligence is abundant, proactive, and cheap. Rather than competing to deliver the fastest chatbot answer, he is trying to create the best economic engine for long-running AI work. The entire company is built around one belief: if you can make tokens cheap enough, you can unlock a new era of software, research, security, and personal assistance.