Overview of 20VC with Lin Qiao, Founder and CEO of Fireworks
Harry Stebbings interviews Lin Qiao, founder and CEO of Fireworks, about the future of AI infrastructure, open-source models, inference economics, and why the market will likely evolve toward millions of specialized models rather than a single dominant AGI system. Lin argues that enterprises will increasingly want to own and customize their intelligence, especially as token costs fall, usage explodes, and AI becomes embedded in production workflows.
Core Thesis: The Future Is Specialized Intelligence
Lin’s central argument is that AI will not converge into one universal model that solves everything.
- He rejects the idea that “one company owns intelligence.”
- He believes every company has:
- unique data
- unique workflows
- unique taste/judgment
- unique user intent
- Therefore, the future is specialized intelligence, not a single AGI layer.
He repeatedly frames frontier-model providers as the “power lines” of AI: essential infrastructure, but not the whole economy.
Open Source vs Closed Models
A major theme is the rapid rise of open-source models and what that means for the AI stack.
Why Fireworks bet on open models
- Fireworks initially debated whether to build its own models or build on open models.
- Lin says open models won because they:
- give users control over weights
- are easier to customize
- make it possible to tune models to proprietary enterprise data
- Fireworks itself uses open models internally for:
- recruiting
- finance workflows
- debugging/coding
- agentic internal processes
Why this matters
- Many enterprise tasks can now be solved well with tuned open models.
- That reduces dependence on expensive frontier APIs for a large share of workloads.
- Lin thinks this is one reason the market is moving toward customized deployment, not just raw model access.
Token Costs, Usage Growth, and the Economics of AI
Lin is very bullish on falling costs and explosive usage growth.
His core forecast
- Token costs could fall 10x in three years
- That cost drop could drive 100x usage growth
Why
- Supply-chain constraints are still keeping prices high:
- GPUs
- memory/HBM
- energy
- data center capacity
- But over time, competition and scale should compress costs.
- As prices drop, more use cases become economically viable.
Important nuance
He stresses that token economics should be evaluated by task, not just raw token price:
- some models are cheaper but more verbose
- some models use fewer tokens and are more precise
- real efficiency comes from:
- tuning the model
- optimizing inference
- matching the model to the workload
Fireworks’ Position in the Stack
Fireworks sits between:
- chip providers / infrastructure
- model providers
- enterprise applications
Lin says Fireworks is not moving into the application layer.
What Fireworks wants to be
- A specialized intelligence platform
- A company that helps customers:
- customize models
- tune them to their proprietary data
- deploy them efficiently
- route tasks intelligently across models
What Fireworks is not trying to do
- Become an application company
- Own the entire stack today
- Over-optimize for margin too early
He says the company may one day invest in data centers if it makes strategic sense, but only when the economics justify it.
Enterprise AI: Why “Build Your Own Intelligence” Matters
Lin argues enterprise adoption is shifting from experimentation to optimization.
The shift
- In the SaaS era, product-market fit was enough to scale.
- In AI, product-market fit and durable business are separate.
- Companies can find demand, but still be unable to scale because AI costs can break unit economics.
The implication
Enterprises increasingly need:
- control over weights
- custom tuning
- guardrails
- workflow-specific orchestration
- multi-model routing
He believes that in the future, every company will own its own intelligence stack, similar to how companies once owned their software stack.
Multi-Model Routing and Automation
Lin sees the “frontier” increasingly shifting from a single model to a routing system.
How it works
- High-complexity tasks may go to expensive frontier models
- Smaller tasks can go to cheaper open models
- Special workflows can be tuned for a company’s own data and evals
He says Fireworks is working toward:
- automatic routing
- automatic tuning
- self-evolving systems that improve based on product usage
In his view, the routing layer itself can be valuable, especially for enterprise workflows that require performance, cost control, and customization.
Why Specialized Models Win in Enterprise
Harry raises examples like legal AI companies and whether they should build their own models.
Lin’s view:
- legal is especially sensitive because hallucinations are costly
- every company has proprietary knowledge and workflow structure
- model choice alone is not enough; the orchestration layer matters too
He believes companies like legal-tech or coding copilots win by combining:
- proprietary data
- workflow-specific tooling
- tuned models
- strong inference optimization
Growth, Customers, and Market Momentum
Fireworks is growing very quickly, and Lin frames that growth as a function of broader AI adoption.
Key points
- Fireworks processes 40 trillion+ tokens per day
- He says most of that is on customized models, not off-the-shelf ones
- He believes token volume could grow 20x to 100x by next year
- Fireworks’ ARR is already at a very large scale, and he expects it to at least double by year-end
Customer mix
- Coding was the big category last year
- “Co-work” is the big category this year
- He sees growth across:
- legal
- finance
- customer support
- recruiting
- sales
- marketing
- healthcare
- consumer recommendation systems
Data Centers, Chips, and the AI Supply Chain
The conversation also goes deep into infrastructure.
Lin’s view on infrastructure
- AI is bottlenecked by the lower layers of the stack:
- chips
- memory
- energy
- data centers
- Specialized data center design matters a lot
- He sees possible future expansion into data centers, but not chips yet
Why not chips?
- Chip development is highly specialized and slow to change
- Workloads are still too dynamic for many companies to justify custom silicon
- Once a company’s workload stabilizes, custom chips make more sense
On data centers
He says data centers are not commoditized:
- power
- cooling
- layout
- network
- deployment architecture
All of it can materially affect performance and economics.
Leadership, Hiring, and What Lin Learned
Lin shares several lessons from building Fireworks and working at Meta and LinkedIn.
Hiring philosophy
He looks for people with:
- extreme ownership
- high velocity
- curiosity
- low ego
- strong accountability
He says Fireworks’ best people are often immigrants and operators who thrive in ambiguity and take full responsibility.
What he learned from Jensen Huang
- Jensen is deeply involved in details
- Great leadership means having the right context to make good judgments
- In a fast-moving market, you cannot rely on slow information transfer
What he would change
- He says he waited too long to invest in marketing
- He now believes marketing is about:
- education
- clarity
- shaping customer understanding of the market
Notable Takeaways
- The AI future is likely many specialized models, not one AGI monopoly.
- Open-source models are becoming good enough to power a large amount of enterprise work.
- Customization and routing are becoming as important as raw model capability.
- Token costs should fall significantly, which will unlock far more usage.
- AI infrastructure remains constrained by chips, memory, energy, and data center capacity.
- Enterprises increasingly need to own their intelligence, not rent it entirely.
- Fireworks is positioning itself as a specialized intelligence platform, not a general application company.
Bottom Line
Lin Qiao’s thesis is that AI is moving from a “model race” to a systems and specialization race. The winners will not simply be the companies with the biggest base models, but the ones that can best combine open models, customization, routing, and efficient inference into production-grade intelligence for each unique workload.
