Overview of 20VC: “Anti-Data Centres is a Chinese Psyop”
Harry Stebbings interviews Thomas Sohmers, co-founder and chairman of Positron AI, about the infrastructure layer behind AI inference, the economics of tokens, the bottlenecks in memory and energy, and the geopolitics of data centers and frontier-model “pacing.” Sohmers is bullish on AI’s trajectory, skeptical of broad anti-data-center sentiment, and argues that the real constraint on AI is increasingly memory, context handling, and economics, not just raw compute.
Who Thomas Sohmers Is and What Positron Does
- Positron AI builds hardware for generative AI inference:
- chips
- software close to the hardware
- full systems and rack-scale deployments
- The company focuses on the infrastructure needed to serve models like ChatGPT and Claude efficiently.
- The conversation frames Positron as part of the broader semiconductor shift back into the center of AI innovation.
Training vs. Inference: Why Inference Is a Memory Problem
Training is compute-bound
- Training benefits from:
- massive parallelism
- more FLOPs
- large datasets already available upfront
- Sohmers says scaling laws have shown that more compute and more parameters generally improve model quality.
Inference is memory-bound
- Inference is different because the model generates tokens autoregressively, one at a time.
- Each token requires repeated access to model weights and prior context.
- That makes inference heavily dependent on:
- memory bandwidth
- latency
- efficient storage of context
- Sohmers argues that as models and usage grow, memory becomes a larger bottleneck than pure compute.
The “Memory Wall” and Why It Matters
- He explains that GPU compute has improved dramatically faster than memory bandwidth.
- Roughly:
- GPU FLOPs improved ~120x
- memory bandwidth improved ~17x
- This divergence has made modern AI systems increasingly constrained by memory access rather than math.
- He ties this to the long-standing “memory wall” in computing, made more acute by transformer-based workloads.
KV Caching, Context, and Compression
What KV caching is
- KV cache stores intermediate attention data so the model doesn’t recompute prior context every time it generates a new token.
- This saves a huge amount of compute, especially in long conversations and agentic workflows.
Why it matters
- KV cache is especially important for:
- long-context models
- coding agents
- multi-turn workflows
- Sohmers cites agentic coding traces where around 96% of tokens may be cached, meaning cache efficiency can massively affect economics.
Compression techniques
- The discussion covers:
- quantization: reducing precision from FP32 → FP16/BF16 → FP8 → FP4 and lower
- memory savings vs. quality loss
- Sohmers notes that advanced quantization can reduce memory use dramatically while keeping performance close to full precision.
The Economics of Tokens
Cache economics are extremely lucrative
- He says providers make very high margins on cached tokens because serving cached context is far cheaper than recomputing it.
- This helps explain why major API businesses can be highly profitable.
Token prices are falling, but value is rising
- He argues that:
- the price per million tokens has dropped sharply over time
- but the value per token has increased even more dramatically because model quality has improved
- His broader point: the market should care less about raw token price and more about value per unit of intelligence.
Data Centers, Energy, and the Anti-Data-Center Backlash
His core view
- Sohmers is strongly pro-data-center and sees the current anti-data-center wave as dangerous and often based on misinformation.
- He claims anti-data-center sentiment has become politically unifying in a way that hurts U.S. competitiveness.
Why he thinks the backlash is misguided
- He argues many complaints about:
- water use
- electricity use
- environmental damage are overstated or misleading.
- He says modern data centers are often:
- closed-loop cooled
- built with dedicated generation
- far less harmful than critics imply
Strategic concern
- He believes the U.S. is handicapping itself with unnecessary restrictions while China builds infrastructure aggressively.
- He supports:
- building on federal/desert land
- cleaner energy sources
- faster permitting
- more data-center construction overall
Pacing the Frontier: His View on the AI Safety Debate
He is skeptical of “pacing”
- Sohmers is broadly opposed to slowing frontier AI development.
- He sees two major risks in “pacing” arguments:
- they can become a pathway to outright anti-AI politics
- they may centralize power in a small number of institutions
His main fear: centralization
- He repeatedly emphasizes that the biggest danger is not just existential risk, but:
- concentration of AI capability in governments or a few firms
- a “modern road to serfdom”
- He argues broad restrictions can entrench the biggest players and block open participation.
China and export controls
- He supports free trade and free exchange of ideas in principle.
- But he argues authoritarian regimes should not get equal access to open ecosystems if they can use them to strengthen totalitarian control.
- He is skeptical that export controls alone can truly “pace” AI globally.
What Happens If the Frontier Slows but China Doesn’t?
- Sohmers argues the U.S. cannot assume it will remain ahead if it slows down.
- He draws a geopolitical analogy:
- if a rival can innovate and scale without similar restrictions, catching up becomes plausible
- He says the idea that the West can simply “hold back” and still win is naive.
Models, Scaling Laws, and the Future of Bigger Systems
He still expects frontier models to get bigger
- Sohmers believes frontier models will continue to scale because:
- scaling laws keep working
- capabilities keep improving
- there is no clear sign of a hard stop
Smaller models still matter
- He distinguishes between:
- frontier models for top-end capability
- smaller or on-device models for local utility and orchestration
- But he thinks small models will often increase overall usage of large models by triggering more autonomous background work.
Local models won’t necessarily reduce cloud usage
- His argument:
- local assistants will keep checking email, calendar, messages, documents
- those agents will send more tasks to larger cloud models
- So the presence of local AI could expand overall token consumption, not shrink it.
Why He Thinks Model Capabilities Are Still Improving Fast
- Sohmers says recent releases have materially improved:
- coding
- debugging
- computer use
- design tasks
- chip/RTL workflows
- He describes one model as being able to take a hardware design from spec to GDS in about two days, a task that would otherwise take humans much longer.
- His takeaway: capabilities are still improving in ways that are functionally transformative, not incremental.
The Chip Layer: Everyone Is Building Silicon Now
- He is enthusiastic about the renewed importance of semiconductors.
- He notes that major AI players are increasingly vertically integrating:
- OpenAI
- Anthropic
- DeepSeek
- others building custom silicon
- He sees this as healthy for the industry:
- more competition
- lower costs
- better capabilities
- But he also thinks many different architectural approaches can coexist.
Context Windows, Long Context, and the Next Constraint
Bigger context matters
- Sohmers believes context length is one of the next major battlegrounds.
- He wants context windows large enough to hold:
- massive codebases
- cross-project reasoning
- full enterprise knowledge graphs
The limitation is not just max length
- He stresses that “advertised” context length is less important than whether the model can actually use that context effectively.
- He cites long-context benchmarks where some models fail badly at recall once context gets large.
Hardware and algorithms both matter
- He expects hardware to keep pushing memory capacity upward.
- But he also credits algorithmic advances from Chinese labs:
- sparse attention
- linear attention
- MLA-style compression
- gated DeltaNet variants
- His view: these lower memory costs, but often with tradeoffs in capability.
Energy as the Ultimate Bottleneck
- Sohmers frames energy as the long-term governor of progress:
- intelligence is tied to compute
- compute is tied to power
- He thinks the bigger practical limiter is not physical energy availability alone, but:
- economics
- debt
- infrastructure buildout
- regulatory friction
- He is more worried about sovereign debt and macro constraints than about AI companies missing revenue targets.
His Most Optimistic Takeaway
- Sohmers ends on a surprisingly broad point: the AI alignment problem is not just a machine problem, but a human alignment problem.
- In his view, the key question is whether society can:
- align incentives
- build enough energy
- permit enough infrastructure
- avoid self-imposed bottlenecks
- He believes smart people should treat these as solvable, systems-level problems rather than purely technical ones.
Main Takeaways
- Inference is becoming the main infrastructure challenge because it is memory-bound, not compute-bound.
- KV caching and quantization are central to making AI serving economically viable.
- Data centers are strategic infrastructure, and he считает anti-data-center politics dangerous and often misinformation-driven.
- Pacing the frontier may sound prudent, but he thinks it risks centralization, bureaucracy, and geopolitical disadvantage.
- Bigger models are still coming, while smaller models may actually increase total cloud usage.
- Energy, economics, and regulation may matter more than raw model architecture in determining how fast AI scales.
Notable Themes
- AI infrastructure as the new industrial base
- Geopolitics of compute and data centers
- The hidden economics of cached tokens
- Frontier-model competition vs. regulatory restraint
- The long-term link between energy production and intelligence
