Overview of #561: TonIO, a Multi-threaded Async Runtime for Python
This episode is a deep dive into TonIO (Tonio), a new async runtime for Python built from scratch for free-threaded Python. Michael Kennedy talks with Giovanni “Joe” Barillari, creator of Granian (the Rust-based app server behind Talk Python), about why traditional asyncio still leaves most CPU cores unused, how free-threaded Python changes the game, and what TonIO is trying to simplify: real multithreading, a smaller set of primitives, and a runtime that refuses to run if the GIL is present.
Why TonIO Exists
Joe’s core argument is simple:
- Classic Python async I/O still runs on one interpreter thread
- Faster event loops like uvloop improve performance, but don’t change the single-threaded shape
- With modern machines having many cores, the bottleneck is often memory and concurrency management, not raw CPU speed
- Free-threaded Python finally makes it realistic to use actual threads without the GIL blocking the model
He sees free-threaded Python as the biggest change in Python in decades, because it fundamentally changes how Python applications can scale.
Granian: The Foundation and Motivation
The conversation starts by revisiting Granian, Joe’s Rust-based Python application server:
- It serves Python web apps through WSGI, ASGI, and Granian’s own RSGI
- Because the HTTP/network layer runs in Rust, it reduces interpreter load and stabilizes latency
- Michael notes that Talk Python’s infrastructure has benefitted from Granian’s reliability and performance
- Granian is already used at companies like Sentry, Microsoft, and Google
Key takeaway: Granian’s architecture made Joe think more deeply about runtime design, which led into TonIO.
What TonIO Is
TonIO is a multi-threaded async runtime for Python designed specifically for free-threaded CPython.
Core ideas
- Uses real OS threads
- Starts with a thread count based on CPU cores
- Has a small, focused set of primitives
- Rejects execution if the GIL is enabled
- Aims to replace the complexity of
asynciowith something cleaner and more predictable for multithreaded execution
Two syntax styles
TonIO supports two ways to write async code:
- Async/await syntax, similar to
asyncio - A generator/yield-based syntax, for people who dislike
async/await
Joe’s point: since the runtime was written from scratch, supporting both syntaxes was feasible.
How TonIO Differs from asyncio
Joe emphasizes several major differences:
1. Spawning is explicit and immediate
TonIO centers around spawn:
spawn(...)starts work immediatelyspawn_blocking(...)sends CPU/blocking work to a separate blocking pool- You can await spawned work later, or store the handle and join later
This makes concurrency feel more “real” and less tied to the old coroutine/task juggling model.
2. Fewer primitives
TonIO intentionally has fewer abstractions than asyncio, because Joe считает asyncio is overcomplicated:
- fewer primitives
- clearer runtime model
- less conceptual overhead
3. Each thread has its own event loop
In TonIO:
- every worker thread has its own event loop
- this is great for parallelism
- but it also means cross-thread sharing is limited by design
That’s one of the big tradeoffs TonIO is trying to address.
Why Free-Threaded Python Matters
Joe is very excited about Python 3.14’s stable free-threaded mode.
Benefits he highlights
- Python can finally use multiple cores directly
- Web apps can avoid multiplying processes just to escape the GIL
- Memory usage can drop significantly because you don’t need 4 copies of the same app process
- Shared state, metrics, and caches become more feasible in a single process model
Practical impact
This matters especially for large Python services like Sentry’s monolith:
- process-based scaling wastes memory
- each process duplicates caches and instrumentation state
- free-threading can reduce that overhead substantially
AsyncIO Complexity and the “Coloring” Problem
Joe’s critique of asyncio is not just performance-related — it’s also about ergonomics.
Pain points he called out
- Too many primitives and concepts:
- futures
- tasks
- handles
- protocols
- transports
- The runtime model is hard to reason about when debugging
- The difference between a protocol and transport is notoriously confusing
- Even simple scripts still require
asyncio.run(...)
He wants TonIO to feel more like:
- “there is a runtime”
- “there are threads”
- “there is a loop”
- “you don’t have to micromanage all of it”
Concurrency, Locks, and New Failure Modes
A major theme of the episode is that free-threading does not remove the need for synchronization — it just makes the rules explicit.
Important rules Joe recommends
- Never
awaitinside a threading lock- this can deadlock TonIO
- Be careful when writing shared state
- Python objects may be thread-safe at the C level, but your logic can still have race conditions
TonIO’s sync support
TonIO includes its own synchronization primitives:
locksemaphorebarrier- other sync tools in a dedicated
syncmodule
So yes, TonIO has async/threading primitives, but the user now has to think like a multithreaded programmer.
TonIO’s Modules and Runtime Structure
Joe described TonIO as a more self-contained runtime with dedicated modules:
tonio.time— timers, timeouts, timing utilitiestonio.sync— locks, semaphores, barriers, etc.tonio.network— sockets and networkingtonio.network.streams— high-level stream APIstonio.fs— async file-system operationstonio.fs.path— async-aware path operations
He also mentioned:
tonio.mainfor entry-point decorationtonio.runfor running a configured runtime
Ecosystem Compatibility and TonIOMonkey
Because existing Python libraries are mostly written for asyncio, TonIO is not drop-in compatible with everything.
Current workaround
Joe built TonIOMonkey, a monkey-patching compatibility package for common libraries, including:
psycopghttpxhttpx2redisfastapi
Future direction
He hopes to eventually get a proper NIO backend or equivalent integration so libraries can support TonIO natively rather than through patches.
Known limitation
TonIO does not support asyncio primitives, so codebases heavily dependent on tasks/futures may need adaptation.
AI as a Development Tool
One of the more interesting side topics was Joe’s use of AI.
He said TonIO itself is mostly hand-written, but he used AI in a clever way:
- not to write TonIO directly
- but to create test applications and stress harnesses around TonIO
- for example, he had Claude rewrite the Py test harness in Python to exercise the runtime
- this helped surface most of TonIO’s bugs
That approach let AI generate workload and ecosystem pressure without handing over core runtime design.
Roadmap: Granian, Windows, and Multi-Runtime Support
Granian future plans
Joe wants Granian to eventually support multiple Python runtimes:
asyncio- Trio
- TonIO
- gevent
- eventlet
- and more
The idea is to separate the protocol layer from the runtime layer.
Windows support
TonIO is currently POSIX-only.
Joe said Windows support is hard because:
- Windows has quirks with sockets and file descriptors
- he doesn’t currently develop on Windows
- the underlying eventing layer behaves differently
Still, he hinted partial Windows support may arrive later, likely through low-level hacks in the I/O layer.
Main Takeaways
- Free-threaded Python is a major shift, and Joe sees it as a once-in-a-generation change
- TonIO is an attempt to redesign async Python around that new reality
- The runtime trades
asynciofamiliarity for:- real threads
- simpler primitives
- explicit spawning
- better multi-core utilization
- It is still alpha, but it’s already being used to test the future of Python concurrency
- The project is tightly linked to Granian, which may eventually support TonIO directly
Call to Action
If you want to try TonIO:
- Experiment with it on free-threaded Python 3.14+
- Check whether your libraries work through TonIOMonkey
- Open issues or discussions if you find gaps in compatibility
- Follow Joe’s work on GitHub and his blog for updates
Joe’s contact points
- GitHub / social handle:
gi0bari(as referenced in the episode) - Blog: blog.baro.dev
- Co-hosted podcast with Marcelo: The WWW Pod on YouTube
Notable Closing Thought
Joe’s big message was that now is the right time to rethink Python concurrency:
If you’re unhappy with
asyncio, or if you’ve hit its scaling and complexity limits, free-threaded Python opens the door to a different model entirely.
TonIO is his bet on what that model should look like.
