#561: TonIO, a Multi-threaded Async Runtime for Python

Summary of #561: TonIO, a Multi-threaded Async Runtime for Python

by Michael Kennedy

1h 15m•September 4, 2026

Overview of #561: TonIO, a Multi-threaded Async Runtime for Python

This episode is a deep dive into TonIO (Tonio), a new async runtime for Python built from scratch for free-threaded Python. Michael Kennedy talks with Giovanni “Joe” Barillari, creator of Granian (the Rust-based app server behind Talk Python), about why traditional asyncio still leaves most CPU cores unused, how free-threaded Python changes the game, and what TonIO is trying to simplify: real multithreading, a smaller set of primitives, and a runtime that refuses to run if the GIL is present.


Why TonIO Exists

Joe’s core argument is simple:

  • Classic Python async I/O still runs on one interpreter thread
  • Faster event loops like uvloop improve performance, but don’t change the single-threaded shape
  • With modern machines having many cores, the bottleneck is often memory and concurrency management, not raw CPU speed
  • Free-threaded Python finally makes it realistic to use actual threads without the GIL blocking the model

He sees free-threaded Python as the biggest change in Python in decades, because it fundamentally changes how Python applications can scale.


Granian: The Foundation and Motivation

The conversation starts by revisiting Granian, Joe’s Rust-based Python application server:

  • It serves Python web apps through WSGI, ASGI, and Granian’s own RSGI
  • Because the HTTP/network layer runs in Rust, it reduces interpreter load and stabilizes latency
  • Michael notes that Talk Python’s infrastructure has benefitted from Granian’s reliability and performance
  • Granian is already used at companies like Sentry, Microsoft, and Google

Key takeaway: Granian’s architecture made Joe think more deeply about runtime design, which led into TonIO.


What TonIO Is

TonIO is a multi-threaded async runtime for Python designed specifically for free-threaded CPython.

Core ideas

  • Uses real OS threads
  • Starts with a thread count based on CPU cores
  • Has a small, focused set of primitives
  • Rejects execution if the GIL is enabled
  • Aims to replace the complexity of asyncio with something cleaner and more predictable for multithreaded execution

Two syntax styles

TonIO supports two ways to write async code:

  • Async/await syntax, similar to asyncio
  • A generator/yield-based syntax, for people who dislike async/await

Joe’s point: since the runtime was written from scratch, supporting both syntaxes was feasible.


How TonIO Differs from asyncio

Joe emphasizes several major differences:

1. Spawning is explicit and immediate

TonIO centers around spawn:

  • spawn(...) starts work immediately
  • spawn_blocking(...) sends CPU/blocking work to a separate blocking pool
  • You can await spawned work later, or store the handle and join later

This makes concurrency feel more “real” and less tied to the old coroutine/task juggling model.

2. Fewer primitives

TonIO intentionally has fewer abstractions than asyncio, because Joe считает asyncio is overcomplicated:

  • fewer primitives
  • clearer runtime model
  • less conceptual overhead

3. Each thread has its own event loop

In TonIO:

  • every worker thread has its own event loop
  • this is great for parallelism
  • but it also means cross-thread sharing is limited by design

That’s one of the big tradeoffs TonIO is trying to address.


Why Free-Threaded Python Matters

Joe is very excited about Python 3.14’s stable free-threaded mode.

Benefits he highlights

  • Python can finally use multiple cores directly
  • Web apps can avoid multiplying processes just to escape the GIL
  • Memory usage can drop significantly because you don’t need 4 copies of the same app process
  • Shared state, metrics, and caches become more feasible in a single process model

Practical impact

This matters especially for large Python services like Sentry’s monolith:

  • process-based scaling wastes memory
  • each process duplicates caches and instrumentation state
  • free-threading can reduce that overhead substantially

AsyncIO Complexity and the “Coloring” Problem

Joe’s critique of asyncio is not just performance-related — it’s also about ergonomics.

Pain points he called out

  • Too many primitives and concepts:
    • futures
    • tasks
    • handles
    • protocols
    • transports
  • The runtime model is hard to reason about when debugging
  • The difference between a protocol and transport is notoriously confusing
  • Even simple scripts still require asyncio.run(...)

He wants TonIO to feel more like:

  • “there is a runtime”
  • “there are threads”
  • “there is a loop”
  • “you don’t have to micromanage all of it”

Concurrency, Locks, and New Failure Modes

A major theme of the episode is that free-threading does not remove the need for synchronization — it just makes the rules explicit.

Important rules Joe recommends

  1. Never await inside a threading lock
    • this can deadlock TonIO
  2. Be careful when writing shared state
    • Python objects may be thread-safe at the C level, but your logic can still have race conditions

TonIO’s sync support

TonIO includes its own synchronization primitives:

  • lock
  • semaphore
  • barrier
  • other sync tools in a dedicated sync module

So yes, TonIO has async/threading primitives, but the user now has to think like a multithreaded programmer.


TonIO’s Modules and Runtime Structure

Joe described TonIO as a more self-contained runtime with dedicated modules:

  • tonio.time — timers, timeouts, timing utilities
  • tonio.sync — locks, semaphores, barriers, etc.
  • tonio.network — sockets and networking
  • tonio.network.streams — high-level stream APIs
  • tonio.fs — async file-system operations
  • tonio.fs.path — async-aware path operations

He also mentioned:

  • tonio.main for entry-point decoration
  • tonio.run for running a configured runtime

Ecosystem Compatibility and TonIOMonkey

Because existing Python libraries are mostly written for asyncio, TonIO is not drop-in compatible with everything.

Current workaround

Joe built TonIOMonkey, a monkey-patching compatibility package for common libraries, including:

  • psycopg
  • httpx
  • httpx2
  • redis
  • fastapi

Future direction

He hopes to eventually get a proper NIO backend or equivalent integration so libraries can support TonIO natively rather than through patches.

Known limitation

TonIO does not support asyncio primitives, so codebases heavily dependent on tasks/futures may need adaptation.


AI as a Development Tool

One of the more interesting side topics was Joe’s use of AI.

He said TonIO itself is mostly hand-written, but he used AI in a clever way:

  • not to write TonIO directly
  • but to create test applications and stress harnesses around TonIO
  • for example, he had Claude rewrite the Py test harness in Python to exercise the runtime
  • this helped surface most of TonIO’s bugs

That approach let AI generate workload and ecosystem pressure without handing over core runtime design.


Roadmap: Granian, Windows, and Multi-Runtime Support

Granian future plans

Joe wants Granian to eventually support multiple Python runtimes:

  • asyncio
  • Trio
  • TonIO
  • gevent
  • eventlet
  • and more

The idea is to separate the protocol layer from the runtime layer.

Windows support

TonIO is currently POSIX-only.

Joe said Windows support is hard because:

  • Windows has quirks with sockets and file descriptors
  • he doesn’t currently develop on Windows
  • the underlying eventing layer behaves differently

Still, he hinted partial Windows support may arrive later, likely through low-level hacks in the I/O layer.


Main Takeaways

  • Free-threaded Python is a major shift, and Joe sees it as a once-in-a-generation change
  • TonIO is an attempt to redesign async Python around that new reality
  • The runtime trades asyncio familiarity for:
    • real threads
    • simpler primitives
    • explicit spawning
    • better multi-core utilization
  • It is still alpha, but it’s already being used to test the future of Python concurrency
  • The project is tightly linked to Granian, which may eventually support TonIO directly

Call to Action

If you want to try TonIO:

  • Experiment with it on free-threaded Python 3.14+
  • Check whether your libraries work through TonIOMonkey
  • Open issues or discussions if you find gaps in compatibility
  • Follow Joe’s work on GitHub and his blog for updates

Joe’s contact points

  • GitHub / social handle: gi0bari (as referenced in the episode)
  • Blog: blog.baro.dev
  • Co-hosted podcast with Marcelo: The WWW Pod on YouTube

Notable Closing Thought

Joe’s big message was that now is the right time to rethink Python concurrency:

If you’re unhappy with asyncio, or if you’ve hit its scaling and complexity limits, free-threaded Python opens the door to a different model entirely.

TonIO is his bet on what that model should look like.