#560: Building a Research OS: From Django to 30,000 Samples

Summary of #560: Building a Research OS: From Django to 30,000 Samples

by Michael Kennedy

1h 2m•August 26, 2026

Overview of Talk Python to Me Episode 560

This episode features Sean Chua, a gastroenterologist and clinical researcher at the University of Glasgow, discussing how he went from minimal programming experience to building Foundry 120: a Django-based “research operating system” for translational science. What began as a practical solution for tracking a projected 30,000 clinical samples has evolved into a platform managing 10 terabytes of clinical and genomics data, backed by Azure infrastructure and powered by agentic AI to help researchers find, join, analyze, and visualize data faster.

The Origin Story: From Excel Pain to Django App

The problem

  • Sean’s research team needed to manage:
    • hundreds of participants
    • repeated follow-ups
    • blood, stool, saliva, and other sample types
    • multiple hospitals across multiple cities
  • The expected volume quickly became too much for manual tracking and spreadsheets.

Why Django

  • After exploring off-the-shelf tools, Sean built a custom web app with Django.
  • Django was appealing because it was:
    • mature and stable
    • “batteries included”
    • strong on authentication, migrations, and admin tooling
  • His programming background was modest:
    • hand-written HTML in Notepad in high school
    • some occasional R/Python scripting for stats and plotting

Learning approach

  • He leaned heavily on:
    • Django documentation
    • YouTube tutorials
    • books like Two Scoops of Django
    • broader software/data engineering books such as Designing Data-Intensive Applications and The Data Warehouse Toolkit

From Sample Tracking to Data Platform

What the first version did

  • The first Django app handled:
    • sample registration
    • QR code labeling/scanning
    • tracking sample movement across sites
  • It solved a very specific operational problem for a research study.

How it expanded

As the project matured, the platform had to manage not just samples, but the data those samples produced:

  • microbiome data from stool samples
  • genomics data from blood samples
  • imaging and clinical outputs
  • large sequencing files, sometimes 5–10 GB per participant

Scale and storage

  • The project now handles about 10 TB of data.
  • Data used to be scattered across:
    • laptops
    • shared drives
    • spreadsheets
  • It has since moved to Azure Blob Storage for scalability and better control.

Foundry 120: A Research Operating System

Core idea

Foundry 120 is designed as a research operating system for translational science teams.

Three main goals

  1. Organize everything
    • participants
    • studies
    • samples
    • access and workflow metadata
  2. Centralize data
    • clinical records
    • radiology images
    • endoscopy videos
    • pathology slides
    • genomics, transcriptomics, microbiome data
  3. Enable AI-assisted research
    • use AI to find, join, and analyze data across modalities

Why this matters

  • Researchers often need to answer questions like:
    • “Do we have samples from participants treated with a certain drug?”
    • “What’s the relationship between CRP and cell-free DNA?”
  • Doing this manually can take days of querying, joining, and checking data.
  • Foundry aims to reduce that to minutes.

How the Data Model Works

Files are the universal denominator

A key design decision was to treat files as the core data unit:

  • CSV for tabular data
  • images, videos, slides, and sequencing outputs as files
  • metadata and relationships layered on top

Why this approach works

  • It generalizes across scientific modalities
  • It matches how researchers already think about data
  • It makes it easier for both humans and AI systems to operate on the same structure

Related technologies mentioned

  • Parquet
  • DuckDB
  • SQLite/DuckDB-style embedded file databases as possible future directions

Agentic AI: Helix

What Helix does

Helix is Foundry 120’s agentic AI system. Unlike a chatbot, it:

  • receives a task
  • chooses tools
  • writes and runs analysis code
  • checks results
  • recovers from errors
  • returns a validated answer

Examples shown in the demo

  • Ask: “What plasma samples do we have for this study, broken down by disease group?”
    • Helix joins clinical and sample data
    • counts disease groups
    • returns a summary such as Crohn’s vs ulcerative colitis sample counts
  • Ask: “Plot CRP vs cell-free DNA”
    • Helix finds the relevant datasets
    • writes Python/matplotlib code
    • produces a graph
    • fixes issues if the first attempt fails

The key distinction

  • This is not model fine-tuning on all research data.
  • It is more like tool-using AI:
    • ground the model with structured tools
    • let it reason and execute workflows
    • keep outputs constrained and checkable

Security, Governance, and Compliance

Why the system is guarded tightly

Clinical research has strict governance requirements:

  • data residency
  • ethics approvals
  • role-based access control
  • privacy and security constraints

How they protect the data

  • Django acts as the security gate between AI and data
  • The AI does not directly access raw files or compute
  • Requests go through Django first
  • Permissions are enforced before compute jobs are launched

Infrastructure details

  • Runs in the University of Glasgow’s Azure tenancy
  • Data and inference stay within approved regional data centers
  • This is important for GDPR and institutional compliance

Architecture and Scaling

Front end / back end

  • Django REST Framework on the back end
  • TypeScript/React on the front end
  • Server-side architecture works well because:
    • hospital computers can be old/limited
    • front-line users need fast, reliable experiences
    • heavy compute should not run on local machines

Compute model

  • Foundry can orchestrate large jobs across about 350 Azure CPUs
  • Many workloads are highly parallel:
    • process many participants at once
    • spin up compute on demand
    • tear it down afterward
  • This is more efficient than buying and maintaining on-prem hardware for sporadic workloads

Broader Takeaways

For researchers

  • Don’t underestimate operational tooling; sample/data tracking can become the bottleneck.
  • Use metadata + files as a flexible foundation for multi-modal science.
  • AI is most useful when it can execute workflows, not just answer questions.

For developers

  • You can build serious research infrastructure even with limited initial experience.
  • Django remains a strong choice for systems that need:
    • reliability
    • permissions
    • database integrity
    • rapid delivery
  • Books, docs, and community resources are still invaluable for self-taught engineers.

For the AI conversation

  • Agentic AI is very different from “chatbot AI.”
  • Reliability is improving, but humans still need to supervise and validate results.
  • The future likely includes:
    • specialized models
    • local and cloud AI
    • tighter orchestration between software tools and models

Final Thoughts

Sean’s story is a good example of how a real-world research pain point can grow into a full platform:

  • from Excel tracking
  • to custom Django apps
  • to cloud-native data infrastructure
  • to agentic AI for scientific workflows

The episode highlights a major theme: in modern biomedical research, the winning system is not just about data storage or analysis alone, but about connecting operational workflows, governance, and AI-assisted execution into one cohesive platform.