Haters think AI agents can't write GPU code? This'll ROCm

Summary of Haters think AI agents can't write GPU code? This'll ROCm

by The Stack Overflow Podcast

26m•September 22, 2026

Overview of The Stack Overflow Podcast

In this episode of The Stack Overflow Podcast, host Ryan Donovan speaks with Anush Elangovan, VP of Software at AMD, about how AMD’s ROCm stack, GPU programming, and agentic AI are changing the way developers write, test, and optimize code for GPUs. The conversation centers on how AMD is making GPU development more open, more accessible, and increasingly AI-assisted—from low-level kernel work all the way up to high-level Pythonic workflows and automated performance tuning.

What ROCm is and why AMD made it open

  • ROCm is AMD’s open software platform for GPU programming, similar in purpose to CUDA, but built around an open-source philosophy.
  • It acts as a unified toolchain and software layer across AMD accelerators, including:
    • GPUs
    • CPUs
    • FPGAs (via Xilinx heritage)
  • AMD’s goal is to provide:
    • open source code
    • open build and development processes
    • a faster feedback loop for community contributions
  • Contributors can now:
    • file issues
    • submit feature PRs
    • have changes rolled into nightly builds

Why open source matters here

  • AMD wants the ecosystem to innovate on top of the platform, not wait on a closed binary roadmap.
  • Because ROCm is open, frontier AI models can also learn from the source code and specs, which helps AI tools generate and understand ROCm-related code more effectively.

GPU programming: from low-level ISA to Pythonic abstractions

  • AMD publishes the ISA specifications for its GPUs, which means developers can work very close to the hardware.
  • The guest explains that GPUs are fundamentally still programmable machines with familiar concepts like:
    • compute
    • load/store
    • loops
  • GPU programming differs from CPU programming because it is built around SIMT execution:
    • thousands of threads
    • highly parallel workloads
    • best suited for matrix ops, HPC, scientific computing, and AI

The abstraction ladder

The episode emphasizes that GPU programming is no longer only for low-level experts:

  • Bottom layer: assembly / ISA-level work
  • Middle layers: ROCm kernels and GPU tooling
  • Higher-level layers: Pythonic DSLs like Triton-style programming
  • Newest layer: agentic AI that can generate or optimize kernels from a natural-language prompt

The key idea is that AMD is giving developers a “layer cake” of access: people can work at the level they’re comfortable with.

How AI agents are changing GPU development

A major theme of the episode is that agentic AI is lowering the barrier to GPU programming.

What AI agents can do now

  • Generate GPU kernels from a specification
  • Help rewrite or optimize code
  • Assist with benchmarking and validation
  • Automate performance analysis
  • Help with driver work and even large engineering tasks

The guest gives examples of AI making previously huge projects feel much more approachable, such as:

  • rewriting GPU-related software components
  • generating thousands of lines of code quickly
  • accelerating internal development workflows

Why ROCm is especially AI-friendly

  • ROCm’s openness means AI tools have more visibility into:
    • APIs
    • specs
    • source code
    • contribution history
  • That makes ROCm easier for frontier models to understand than closed systems.

AMD’s internal use of AI

AMD is not just building AI tools for customers; it is also using them internally.

How AMD teams are using agents

  • Code generation
  • Benchmarking and validation
  • Long-running analysis tasks
  • Hardware/software co-design assistance
  • Performance and power optimization
  • Scheduling and infrastructure work

The guest says that at AMD, the mindset has shifted from “AI as an assistant” to AI as an agent.

The impact on engineering velocity

  • The team is moving toward a world where:
    • software is written earlier
    • validation starts sooner
    • hardware ideas are tested in software almost immediately
  • This “shift-left” approach means AMD often reaches an optimized software state before the silicon even tapes out.

Hardware, simulation, and the shift-left workflow

Because hardware cannot be changed instantly, AMD relies on several stages:

  • Functional verification
  • Simulation
  • Emulation
  • FPGA-based prototyping
  • Final silicon validation

The main point is that software and hardware are now tightly coupled:

  • hardware ideas should be paired with software prototypes immediately
  • AI helps compress the feedback loop
  • teams can validate design choices faster and earlier

Why this matters

  • Faster iteration reduces risk
  • Bugs and performance issues are caught earlier
  • Engineers can make better decisions before manufacturing is locked in

The practical value of AI on AMD platforms

The episode also highlights the user-facing value of AI on AMD hardware:

  • Client/laptop: local inference without cloud dependency
  • Desktop: AI for gaming, simulations, and advanced workloads
  • Cloud/data center: large-scale inference and frontier model deployment

A specific example mentioned is Strix Halo, positioned as a strong local inference platform thanks to high memory capacity, making it possible to run very large models on a laptop.

Main takeaways

  • ROCm is AMD’s open-source GPU software stack, designed to make hardware programming more accessible and more customizable.
  • GPU programming is becoming easier thanks to higher-level abstractions and Pythonic tooling.
  • Agentic AI is dramatically lowering the barrier to writing and optimizing GPU code.
  • Open source helps AI tools learn faster, because models can ingest the code, specs, and contribution history.
  • AMD is using AI internally to accelerate both software and hardware development.
  • Hardware and software are converging, with AI-driven workflows helping AMD optimize earlier in the product cycle.

Notable closing thought

The guest’s core message is that the industry is moving toward a world where developers can choose the right abstraction level for the job—and AI now helps them operate effectively at levels they may not have had the expertise or time to reach before. The result is faster engineering, broader participation, and more ambitious projects across AMD’s stack.