Professional skepticism is a dev’s best skill

Summary of Professional skepticism is a dev’s best skill

by The Stack Overflow Podcast

28m•September 25, 2026

Overview of Professional skepticism is a dev’s best skill

In this Stack Overflow Podcast episode, Ryan Donovan talks with David Burns, chair of the W3C Browser Testing and Tools Working Group and co-editor of the WebDriver specification, about why professional skepticism is one of the most important skills in software development. The conversation focuses on QA, testing, agentic engineering, and why human judgment still matters even as AI tools generate more code and more automated tests.

Why skepticism matters in QA and software engineering

David argues that QA professionals are valuable because they approach software differently:

  • They think from the user’s perspective: “If I used this, would it make sense?”
  • They are trained to question assumptions rather than trust the happy path.
  • They tend to notice things like slow load times, awkward flows, and architecture that “feels wrong.”
  • They help teams surface bugs earlier by challenging systems instead of accepting them.

His core point: QA is really about confidence—confidence that software does what it should, in the way it should, and under the conditions that matter.

AI, agents, and test-driven development

A major theme is how QA fits into AI-assisted development and agentic workflows.

Test-driven development still applies

David says agentic engineering should still follow a familiar TDD-style process:

  1. Write the tests first based on the requirements.
  2. Have the agent generate code to make the tests pass.
  3. Require human approval for any changes to the tests.

He stresses that tests are not just validation—they also act as documentation of intent and help protect against bad assumptions.

Human review remains essential

David is skeptical of handing test creation entirely to AI because:

  • Hallucinations still happen.
  • AI tends to be a “people pleaser” and may agree too easily.
  • A PRD alone is not enough to capture intent.
  • Humans need to verify whether the output matches what was actually meant.

Looser vs tighter tests

The discussion also covers how strict tests should be.

  • Some checks should be very tight, especially for critical systems where failure could be severe.
  • Other checks can be looser, especially for visual or layout-related behavior.
  • The standard should be risk-based: tighter where the business or security impact is higher.

David’s example: a login flow redirecting users to localhost should absolutely have been caught by testing, while some UI positioning checks can tolerate more flexibility.

Vibe coding, misuse, and safety boundaries

The episode spends time on the risks of “vibe coding” and over-trusting AI-generated changes.

Problems David sees repeatedly

  • AI accidentally touching production systems
  • People clicking through prompts without reading them
  • Dangerous scripts or actions being approved too casually
  • The same old failures reappearing under new tooling

His recommendation

The industry should reintroduce stronger guardrails:

  • No broad root access by default
  • Explicit permission for dangerous actions
  • Better orchestration and policy controls for agents
  • Auditing and supply-chain awareness for code/package actions

He argues that many of the safeguards we used to have should now be doubled down on, not relaxed.

Testing strategy and flaky tests

David gives a strong critique of how teams often handle flaky tests and large UI test suites.

Key observations

  • Flakiness often gets blamed on the framework, but the root issue is usually deeper.
  • The real problem is often a mismatch between:
    • synchronous test scripts, and
    • asynchronous applications and systems
  • Teams keep switching frameworks—Selenium, Cypress, Puppeteer, Playwright—but the underlying testing problems remain.

What good automation looks like

A strong automated test should have:

  • a known starting state,
  • one clear purpose,
  • a clean return to the original state.

He also emphasizes measuring flakiness over time and tracking the delta, rather than chasing an unrealistic “100% green” standard in huge suites.

The testing pyramid and practical QA coverage

David supports using the testing pyramid thoughtfully:

  • Keep the bulk of tests low-cost and focused where possible.
  • Use component and unit tests to isolate behavior.
  • Reserve end-to-end tests for a few high-value, risk-based flows.
  • Add observability and monitoring to improve quality gates over time.

He warns against inverted pyramids and “testing Christmas trees” where teams overinvest in expensive end-to-end coverage and lose speed and confidence.

Automation is growing, but judgment is more valuable

One of the episode’s strongest takeaways is that as automation becomes easier, judgment and taste become more valuable.

David believes senior QA engineers, exploratory testers, and automation specialists will remain important because they can:

  • recognize what matters most to users,
  • decide what should be tested deeply,
  • know when a result is good enough,
  • and identify gaps in knowledge about how the system works.

He sees exploratory testing and automation as complementary: exploratory testing builds understanding, which then improves automation.

Notable takeaways

  • Professional skepticism is a core engineering skill, not just a QA skill.
  • AI can speed up development, but it does not remove the need for verification.
  • Tests should encode intent, not just validate syntax or behavior.
  • Dangerous operations require stronger guardrails, not looser ones.
  • Flaky tests are often a sign of bad assumptions or poor test design, not just bad tooling.
  • In an AI-heavy world, confidence, judgment, and taste are increasingly important.

Closing thoughts

The episode argues for a balanced approach: use AI to accelerate work, but keep humans firmly in the loop for review, intent, and risk management. David’s bottom line is that engineers should not stop being skeptical—if anything, they should become more so, because software systems and agentic tools are moving faster than ever.