Building an agentic SDLC with a QA engineering mindset

Summary of Building an agentic SDLC with a QA engineering mindset

by The Stack Overflow Podcast

28m•August 18, 2026

Overview of Building an agentic SDLC with a QA engineering mindset

This episode of the Stack Overflow Podcast explores how agentic AI is changing the software development lifecycle, especially through the lens of QA engineering. Host Ryan Donovan speaks with Sunit Malhotra, Senior Manager of Test Engineering at Motorola Solutions, about building an end-to-end, AI-assisted SDLC pipeline that goes from Figma designs to requirements, Jira tickets, developer work, PR review, QA, automation, and reporting. The conversation centers on how to use LLMs and agents effectively while still keeping humans in the loop for judgment, context, and quality control.

Key Themes and Takeaways

  • The SDLC is becoming more agent-driven: Instead of using AI only for coding assistance, Sunit describes a broader pipeline where agents help with planning, requirements, implementation, review, test generation, and reporting.
  • QA thinking is central to agentic workflows: His background in QA shapes how he approaches AI systems—by focusing on edge cases, failure modes, and validation rather than assuming the model will get everything right.
  • Context matters more than ever: Generic prompts or specs produce weak results. The best outcomes come from giving agents detailed context about the app, platform, codebase, conventions, and launch goals.
  • Human review is still required: Even with strong automation, the system typically reaches only about 70–80% accuracy at first. Humans are needed to review, refine, and catch missing cases.
  • Evaluation is the real bottleneck: As AI generates more of the work, the hardest problem becomes reliably judging whether its output is actually correct.

The Agentic SDLC Pipeline

End-to-end flow described in the episode

Sunit outlines a multi-agent workflow that can support feature delivery from start to finish:

  1. Design intake

    • Start with a Figma design.
    • Use agents connected through MCP (Model Context Protocol) to pull in supporting context from tools like Confluence.
  2. Requirements generation

    • A product manager agent drafts a PRD/requirements document.
    • The agent is enriched with platform, launch, and business context.
    • A human reviews and fills in gaps such as negative cases or strategy concerns.
  3. Jira planning

    • Another agent creates Jira epics and tasks from the requirements.
    • This reduces manual ticket creation and keeps work aligned with the spec.
  4. Development

    • A developer agent uses the requirements, Figma, and Jira context to generate code and PRs.
    • In the example discussed, AI writes a large amount of the code.
  5. PR review

    • A PR reviewer agent checks the generated code to catch issues a human might miss at scale.
  6. QA and automation

    • A QA agent creates test cases using tools like TestRail MCP.
    • An automation agent writes scripts, using Appium + TypeScript for mobile automation.
  7. Reporting

    • Final results are pushed to tools like Slack for visibility and coordination.

LLMs as Judges and Evaluation Strategy

A major topic in the episode is how to evaluate AI-generated work without making humans review everything manually.

Why LLM judgment matters

  • Manual review of hundreds of test cases or checks is too slow and not scalable.
  • Sunit experimented with using models like Claude, Codex, Ollama, and DeepSeek as evaluators.
  • The goal is to have one AI model judge the output of another AI model.

Confidence and agreement

  • He uses Cohen’s Kappa to measure agreement between AI judges.
  • A threshold of around 0.7 is treated as a meaningful level of agreement, beyond random coincidence.
  • This allows humans to focus on the uncertain or failing cases rather than reviewing every single result.

Why this helps

  • Frees humans from routine grading.
  • Makes AI-assisted QA more scalable.
  • Creates a more practical hybrid model: AI does most of the work, humans handle edge cases and final judgment.

Challenges, Limits, and What Still Needs Humans

Common failure modes

  • LLMs often miss:
    • Security concerns
    • Localization needs
    • Accessibility requirements
    • Other app-specific or company-specific design constraints

The “context vs. spec” distinction

  • A spec alone is not enough.
  • Agents need:
    • the product goals,
    • the codebase,
    • supported platforms,
    • testing conventions,
    • and launch constraints.
  • Without this, the output can devolve into low-value “AI slop.”

Sycophancy and over-approval

  • Some models may try too hard to satisfy the goal, even if that means incorrectly passing something that should fail.
  • This makes careful prompting, constraints, and validation checks essential.

Specification Enrichment: Shifting QA Earlier

One of the most interesting ideas in the episode is specification enrichment.

What it does

  • Adds an early LLM stage that inspects a design before requirements are written.
  • It checks for missing assumptions and hidden constraints.
  • It can ask targeted questions back to a human, such as whether the product needs:
    • localization,
    • accessibility coverage,
    • security scenarios,
    • or other launch-specific concerns.

Why it matters

  • Moves QA further left in the lifecycle.
  • Helps avoid late-stage requirement changes and bug discoveries.
  • Improves downstream quality for development, testing, and automation.

Practical Advice and Open Source Notes

How to get started

  • Sunit says the open source projects are simple to run:
    • download the code,
    • connect access to a tool like Cloud Code,
    • and evaluate the workflows.
  • He includes instructions in the README for users who want to experiment.

What he’s looking to improve

  • Better meta-prompts
  • More robust spec enrichment
  • Stronger AI evaluation methods
  • Better handling of missing or implicit requirements

Notable Insight

“As much context as I can provide to the AI, that will help improve accuracy and make the process smoother and faster.”

The episode’s core message is that agentic SDLC works best when paired with a QA mindset: build with structure, verify with judgment, and assume that human oversight is still necessary—especially when quality, scale, and edge cases matter.