Designing agentic AI for the right moment

how research redirected an AI product roadmap 


 

TL;DR:

Role: Lead UX Researcher

Product: IBM Test Accelerator for Z — agentic AI test generation for mainframe development

Timeline: June - October 2026

Methods:

  • End-to-end workflow mapping

  • Six moderated interviews

  • Cross-product research alignment

Outcome: Research findings shifted the agentic workflow integration from Functional Testing to Unit Testing before launch, redirecting the roadmap toward the users with the highest need and established a cross-product research alignment model across IBM's testing portfolio.

 


Research Context:

IBM was investing in agentic AI across its development tooling, specifically features where an AI agent doesn't just assist but takes on work autonomously.

Test Accelerator for Z was scoped to bring agentic test generation to mainframe development by generating test cases that developers would review and run.

The real risk wasn’t that the technology wouldn't work, but whether developers would actually adopt it. Agentic features often fail subtly: when users lack trust in the agent’s output, or it appears at an inconvenient point in the workflow, adoption falters, and the investment yields a feature that no one relies on.

The roadmap had scoped the agentic workflow to the Functional Testing stage. Whether that was the right moment in the developer's workflow was an assumption, not a finding.

As the lead User Researcher, I didn’t want to answer "Is the feature usable?" but rather, “When in a developer's workflow is an agentic AI feature most likely to be adopted, and are we designing for that moment?”

Research phase 1

Workflow mapping with the development team to establish the baseline developer journey and locate candidate moments for agentic intervention

To identify the areas of opportunity, I needed to understand the workflow before the feature entered it. I partnered with the development team to map the existing path end-to-end, from writing code to executing tests, identifying where agentic assistance had genuine opportunity versus where existing tools already covered the job.

Research Phase 2

Six moderated interviews with developer/tester participant description mix of experienced mainframe developers and early-career engineers. Our conversations focused on how they currently generate, select, and trust test cases

Research Phase 3

Cross-product alignment: I identified a parallel IBM team building a related testing experience and initiated shared research goals across the two efforts



Findings

Finding 1: USEr Need - Broadening Scope Beyond Functional Testing to Early Development Testing

Enthusiasm for Agentic-assisted test case generation, given that existing processes remain manual.

Impact: This approach aimed to accelerate the progress of early-career testers who may lack familiarity with Cobol or existing test cases

 

Finding 2: User Value - Test Case Management was a key value driver

The ability to directly review and modify test cases before bulk generation was significant during the initial adoption phase.

Impact: Allow users to manage their test at all stages of the life cycle, increasing likelihood of user adoption

 

Finding 3 : cross-product fragmentation

During the research, I identified a parallel product team developing a similar testing experience. Left unaligned, the two efforts would have shipped fragmented, inconsistent experiences across IBM's testing portfolio. This kind of seam that users encounter organizations don't see. I initiated a cross-team alignment process with shared research goals.

Decision this drove: The alignment created a new collaborative model for testing experiences across product lines, with research operating at the portfolio level, not the product level.


Outcomes

  1. Agentic workflow integration shifted from Functional Testing to Unit Testing before launch, redirecting the roadmap toward the highest-need user segment

  2. Cross-product research alignment model established across IBM's testing experiences via joint sessions across our design and development teams on an ongoing cadence.

Reflection

What I’ve learned from this work is that agentic AI features tend to fail not because of limitations in the model, but because of timing. The important question isn’t whether the agent can perform the task, but whether it appears in the workflow at the exact moment when the user's need outweighs their doubt. For early-career developers confronted with unfamiliar COBOL test cases, that critical moment existed. For seasoned developers at the functional testing stage, it did not.

That's a research question, not an engineering one, and it's the kind of question I think every team shipping agentic AI needs answered before launch, not after.