AI can speed up CRO research. But which parts should you actually trust it with?
That’s the question we set out to answer in last week’s AI Native Marketer workshop, hosted by CXL President Hesh Fekry and CRO consultant Ruben de Boer, who has spent more than 15 years running experimentation programs.
The goal was to move beyond the theory and see what happens when you actually apply AI to a proven CRO research process. We built an AI-powered workflow that automates three parts of ResearchXL: heuristic review, thematic analysis, and test idea generation.
Some parts worked surprisingly well. Others showed exactly where human judgment still matters.
Table of contents
How we mapped the workflow before touching an agent
The starting point was simple: don’t start with the AI. Start with the process.

We used a swimlane to map each step, the people involved, the tools they use, and roughly how long each task takes. From there, we separated the steps that require human judgment from the operational work AI can take over.
Once the workflow is visible, it becomes much easier to turn those individual steps into skills and commands for an agent.
What we built
Three chained pieces, built to mirror three stages of ResearchXL:

- Heuristic review: Audits a user journey against UX heuristics (relevance, clarity, friction, motivation) and produces both an annotated audit and a scoring matrix.
- Thematic analysis: Takes research inputs from any method (heuristic review, surveys, analytics exports, session recordings) and groups the findings into themes, showing which sources back each theme.
- Generate test ideas: Turns those themes into Spearo’s five buckets: just-do-it, instrument, test, hypothesize, investigate, with a PXL-style priority score attached.
Why start with heuristic review specifically
In ResearchXL, heuristic review sits at the bottom of the evidence hierarchy: it’s usually one or two people, it lacks data volume, and it carries obvious individual bias. That makes it the lowest-risk research method to hand to AI.

Synthesized interviews or anything that replaces real human feedback stays off the table for now.
The speed difference is real. A human heuristic workshop takes hours to set up and run. The agent went through five full journeys in about 30 minutes.
AI scores look more objective than they are
The faster review is useful, but Ruben’s pushback points to a bigger lesson: we need to be careful about how much we trust AI when it scores or prioritizes research.
The AI gives very specific scores, like 42% or 53%, which makes the results look more scientific than they really are. Underneath, they’re still based on subjective judgment, not actual measurement.
AI also doesn’t really know what motivates a user. It can check for things like urgency, social proof, or authority, but it can’t know whether those things actually matter to a specific visitor. Something that looks like a problem according to a framework might have little impact on conversions.
The biggest weakness is prioritization. AI is useful for finding and organizing potential issues, but deciding what actually matters most still requires human judgment. Treat AI-generated scores as a starting point, not the final answer, and validate them against real user behavior and test results. facts. Use them as a starting point, then validate them with real user behavior and test results.
Turning research into themes
The second part of the AI workflow moves beyond individual observations.
It takes inputs from different research methods, including heuristic reviews, surveys, analytics, copy testing, heatmaps, and more, and groups the findings into common themes.
Instead of ending up with hundreds of disconnected observations, you can start seeing patterns such as:
Theme → supporting research sources → underlying customer problem
And when the same theme is supported by several independent research methods, you have more evidence behind it.
This is also where the output becomes more strategically useful.
Instead of telling stakeholders “we should add social proof,” you can discuss the larger customer problem your experiments are trying to solve.
Turning those themes into test ideas
The final part of this workflow takes those themes and turns them into potential experiments and next actions.
But there’s another AI failure mode worth watching for here: research as decoration.
As Ruben described it, sometimes AI comes up with an answer or best practice first and then searches the research for evidence that supports it.
That’s backwards.
The research should lead to the insight, which should lead to the hypothesis, which should lead to the test idea, not the other way around.
The human shouldn’t disappear from the workflow
The goal isn’t to build one giant agent, give it all your research, and blindly accept whatever comes out.
A better workflow is:
- After the heuristic review, add your own research and context.
- After thematic analysis, review whether the themes actually make sense.
- Before generating test ideas, add anything the AI missed.
- Then review and prioritize the final backlog yourself.
Breaking the workflow into stages makes it easier to catch mistakes early rather than discovering at the end that the entire analysis was built on a bad assumption.
Ruben framed the AI workflow as something closer to another member of your brainstorming session. It can work quickly, surface ideas you hadn’t considered, and challenge your thinking, but it shouldn’t replace the diverse human perspectives around the table.
The main takeaway
AI can dramatically speed up CRO research workflows. It can review more journeys, organize more research, surface patterns, and generate test ideas much faster than doing everything manually.
But faster analysis isn’t the same as better judgment.
The strongest workflow combines AI’s speed and scale with human context, expertise, and quality control.
The bigger lesson is that the better your process and inputs, the better AI can amplify them. If the underlying process is bad, you’re simply automating bad work faster.
What’s next?
This workshop is part of our new AI Native Marketer program which was built because marketing is changing fast. There’s more pressure than ever to ship faster, wear more hats, and deliver better results, often with the same team or even fewer resources.
Throughout the program, you’ll build practical AI systems you can immediately apply to your own roles. If you’d like to continue building with us, we have an upcoming live workshop coming up in September:
Build your team AI operating systems: Learn how to build a shared AI operating system for your team, with the right structure, workflows, context, and governance to make AI actually work across the organization.
This is a two-session workshop, with 90-minute sessions on September 7 and 19.
See you there.