I Built an AI Agent to Do UX Research. Here's What It Couldn't Do.
AI won’t shrink research. It will split it.
It started with a belief
Every website problem I’ve encountered — across two decades of UX strategy — traces back to the same root. The design isn’t the problem. The problem is that nobody was clear about what the website was supposed to do for the business in the first place.
So I built something to test that belief.
The target was a specific type of company: Taiwan’s “hidden champions” — small-mid-sized B2B manufacturers that dominate niche global markets but rarely get seen. Exceptional at making things. Often invisible online. Websites built to exist, not to convert.
I wanted to build an AI agent that could audit these sites. Scoped tightly. Purposefully narrow. And eventually, self-service.
The first design decision wasn’t about AI at all
Before the agent could do anything, I needed to define what it was auditing — and why.
So I designed an intake form. A set of questions built from years of stakeholder interviews, translated into a Google Form that clients fill out themselves before anything begins. What’s your primary conversion goal? What pages matter most? What problems have you already noticed?
It sounds simple. But that form is doing something that took me years to learn how to do in person — scope the audit to the actual business problem, not just the visible symptoms on screen.
The AI doesn’t decide what to look at. The human asking for help decides. The form just makes that process structured, scalable, and honest.
That was the first thing I learned. The thinking that has to happen before research begins — that part didn’t get easier with AI. It got more important.
Training it felt like onboarding a junior researcher
Not because AI is human. But because the process was almost identical.
First, frameworks. Then references — the right ones, chosen deliberately. Then clear guidelines for what a good output looks like, what the report is for, and what it isn’t for.
Then I ran it. Watched what came out. Adjusted. Ran it again.
Anyone who has trained junior UX researchers will recognise this loop immediately.
The difference was the timeline.
In my experience, a new UXer needs at least six months before they can produce the kind of audit report this agent was generating after two weeks of core development and two weeks of pilot testing. And that’s with consistent training — same method, same guidance, same feedback.
The outcome with humans is still unpredictable. Some are ready in six months. Some still need close supervision at two years. The same input produces very different people.
The agent doesn’t have that problem. Within its defined scope, the consistency is something a human team simply can’t replicate.
But then I hit the boundary
Not a failure. More like a wall you walk into slowly, only realising it’s there when you stop moving forward.
The system I built isn’t just an AI generating a report. It’s a hybrid — AI analysis first, then expert review, correction, and supplementation, then a weighted integration into a final recommendation. The human layer isn’t optional. It’s structural.
And that structure exists for a reason.
When a site sat comfortably within the defined scope — B2B manufacturer, clear conversion goal, overseas market — the outputs were genuinely good. Specific findings, prioritised by severity, tied to business impact. The kind of report that takes a junior researcher months to learn to write.
But the moment context got ambiguous — a company that didn’t quite fit the profile, a business model that was harder to read, a finding that pointed somewhere the framework hadn’t anticipated — the AI kept moving. Confidently. Without flagging that it might be on uncertain ground.
It was producing findings without sensing what was actually at stake.
That’s not a technology problem. That’s a different kind of work entirely.
Research won’t shrink. It will split
Here’s the frame I keep coming back to.
There are two layers to research work. There always have been — we just never needed to separate them because humans were doing both.
The execution layer: applying frameworks, recognising patterns, maintaining consistency, producing structured outputs from defined inputs. Within clear scope, AI does this faster, more consistently, and more tirelessly than any human researcher. That’s not a prediction. I watched it happen.
The thinking layer: deciding what question is actually worth asking, defining scope before the work begins, sensing when a finding is pointing somewhere unexpected, interpreting results in context that no framework fully captures. This is where human judgment doesn’t just survive — it becomes the thing the whole system depends on.
The researchers who will struggle are the ones whose value lived entirely in execution.
The researchers who will matter more are the ones who were always doing the harder thinking — they just spent too much time buried in tasks that AI can now handle.
And yes — AI is moving fast. The boundary between these two layers will shift. What counts as “thinking work” today may not be in three years. I don’t know exactly where that line will move.
But I don’t think it disappears. I think it migrates.
The risk isn’t replacement. It’s acceleration in the wrong direction
AI is not cheap. Building this system took real investment — in time, in iteration, in the kind of domain knowledge that took twenty years to accumulate. Anyone who tells you a capable AI research system is easy to spin up hasn’t actually tried.
But here’s what worries me more than the cost.
As these tools get cheaper and more accessible, teams will use them to do research faster. That part is straightforward.
What’s harder — and what I don’t think the industry is talking about honestly enough — is that speed doesn’t fix a bad question. It just gets you to the wrong answer more efficiently.
The intake form I built wasn’t a technical feature. It was a forcing function — a way to make sure the business problem was actually defined before the AI touched anything. Without it, the agent would still run. It would still produce something that looked like a report.
It just wouldn’t be answering the right question.
That’s the real risk. Not AI doing research instead of humans. AI doing faster research for teams that never stopped to ask whether they were researching the right thing in the first place.
I completed the pilot. The agent works.
And the thing it taught me most clearly had nothing to do with AI.
It taught me that the hardest part of research — defining the right question, scoping to the real problem, knowing what your findings can and can’t tell you — that part doesn’t get automated.
It just gets more exposed.