Moderated or unmoderated, on a prototype or on what’s already live. AI handles transcription and the first pass of synthesis now, so the effort goes into designing the right tasks and interpreting what actually happened.
A design is about to be built and the team is arguing from opinion. A flow tests fine internally and fails with real users. Or something is live, the metric is poor, and nobody can see why from the data alone.
Small numbers, done properly. Five to eight participants per round surfaces most of what's badly wrong; the instinct to recruit thirty usually delays the study past the point where it can change anything.
Recruiting is where studies are won or lost. Testing with people who resemble your users is the whole basis of the finding, and a convenient sample produces a comfortable and useless result.
I run moderated sessions when the reasoning matters and unmoderated when the task is simple and the volume helps. Either way, tasks get written to avoid leading, and findings get separated from recommendations — what happened is not the same as what to do about it.
Two to three weeks per round including recruitment. Recruitment is usually the long pole.
Is five participants really enough?
For finding serious usability problems, largely yes. For measuring anything, no — that's a different kind of study.
Can you test a prototype rather than a build?
Usually better. Testing before the code exists is when findings are cheapest to act on.