Thematic Analysis: Six Phases on UX Interview Data
Thematic analysis on UX interviews: six current phases, codes vs themes, and one published-quote walkthrough. Includes saturation and common mistakes.

Thematic analysis on UX interviews: six current phases, codes vs themes, and one published-quote walkthrough. Includes saturation and common mistakes.

Thematic analysis is a method for finding patterned meaning in interview transcripts so those patterns can inform a product decision. Virginia Braun and Victoria Clarke set out the approach in 2006; they now teach six renamed phases. A theme is a shared-meaning pattern with a central concept, not the line people mentioned most.
Ranking how-tos still use the 2006 labels ("searching for themes," "producing the report"). This page teaches the current names, how coding differs from theming, and one walkthrough on published UX interview quotes.
If you already collect qualitative research, this is the named method for turning that corpus into a narrative a stakeholder can challenge.
You move from raw talk to patterned meaning. You read the dataset, code interesting stretches, then construct themes that say something about those codes as a set.
Maria Rosala at NN/g describes the UX job as tagging individual observations and quotations so themes become visible across participants. That is the practitioner version of the same method. It is not a CAQDAS product, not ChatGPT, and not the whole of UX research methods.
The official definition of a theme, from Braun and Clarke's site, is a pattern of shared meaning underpinned by a central concept or idea. "Pricing" is a topic. "First-time users do not trust the price until they see the confirmation email" is a theme, as Talkful puts the contrast.
Understanding TA splits the field into three clusters. This guide teaches reflexive TA on UX interviews. Stay inside that cluster once you pick it.
Cluster | Validity story | How you code | What a theme is |
|---|---|---|---|
Coding-reliability | Small-q; intercoder agreement | Structured codebook, dual coding | Often a topic summary |
Codebook (framework / template) | Hybrid | Pre-defined frame plus new codes | Mix of topics and meaning |
Reflexive | Big Q; subjectivity as a resource | Researcher constructs codes across two or more rounds | Shared meaning plus a central concept |
NN/g's "invite others / is the theme saturated" checks, and in-session lists of predefined themes, are codebook or reliability moves. They are legitimate on those projects. They are not proof you did reflexive TA.
CASRAI notes that the literature treats kappa of about 0.61-0.80 as "substantial" and above 0.80 as "near-perfect." Reflexive TA explicitly rejects treating a kappa or alpha score as evidence of coding quality.
Product teams still ship Sunday-night Miro boards with labels decided before the last interview. That is a confirmation deck. It is not analysis.
Braun and Clarke's practical guide (Sage, 2022) is the current book-length statement of reflexive TA. The 2006 paper remains the origin citation. Cite 2006 for history; practice the 2019-2022 update.
On r/UXResearch, throughput is the live complaint: coding one session can take an afternoon, and a 300-sticky board turns to soup when the room is watching. The method's answer is not a faster tool. It is two or more coding rounds, a negative-case review, and a write-up that can show its work.
JMIR 2024 is the other 2026 pressure. ChatGPT and a human each identified five themes on a Diabetes UK podcast transcript, and the sets did not match. Treat the model as a junior assistant for candidate codes, not as the owner of the theme list.
Doing reflexive TA lists six phases and states that the names changed since 2006. The phases do not prescribe a rigid ladder. You recurse: a weak theme name sends you back to the extracts.
Phase | Current name (use this) | 2006 name still circulating |
|---|---|---|
1 | Familiarising yourself with the dataset | Familiarization |
2 | Coding | Generating initial codes |
3 | Generating initial themes | Searching for themes |
4 | Developing and reviewing themes | Reviewing themes |
5 | Refining, defining and naming themes | Defining and naming themes |
6 | Writing up | Producing the report |
Skip five-step condensations that merge define and write. Do not invent a seventh step. Keyword-to-conceptual-model recipes are a different school with similar nouns.
Listen to the audio once without coding. Read the transcript as a document, not as a quote mine.
Talkful puts this on Tuesday: one pass through the audio, no codes. The point is to hear hesitation, sarcasm, and the joke that will not survive a spreadsheet cell.
u/malrat72 in r/PhD (June 2026) described the same job as organizing a messy closet rather than hunting a needle: re-read until the content feels already known.
In the same thread, u/commentspanda could not run background audio while coding, and aimed at 1-2 interviews a day over a month.
If your transcripts are free-flowing, do not cluster later codes by interview question number. u/GalwayGirlOnTheRun23 in r/AskAcademia (March 2026) tried that and called it meaningless. Topics land in random places, so code the idea, not the prompt.
A code marks a stretch of text that matters for your research question. Keep it specific and single-idea. Do at least two rounds, then collate the extracts under each code.
The deliverable here is retrieval tags on real lines, not a theme list. Keep codes close to the participant. "Gave up and emailed support" beats "usability."
You can write a more interpretive code in round two. You cannot recover a quote you never tagged.
Cluster codes that share a central concept. You are generating candidate themes, not discovering objects that were waiting in the file.
Virginia Braun, in the Sage webinar (27:30), is blunt about n=1:
"I probably wouldn't report that as a theme because it's only evident in one particular participant's experience, but that doesn't mean it should be erased, cleaned up from the data set and not mentioned."
Demote the striking one-off. Park it for the write-up. Do not inflate it into "the theme" because it will play well in a slide.
Talkful reports a week in which 287 codes became seven candidates, then two single-voice candidates were demoted. That 3-7 range is a product-research heuristic, not a Braun rule.
A candidate has to survive its extracts, then the whole dataset. Look for the negative case: the interview that should fit and does not.
A Looppanel post (July 2026) describes a researcher defending one quote for an hour. The summary had logged "fine" as positive. She had been in the room, heard the flat tone and the pause, and knew "fine" meant they had given up months ago, so she recoded.
This is also where a topic bucket dies. If the only thing the extracts share is the word "onboarding," you have a heading, not a theme. Ask what claim the heading is making; if you cannot finish the sentence, collapse or split.
Write a short definition for each surviving theme: the central concept, what belongs, what does not, and why it matters for the product.
Names should be specific. "Trust" is a bucket; "Trust collapse at price change" is a theme name you can test. Clarke's anti-example is "barriers to exercise" (a dump) versus "exercise is costly" (a claim).
Inside this phase, a thematic map is the artifact that shows how themes relate. GIS results for "thematic map" are a different field. Do not hunt a cartography tutorial.
Clarke's lecture heuristic: analyses with 28 themes are over-fragmented, and more than about six themes in an 8,000-10,000-word write-up starts to go thin. Treat that as a density check, not a quota.
Each theme gets evidence, not a vibe. Verbatim quotes with timestamps beat paraphrases. Lieven Aerts (UX researcher, ING Belgium) keeps every quote in the language it was said, because translation strips nuance.
Report in English if you must. Do not "clean" the evidence.
On r/UXResearch the useful deck rule is one slide per theme: the quote as the visual, the insight, then the so-what for the product. Talkful ends the week on a six-slide deck after naming, with 4-6 quotes per theme.
NN/g's time rule still applies. If collection took two weeks, analysis that fits in Thursday afternoon is the tell that themes were pre-decided.
CASRAI (updated 24 August 2026): "A code is not a finding and it is not yet a theme. It is a retrieval tag." On the same page: a code describes; a theme claims something analytically. Confusing the two is one of the most common weaknesses reviewers flag.
Clarke, in this lecture (37:26):
"A code is a label that captures something that's interesting in the data. It really is that simple. We'd avoid one-word code names because you're trying to get deeper into the data and it's unlikely that one word like stigma or gender is really getting in there."
A code is generally single-facet (a brick). A theme is multi-facet and made of many codes (the wall).
You can code at two levels on the same line. Descriptive codes stay close to the event. Interpretive codes name the meaning you are willing to defend.
NN/g's UX-careers interview fragment uses how skills are acquired (descriptive) versus self-reflection (interpretive). Maze (Louise Jones, 2026) publishes a fuller pair:
Quote | Descriptive code | Interpretive code |
|---|---|---|
"I never know where to find anything in the app. I just give up and email support." | Navigation difficulty | Loss of trust in product discoverability |
The richer theme name is not "Navigation." It is "users lose confidence and abandon tasks when they cannot reliably find key actions."
Do two rounds: round one harvests, round two visits the marks with the whole dataset in view. Looppanel's LinkedIn line: synthesis is two jobs, pile-sort then judgment.
Reflexive TA on UX interviews is inductive and experiential: codes come from the dataset, not from a stakeholder's priority list. Deductive coding from a pre-defined theme list is codebook TA. Hybrid is allowed if you stay theoretically congruent; do not claim reflexivity and then freeze the frame in week one.
Semantic codes stay with explicit meaning. Latent codes go after implicit meaning. Clarke's correction: latent does not mean Freudian unconscious; it means the assumption under the sentence ("fine" as surrender).
If the team needs a shared frame, write codebook fields: name, definition, inclusion, exclusion, exemplar quote. Version that file. That is the codebook-TA aside, not the reflexive default.
Indexing passages so you can retrieve them is not the six-phase method. Indexing is phase 2. The method continues through write-up.
Most worked examples never walk a UX interview through all six phases. Student PDFs stop at stage 2, and health-education papers (Byrne's 2022 illustration is the canonical method walkthrough) are the wrong corpus for this audience. The lines below are published fragments, not a transcript invented for this page.
Familiarisation. You have eight onboarding interviews for a B2B product. Tuesday is audio only. You notice people talk around price without naming it.
A later Looppanel write-up records that pattern: nobody said the word pricing; they said budget; they would have to justify it internally; one person just laughed. You write a memo. You do not code yet.
Coding, two rounds. On the Maze line, round one gets navigation difficulty. Round two gets loss of trust in product discoverability. A second published shape, from Talkful's teaching contrast, would tag waits for confirmation email rather than pricing.
No one-word codes. After the full set you might hold a few hundred codes.
Generating initial themes. Cluster those codes around a central concept: trust does not attach to the price point; it attaches to whether the product will still be there after checkout. A single vivid quote about a dark-pattern cancel flow is not this theme; park it. Braun's rule: n=1 is not a theme, and it is not trash.
Developing and reviewing. The negative case is the participant who called the plan "fine." The spreadsheet logged positive sentiment; the room-tone was flat, so recode. In Talkful's published week, review collapsed the candidates to four themes.
Naming. Discard "Trust" and "Pricing."
Keep a name that states the claim, such as Trust collapse at price change. Write the definition, the exclusion (billing-page copy bugs that never touch trust), and the product so-what (show the confirmation path before the paywall).
Writing up. One slide, one theme. Put the verbatim line on the slide, with the timestamp, then the insight and the change you are asking for. If research ops will store this in a repo, keep the quote in the original language and keep the link back to the session.
That is the method on real interview material. A spreadsheet whose final headers are Customization, Data Usage, Data Stories, Current Product is a topic sort, not this walkthrough.
Affinity mapping (KJ Method) is early, visual, collaborative, and fast. Sticky notes are the unit. You use it in a workshop to surface clusters while the room is still in discovery.
Thematic analysis starts after collection. The unit is a coded extract from a transcript. It is slower, often solo, and the output is a defensible narrative rather than a wall.
Combine them when you mean both: affinity to organize the room, then coding to interpret the corpus. NN/g lists affinity-diagramming as a technique you might use inside analysis. That is a visual habit, not the six phases.
u/poodleface in r/UXResearch (July 2026) put the failure mode in one line: affinity mapping shreds the context of the quotes, and the words alone can reverse the intent. In that same thread, a 14-interview onboarding study redone as a mind map centered on abandonment surfaced an admin-versus-end-user dependency the affinity wall missed.
Card sorting belongs with information architecture, not with this method.
Content analysis scores presence and frequency, often against a predetermined attribute list. Thematic analysis interprets patterned meaning. Vaismoradi 2013 already noted that the boundaries have not been clearly specified and the two are often used interchangeably, so a complaint-frequency chart is already brushing content analysis.
Tanner Kohler (26 February 2024): counting how frequently something is mentioned will not always show you what is most important, and sometimes the most valuable insight is mentioned once. You can combine the two methods. Do not let the bar chart pick the roadmap.
Grounded theory is theory-building with theoretical sampling. One contrast is enough: you do not sample until a theory saturates. Open coding, as a search query, is a different job, so name it and leave it.
Atomic research (experiments → facts → insights → recommendations) is a storage and synthesis pattern across studies. It is adjacent to research synthesis, not a replacement for coding one study's interviews.
You can code in a spreadsheet. Cloud CAQDAS is optional: Delve lists $50/user/mo, ATLAS.ti joined NVivo under Lumivero after the 12 September 2024 acquisition, and MAXQDA remains independent.
A UX repo such as the one in this Dovetail review is another slot. Then return to the method. Software is not a phase.
Product teams say they hit saturation at 8. That sentence almost always imports a codebook study into a reflexive project. Report the disagreement; do not pick a magic n.
Source | What they measured | n |
|---|---|---|
Guest, Bunce and Johnson 2006 (via Hennink 2017) | Theme presence, 60 West African interviews | Basic elements at 6; codebook stable by 12 (88% of themes, 97% of important themes) |
25 in-depth interviews; code vs meaning | Code saturation at 9 ("heard it all"); meaning saturation 16-24 ("understand it all") | |
Review of empirical saturation tests | Average 12-13 interviews for code saturation; outliers 20-40 | |
Base size + run length + new-information threshold | No single n. Homogenous datasets | |
Argument, not an experiment | Avoid saturation claims in reflexive TA. Saturation-n papers used coding-reliability assumptions |
Hennink is Monique. Guest 2006 is not Guest 2020.
The 2021 Braun and Clarke paper is the match for reflexive work. Claims about saturation do not translate to reflexive TA. Use information power (sample adequacy for this question) and show your decisions, including the mess.
Usability testing has its own sample-size folklore. Do not paste it into interview theming.
Braun and Clarke flagged these when they reviewed papers that cited 2006 while drifting off the method. UX practitioners keep repeating the same ones.
Clarke, in the Sage webinar (22:35), calls this a bucket theme: slap a label on a bucket, chuck the data in, pour it onto the page. "Onboarding," "Navigation," "Trust" are buckets. A theme has a central concept and a story.
Clarke, in the introduction (27:07):
"Please don't say they emerge. Please. We get grayer every time we read something that says themes emerge."
Themes are not sitting in the data waiting to be scooped out. You construct them. The active verb is the method.
Using 2006 as the reference while dual-coding to a kappa target, or while teaching "searching for themes" as a hunt, is methodological incoherence. If you ran codebook TA, say so. If you ran reflexive TA, use the current names and the 2019-2022 papers.
The "fine" recode is the review phase. Shipping a single dramatic quote as the finding is the generating-themes failure. Both make the deck more confident than the dataset.
Kohler's frequency warning belongs here, and so does NN/g's budget rule. A 40-interview study analyzed on Friday is a pre-decided story. Confirmation decks and Sunday-night Miro with themes chosen in the kickoff are the same bug with nicer stickers.

Statistical significance means a result would be rare if nothing changed. P-values, 0.05, practical vs statistical significance, and peeking's false positives.

The Kano model classifies how customers feel if a feature is present versus absent. Satisfaction does not rise in a straight line with more functionality.

UX ResearchOps as four systems: first-party panels, consent ops, repositories, and democratization guardrails. NN/g data and a GitLab SOP.