Going from raw interview transcripts to a credible set of themes is a structured process, not an intuitive leap. Committees want to see the steps, initial codes, how they grouped into categories, how categories became themes, not just a polished final list that appears from nowhere.
| Stage | What Happens |
|---|---|
| Familiarization | Reading and re-reading transcripts before any coding begins |
| Open/initial coding | Labeling meaningful segments with descriptive codes, close to the data |
| Axial/focused coding | Grouping related initial codes into broader categories |
| Theme development | Combining categories into overarching themes that answer your research questions |
| Refinement | Checking themes against the full dataset, revising boundaries as needed |
Open coding (sometimes called initial or first-cycle coding) is the deliberately unglamorous first pass through a transcript. The goal at this stage is not to be clever. It's to label meaningful segments of text with short, descriptive tags that stay as close as possible to the participant's own words and meaning, resisting the urge to jump straight to abstract, theory-laden labels. A participant describing "staying late most nights because there's no one else to cover" might get coded simply as working beyond scheduled hours rather than immediately labeled burnout. The more abstract, interpretive label comes later, once the pattern across many participants earns it. Open coding is intentionally messy and produces far more codes than will survive to the final theme structure; a single 45-minute interview transcript can easily generate 40-60 initial codes, most of which will later be merged, renamed, or discarded as redundant. That messiness is a feature, not a failure. A coding pass that produces a tidy, small set of codes on the first try is usually a sign the researcher skipped ahead to conclusions rather than genuinely sitting with the data.
Where open coding fragments the data into many small pieces, axial coding puts some of it back together by asking how the initial codes relate to one another. This is the stage where codes like working beyond scheduled hours, skipping breaks to finish charting, and coming in on days off to catch up get grouped into a broader category. Something like absorbing unaccounted-for workload, because they all describe variations on the same underlying phenomenon. Axial coding also surfaces relationships between categories: does one category function as a condition that produces another, a consequence that follows from another, or a strategy participants use to cope with another? Mapping these relationships, even informally, is what eventually lets a set of categories become a genuine theme rather than just a longer list of labels. It is normal for this stage to require several passes back through the transcripts, since a category that looked solid in isolation sometimes turns out to be two distinct things once compared across the full dataset.
A codebook is the single document that makes a coding process auditable rather than something the researcher simply has to be trusted about. At minimum, it should record, for every code: a clear working definition, the criteria for when a segment does or doesn't qualify for that code, and at least one or two example excerpts. As coding progresses, the codebook needs to be a living document. Codes get renamed for clarity, split when they turn out to cover two distinct ideas, or merged when two labels turn out to describe the same underlying concept. We keep a version history of the codebook so the evolution of the coding scheme is itself part of the audit trail: a committee member asking "how did you decide code X applied here and not there?" should be answerable by pointing directly at the codebook entry, not by re-litigating the decision from memory.
Don't force data into a predetermined framework. A common credibility problem is coding transcripts to match themes the researcher expected to find rather than what the data actually shows. If your final themes look suspiciously identical to your literature review's expectations, that's worth a second, more skeptical pass.
Saturation is the point at which continuing to code additional transcripts stops producing new codes or meaningfully refining existing categories. It is a judgment call, not a number you can calculate, which is exactly why it needs to be described carefully rather than simply asserted. In practice, we track saturation by noting, transcript by transcript, whether a new interview contributed a genuinely new code, merely reinforced an existing one, or added a variation worth capturing as a sub-code. When several consecutive transcripts stop contributing anything new, saturation has plausibly been reached. And that pattern, not just a stated sample size, is what belongs in the methodology or results chapter as evidence. A common mistake is treating saturation as automatic once a target number of interviews (10, 15, 20) is hit; the number is a planning estimate, not proof, and a defensible saturation claim describes the pattern of diminishing new codes across the actual data collected.
Take a qualitative study of nine hospice nurses discussing moral distress. Open coding on the first three transcripts produces codes like disagreeing with a family's care decision, following an order I thought was wrong, and staying quiet to avoid conflict with a physician. Axial coding groups these into a category tentatively called silenced disagreement. By the sixth transcript, no genuinely new variation has appeared within that category for two interviews running, though a related but distinct category: disagreement voiced but overridden. Has emerged and needs its own definition in the codebook. By the ninth and final transcript, both categories are well saturated, and together with two other categories they combine into the overarching theme navigating powerlessness in end-of-life decisions. Every step of that path, from individual code, to category, to theme, is traceable back to the codebook and the transcript excerpts that generated it, which is exactly what a committee expects to be able to follow.
At the master's level, a coherent set of themes is often enough. At the doctoral level, committees expect to see, and sometimes ask to review directly, the codebook itself, the progression from open to axial coding, and an explicit account of saturation, along with a stated position on researcher reflexivity (how the researcher's own perspective might have shaped which codes felt salient). Doctoral qualitative work is also increasingly expected to address trustworthiness explicitly: credibility (does the coding plausibly represent participants' meaning), dependability (would another careful coder follow a similar process from this codebook), and confirmability (is the audit trail clear enough that conclusions can be traced back to data rather than to the researcher's assumptions). We build all three considerations into the coding process itself rather than trying to retrofit them into the write-up afterward.
A traceable path from transcript to theme, ready for committee scrutiny.
Whichever fits your access and preference. We work in NVivo or Atlas.ti when available, or build manual coding tables that document the same logic just as rigorously.
There's no fixed number. It depends entirely on what the data supports. A forced count (too many or too few) is a red flag; the right number is whatever your data genuinely produces.
Yes, this is one of our most common qualitative requests. Send the transcripts and we build the full coding process from there.
A code is the smallest label attached to a meaningful segment of text. Categories group related codes together. A theme is a broader pattern, built from one or more categories, that directly answers a research question. Themes are what actually appear in your findings chapter, while codes and categories are the scaffolding that gets you there.
Yes, when a second coder is available. We calculate a simple percentage-agreement or Cohen's kappa figure on a subset of the data and report it alongside the coding process, which strengthens the credibility of the final theme structure.