Mirjam Sophia Glessmer

Currently reading Nguyen-Trung & Friese (2026) “Towards a methodologically congruent framework for GenAI use in nonpositivist qualitative research”

I’ve written repeatedly about how using GenAI for qualitative analyses is a bad idea. In this paper, Nguyen-Trung & Friese (2026) present a different, very interesting and nuanced perspective. My summary below.

Nguyen-Trung & Friese (2026) agree that using GenAI to replace traditional qualitative analysis is not possible, for all the reasons we also mention, and in a very nice table where they give parts of prompts and explanations for why they cannot work. For example, prompting to “not[e] reucurring ideas” cannot work because the model cannot store anything for subsequent steps; “familiari[se] yourself” begs the question “who is ‘yourself’? Can Copilot have positionality?“, “Provide between four and eight themes” runs into the question “What if the number of themes exceeds eight? On what criteria should the Copilot prioritize?“, and many more (well worth a read!!). They conclude this exploration of critique “In sum, rather than recognizing that these outcomes stem from a prompt that was never technically feasible (Friese et al., 2026), the researchers perceived this not as a misalignment of frameworks, but as a failure of the tool.” True.

Even when GenAI is used in ways that are technically feasible, according to Nguyen-Trung & Friese (2026), often they are used to explore qualitative data without proper understanding of qualitative methods. Questions are approached as “small q” questions, i.e. approaching qualitative data descriptively and mechanically, by for example counting occurances of phrases, rather than “Big Q” questions, i.e. fully quantitative traditions like reflexive thematic analysis, or whatever other method would be appropriate in any given case. Not deeply engaging with the data by doing the coding manually also means that researchers won’t know their data well enough to determine whether any conclusions can actually be drawn from it or whether more, or different data is needed to answer their question, or if they are even asking the right questions. They also point out that for example interrater reliability, done to keep researcher bias in check, is completely misaligned with the ideas of reflexive thematic analysis where researchers are seen (and also embrace themselves as being) subjective.

As a refresher of how reflexive thematic analysis should actually be used, I at this point decided to read one of Braun & Clarke’s more recent articles (I had no idea there were so many) and found their editorial “Toward good practice in thematic analysis: Avoiding common problems and be(com)ing a knowing researcher” (Braun & Clarke, 2023), where they — on the basis of 20 recent articles in that journal — explain what to look out for.

First, Braun & Clarke (2023) point out that thematic analysis (TA) is a family of methods that shares “practices of coding and theme development; the possibility of capturing semantic and/or latent meaning, and orienting to data inductively and/or deductively; and the designation of TA as a theoretically flexible method, rather than a theoretically informed and delimited methodology“. They write “Reflexive TA approaches embrace researcher subjectivity as a resource for research (rejecting positivist notions of researcher bias, see Varpio et al., 2021), view the practice of TA as inherently subjective, emphasize researcher reflexivity, and reject the notion that coding can ever be accurate—as it is an inherently interpretative practice, and meaning is not fixed within data.

How exactly things are enacted in detail is very different between different TA methods, and mixing Big Q reflexive TA approaches with small q consensus coding, or use of codebooks, or the idea that a correct interpretation of data is possible is just methodologically incoherent: “Researchers cannot coherently be both a descriptive scientist and an interpretative artist (Finlay, 2021)“. However, Braun & Clarke (2023) point out that their goal is not to get everybody to exactly follow a fixed method, but rather be thoughtful about why they are doing what.

Another really interesting point in Braun & Clarke (2023)’s discussion is the “conceptualization of themes and whether themes are understood as a) summaries of topics or categories (what is shared and unites the observations in the theme is the topic, such as “good experiences in healthcare”); or b) capturing a core idea or meaning (what is shared and unites the observations in the theme is meaning), and the telling of an interpretative story about it.” in a nutshell, “a topic summary or a meaning-based interpretive story“. They explain that themes of the “a” category are the ones that you could come up with before looking into the data, whereas the type “b”, “themes as interpretative stories built around uniting meaning cannot be developed in advance of analysis. They contain diversity, but they have a central idea that unifies the diversity (instead of “good experiences of healthcare” you might have the theme “validation of my personhood”).” In reflexive TA, themes will be of type “b” and therefore cannot come out of a codebook or have validated interrater reliability.

Super interesting is also the point that “[l]anguage around theme development is also important to signal the researcher’s active role in generating their themes (and also to clearly signal that themes are not implicitly conceptualized as real things that exist within data prior to analysis). In reflexive TA, themes are generated, created or constructed (for example), they are not identified, found or discovered, and they definitely don’t just “emerge” from data like a fully-grown Venus arising from the sea and arriving at the shore in Botticelli’s famous painting“. And since the researchers are actively constructing, it is super important to describe their positionality, and their theoretical approach: “TA cannot be conducted in a theoretical vacuum—researchers inescapably bring in assumptions about the nature of reality, about what constitutes meaningful knowledge and knowledge production, and what their data represent or give them access to, even if these are not discussed. Ideally, the reader should not be left to detect what the researcher’s assumptions are—they would explicitly be discussed in the paper.

Braun & Clarke (2023) also suggest to use participant validation of analysis to make sure that the researchers’ interpretation represents their experience and is not biased in some way. But there is a conceptual tension with the assumption in that “that there is a truth of participants’ experiences that we can access if we can keep the potentially distorting effects of researcher influence in check (see Smith & McGannon, 2018). Reflexive TA is premised on the researcher always shaping their research; it will always be infused with their subjectivity, and they are never a neutral conduit, simply conveying a directly-accessed truth of participants’ experience.

The article ends with 10 “snappy” recommendations for researchers using TA, and a shoutout to the very helpful website https://www.thematicanalysis.net/. Phew! There is more to read about reflexive TA, but that’s for another day…

Now back to the GenAI in qualitative research paper, where they point out that interrater reliability and reflexive TA don’t work together, so “[b]y using [interrater reliability] to ‘prove’ the AI works, researchers are imposing a (post) positivist metric onto a constructivist process – a textbook definition of methodological incongruence.” So again, it is, according to Nguyen-Trung & Friese (2026), not the tool doesn’t work in itself, but that people are using it in a way that cannot possibly work in the first place. “A chatbot, when asked to ‘code’ a dataset, generates an output that looks like coding, but it does not necessarily create a coded dataset“. “In short, replicating the appearance of qualitative analysis through automating the coding process does not reproduce its essence.

But, they argue, coding was only ever useful because of the way human brains work, to make data processable for us without relying on working memory that cannot possibly hold all the data simultaneously. But GenAI works differently, so instead of using it to replace a human brain doing human-brain-style tasks, they suggest “treat[ing] the LLM as a non-human dialog partner, capable of holistic retrieval and pattern recognition“: “The AI can help surface the ‘what’ of the analysis: summaries, quotations, and contrasting perspectives. The researcher remains responsible for the ‘how’, ‘why’, and ‘so what’: contextualizing meaning, developing the explanatory account, and linking the findings back to the original data and their research purposes (e.g., representing participants’ voices).” And they remind us that “[w]e are not ‘finding’ the truth with GenAI; we are using GenAI to provoke deeper, more dialogic, and more multifaceted human construction of meaning.” And that should be ok: “As Friese et al. (2026) argues, reflexivity is not a state of cognitive isolation. We already ‘think with’ literature, diagrams, and peers. Extending this to ‘thinking with’ an LLM does not violate the core value of reflexivity, which is to critically interrogate how knowledge is made. It simply expands the network of actors involved in that making.

Nguyen-Trung & Friese (2026) then go on to point out that analysis using GenAI needs to be done ethically, responsibly and rigurously, just like any non-GenAI analysis, only involving some additional steps. Then, GenAI “retrieves relevant passages, summarizes content, and identifies surface-level patterns (topic summaries)“, but, for example, themes only become themes when the researchers have decided that they are meaningful. “Verification, reflection, and adaptation are not separate post-processing steps, but continuous activities within the interaction itself. Researchers need to assess whether an AI-generated summary, contrast, theme candidate, or interpretation is sufficiently grounded in the data, whether it fits the research question, whether it overlooks relevant tensions or negative cases, and whether it requires further interrogation.” Also, the process itself must be documented and described appropriately so it becomes traceable and understandable for others.

They summarize their two main points:

  • First, we shift from line-by-line coding and tagging segments by segments to a dialogic mode of deep thinking and questioning. This mode emphasises leveraging the actual strength of LLMs – not coding automation, but enabling a reflective, interactive conversation about the data.
  • Second, we shift toward an abductive and dialectical process of intellectual engagement with GenAI“; “This abductive, dialogic process preserves the iterative, recursive nature of qualitative work – the ‘deep immersion’ – without the mechanical burden of manual coding. The analysis evolves through conversation, not accumulation.

So, in a nutshell, instead of using GenAI to replicate how we traditionally analyse GenAI, they suggest a methodological evolution.

Phew! Done reading this article! And I think they make a lot of good points about how NOT to use GenAI (and I especially like their table with why typical prompts cannot possibly work). But I think I would be a bit more careful in embracing GenAI use in the way they suggest than they are. I kind of like the idea to use GenAI to find quotes — sometimes you have a vague recollection that someone said something, but you don’t remember the exact words, so having GenAI search for passages that say the same thing in different words could be very helpful indeed (if you didn’t code your data properly already…), or to find more passages that express a similar thought. At the same time, I would be very very careful with using GenAI to summarize anything because it cannot actually do that, and I think that there is a dangerous slippery slope where GenAI output sounds plausible enough that it is very easy to not be critical enough with it when under time pressure (and aren’t we all under time pressure most of the time?). At the same time, maybe that’s a  researcher integrity thing that we just need to be aware of and work on? Anyway, very interesting and thought-provoking article!


Braun, V., & Clarke, V. (2023). Toward good practice in thematic analysis: Avoiding common problems and be (com) ing a knowing researcher. International journal of transgender health, 24(1), 1-6.

Nguyen-Trung, K., & Friese, S. (2026). Towards a methodologically congruent framework for GenAI use in nonpositivist qualitative research. International Journal of Social Research Methodology, 1-22.


P.S.: Re: Featured image: The commune where I live posts instructions on social media for how to use and throw life buoys (important: Step on the other side of the rope and don’t throw the whole thing in the water with the buoy…) and encourages people to give it a try so they know how to use it in an emergency. That was fun! I used to be so annoyed when I saw the life buoys in the water or with the ropes out of the casing, and now I am still annoyed, but at least I know that trying them is encouraged and I just need to be annoyed about people not putting them back properly, not with touching them in the first place. I have detangled and tidied up sooo many ropes the last summers to put life buoys back in place… But then I like playing with ropes :-D

Leave a Reply

    Share this post via

    Contact me!

    Adventures in Oceanography and Teaching © 2013-2026 by Mirjam Sophia Glessmer is licensed under CC BY-NC-ND 4.0

    Search "Adventures in Teaching and Oceanography"

    Archives