Category Archives: ai
In April 2026, the Association of Internet Researchers organized an online roundtable on methods (if you are a member, you can watch it back here) with Annette N. Markham, Axel Bruns, and myself. This is the transcript of my opening statement.
Thank you, and thank you to the organizers for putting this roundtable together. We’re here to discuss the how Internet research methods are evolving, and I want to use my ten minutes to talk about AI as part of our analytical setups and take the quantitative-qualitative divide as an entry point.
My intuition – and I offer it as a discussion point, not a settled thesis – is that AI is reshuffling that divide in ways that are both complex and evolving. As computational methods gain semantic capabilities, they take on an increasingly qualitative character, I want to argue, and this demands new vocabularies for methodological decision-making and accountability.
I’m tempted to call this a “post-quantitative” turn, but I don’t think this is the right term, so I want to explain more specifically what I mean and where this leaves us methodologically.
I think that for a long time, the dominant framing in internet research and elsewhere framed “computational methods” as synonymous with quantitative analysis. Scraping, counting, plotting. Running corpora through pipelines that produce frequencies, network graphs, and co-occurrence matrices. Quantitative tools, quantitative outputs, quantitative epistemological assumptions.
But this was of course never the whole story. Those of us working in the digital methods space and other traditions that draw on STS and media studies have always had a more entangled practice. We have argued for a long time that computational tools are not neutral instruments but that they shape what becomes visible and how we create knowledge. There has been a quali-quanti sensibility running through these traditions, where we treat algorithmic outputs as inscriptions that require reading, and this was already troubling the quantitative-qualitative divide before anyone had ever heard of a transformer.
So when I say that something has shifted, I don’t mean we’ve moved from a pure quant world to a blurred one. The blurring was always there, but it lived in the researcher’s practice. The tools themselves were still largely quantitative.
What transformer architectures have changed is that the tools we can use for analysis now have much deeper semantic capabilities. They handle meaning, context, ambiguity – and with that, qualitative judgment has migrated into the computational process to a much higher degree.
And I think that the more semantically capable the tools become, the more qualitative the work of using them becomes. To show what I mean, I want to talk about three distinct ways researchers now integrate AI and transformer models into their analytical process, three options (and there are more!) where this qualitative character increases progressively.
Here is the first option: The deductive approach, or “the prompt as codebook”, where a researcher working with fifty thousand social media posts about a public health crisis writes a detailed prompt defining five categories of vaccine stance. They classify every post, they report inter-coder reliability against a human-coded sample. This is recognizably confirmatory and quantitative in its logic, even though the mechanism underneath is a language model performing interpretation. The epistemological challenge here is real, especially if you consider the importance of model choice and prompt, but also kind of manageable. You can audit the model’s classifications against your own, and the framework of content analysis gives you established criteria for validation. The interpretive labor is delegated, but it is also tightly constrained. That is the first option.
The second option moves us deeper. This is what you could call the “controlled pipeline” – and you can see it quite clearly if you consider how topic modeling has evolved. With LDA, the dominant algorithm for “traditional” topic modeling, the pipeline is relatively stable: you choose a number of topics, you run the algorithm, you get word distributions, and then the qualitative moment arrives: you decided what to call topic 7. Now, take something like BERTopic applied to those same social media posts. BERTopic is itself built on a transformer architecture. And the pipeline has become a cascade of consequential, human choices: which sentence transformer you embed with; how you parameterize the dimensionality reduction; which clustering algorithm you select; how you handle outliers. Each choice shapes the output, and none has a strictly “correct” setting. Using something like BERTopic requires a lot of judgment, familiarity with the data, and iterative exploration. Qualitative decision-making is present at every step, not just at the end. But I think that we can say that, these choices are still our choices, they are visible, adjustable, and accountable to a large degree. You can trace why you got the output you got. Rigor here means documenting the full pipeline and demonstrating sensitivity to parameter choices. It’s demanding, but the researcher remains in largely in control.
Which brings us to the third option: The inductive approach. Open-ended prompting where the researcher delegates not just classification, but more involved forms of interpretation. Imagine the prompt: “Read these posts and identify the arguments and rhetorical strategies being used.” The researcher then iteratively refines what the model surfaces, treating outputs as interpretive proposals rather than data points. This is recognizably qualitative, and the epistemological challenges are the most severe. When a model surfaces “emergent themes,” how do you distinguish patterns in your data from patterns in the model’s training data? The model might be “finding” strategies that reflect how it was trained and tuned, rather than the community you’re studying. And since the output is well presented, it looks like insight. It reads well. It sounds plausible. And that is what makes it dangerous. The question of rigor then becomes much more fundamental: what does reflexivity even mean when you’re not the interpretive agent?
These three options are not a historical sequence. They coexist. But they carry very different epistemological commitments, and demand very different forms of accountability.
For the deductive case, we can draw on established validation frameworks. For the controlled pipeline, we need transparency, discussion of choices, and ways to come up with best practices. For the inductive case, we need something we haven’t invented yet. Perhaps something we could call model positionality – ways of documenting, probing, and comparing models to better understand and account for their epistemic character.
That is what I mean by the post-quantitative turn. Not that the quantitative-qualitative divide has vanished, but that it has been relocated inside computational methods themselves, in forms that are hard to keep track of.
If AI is transforming how we do research, this is one of the places where it hits: the old assumption that computation is quantitative and interpretation is human no longer holds. What should concern us is not that researchers will choose the wrong option among the three I’ve sketched, but that they will more and more drift toward the third one and make no conscious choice at all.
I think that we should try to make sure that this doesn’t happen.