Grounded theory (method) was built to study talk and text; but now social media, AI and algorithms are not far

Grounded theory codes a single stream of talk. A reel is not one: it arrives with its synthesised voice, its caption, its hashtags, its metrics and a thread that talks back. MGTM is a dual-track extension for coding such artifacts, and their reception, without flattening either.
Grounded theory (method) was built to study talk and text; but now social media, AI and algorithms are not far
Like

Share this post

Choose a social network to share with, or copy the URL to share elsewhere

This is a representation of how your post may appear on social media. The actual post will vary between social networks

Explore the Research

Springer Netherlands
Springer Netherlands Springer Netherlands

Multimodal grounded theory method: an extension framework to analyze layered social-media data in the current algorithmic culture

The proliferation of convergent social-media data, where the audio, visual, textual and other interactive elements from a YouTube deepfake, a TikTok reel, a Facebook post or an Instagram short do not arrive as a single stream but tightly coupled composite of moving image, sound, on-screen and embedded text, hashtags, captions, engagement metrics, and comment threads that talk back to the artifact in real time, produces artifacts that are dynamically integrated with audience commentary and algorithmic metadata, pose a fundamental methodological challenge for qualitative research. While the classical Grounded Theory Method (GTM) has been extended to visual and audiovisual data, it lacks explicit analytic procedures for systematically disaggregating, analyzing and reconverging these layered platform-centric multimodal artifacts. This paper addresses this gap by introducing the Multimodal Grounded Theory Method (MGTM), a formal methodological framework that operationalizes GTM’s inductive logic for the complexities of convergent digital data. MGTM makes three core analytical operations explicit: (1) modal differentiation, where artifacts are analytically segmented into visual, auditory, textual and receptional streams before coding; (2) convergent axial coding, which theorizes the relationships and tensions between codes across modes; and (3) the constitutive treatment of audience reception and platform metadata as primary data for theory generation. Developed through a double-track (Tracks A and B) systematic analysis, beyond eight major visual and audiovisual GTM adaptations, MGTM provides a structured yet flexible workflow, from data capture and multimodal open coding to iterative theoretical sampling. The extensive method is demonstrated first on a single test case, a YouTube deepfake with 258 viewer comments, and then on two further cases, from TikTok and Instagram, with respectively 448 and 737 viewer comments on each (totaling n = 1443 reception reactions overall) to generalize the apparatus. Across the three, the decisive variable demonstrates MGTM as a structured extension framework of doing grounded theory when the data are layered and algorithmically mediated.

When Glaser and Strauss developed grounded theory in 1967, the datum was an interview transcript: a single linguistic stream that could be coded line by line. Later extensions carried the method to photographs and to video but the unit of analysis in each remained a discrete object, something that sits still while you code it.

Yet, the artifacts that now organise public attention do not sit still. A YouTube deepfake, a TikTok fact-check, an Instagram reel: each arrives as a convergent composite of moving image, synthesised voice, on-screen text, hashtags, engagement metrics and a comment thread that talks back in real time, all shaped by the same algorithmic infrastructure. Any researcher who has tried to code such an object knows the two losses on offer. Either the multimodal complexity is diluted into verbal description, or the artifact is broken into modal pieces that lose exactly the convergence that made it worth studying.

The method in its first working form, drawn by hand while the analysis was still finding its shape.
Ideating the model

What the method gives you

MGTM closes that gap with three operations that grounded theory has been performing implicitly and that this paper renders explicit and replicable.

Firstly, modal differentiation is the pre-coding move. The coupled artifact is bracketed into visual, auditory, textual and reception streams, each coded in its own semiotic terms before any synthesis is attempted. In practice this means watching with the sound off, then listening with the eyes closed, then reading the captions as a discourse analyst would, and only then watching the whole. It refuses the single fused judgment , it looks real, it does not, long enough to specify which mode is doing what work.

And then, convergent axial coding recomposes the streams, asking which artifact features systematically produce which reception patterns.

Constitutive reception is the consequential one. Comments, emojis, engagement metrics and moderation actions are treated as primary data for theory generation, not as background illustration. The workflow is therefore dual-track: Track A codes the artifact mode by mode, Track B grows one- and two-word tiny codes out of individual comments into emergent families, and the two tracks meet in convergent axial coding.

Track A codes the artifact mode by mode; Track B grows tiny codes into reception families; the two tracks meet in convergent axial coding.

The MGTM Dual-Track Workflow: Track A codes the artifact mode by mode; Track B grows tiny codes into reception families; the two tracks meet in convergent axial coding.

What that buys

In the demonstration case, when I studied a 258-comment thread beneath an AI-voiced Trump–Biden "interview", exactly one comment raised an AI-safety concern. Under one per cent. The largest family by a wide margin was comic connoisseurship. That absence is the paper's finding, and it is legible only because reception is coded as data in its own right. Two further cases, a CBC fact-check and a Lil Miquela reel, yielded two further theories from the same apparatus: what governs public response is neither detectability nor disclosure, the levers policy habitually reaches for, but the frame.

For whom

Now, if your data are a thread, a reel, a duet or a comment war, your object is multimodal and your audience is part of it. MGTM is structured enough to be replicable and loose enough to let the data drive the categories, and the full coding workbooks behind all three cases sit in the appendix, so the procedure can be inspected, contested and reused. Grounded theory has always insisted that all is data. The question this paper puts is what that principle now requires, once the data arrive already convergent, already algorithmically shaped, and already talking back.

This work can be found here: Bashir, S. (2026). Multimodal grounded theory method: an extension framework to analyze layered social-media data in the current algorithmic culture. Quality & Quantity. https://doi.org/10.1007/s11135-026-02948-y