LLMs for Survey Text Analysis: A Performance Comparison Between Humans and GPT-5 on Inductive Content Analysis
Synopsis
Using 903 open-ended responses across six variables from a European PhD student survey, this study had five human coders and GPT-5.4 each perform the same inductive content analysis procedure to produce codes and themes, and measured agreement with the Adjusted Rand Index (ARI), finding average human-LLM agreement of 0.61 for coding and 0.54 for themes, close to within-human consistency (0.68) and within-LLM consistency (0.76), with wide variation across variables and low within-entity consistency consistently accompanying low between-entity agreement.
Interpretation
On inductive content analysis, human-LLM agreement (ARI = 0.61 for coding, 0.54 for themes) is close to the alignment between coding and themes within humans (ARI = 0.68) and within the LLM (ARI = 0.76). Prior evidence on LLM text coding has largely concerned deductive coding, especially binary categories, while data on the inductive approach and its comparison with humans remain rare; this study runs the same inductive procedure with both humans and a model and reports comparable quantitative agreement. Based on 903 responses across six variables, with five human coders and GPT-5.4 each producing codes and themes, compared using the Adjusted Rand Index, with multi-label code assignments transformed into co-assignment matrices before calculation; the authors note the index is 0 for random partitions and bounded above by 1 for perfect agreement.
Human-LLM agreement varies widely across variables, with ARI values ranging from 0.31 to 0.89, and low internal consistency of a coding entity is always associated with low agreement between coding entities. The study links reliability differences to data characteristics and individual performance, showing that variation is not evenly distributed, for example all values for Financial Situation are at or below 0.71 while all values for Anything Else are at or above 0.70. Based on the per-variable ARI values in Table 1 and the authors' corresponding explanations of low LLM internal consistency for Financial Situation and Well Being and low human internal consistency for Personal Experience.
Alignment at the theme level is lower than at the coding level, which the authors attribute to more degrees of freedom in the additional abstraction step, noting this reduction also occurs between humans. It reframes the coding-versus-theme gap from a model-specific issue toward a feature of the inductive analysis steps themselves, consistent with related work showing lower theme agreement than coding agreement. The contrast between coding ARI = 0.61 and theme ARI = 0.54, together with the related work the authors cite.
Human coders, after seeing the model's output, indicated that the LLM coding reflected their own categorizations and supported using this approach for further data, while raising concerns about overinterpretation, reduced precision, and broadened themes. It supplements the quantitative indices with coders' direct judgments of the model output, providing qualitative grounding for positioning the model as a support tool for early stages. The per-variable coder opinions in Appendix D, including specific observations about short responses being overinterpreted, different responses being treated identically, and multiple codes being merged into a single theme.
Perspective
This work addresses inductive content analysis of open-ended survey responses, particularly the early, labor-intensive coding stage, and is meant for qualitative researchers and teams handling large text volumes who want LLM support for initial categorization. The authors stress that the results should not be read as suggesting LLMs replace human qualitative judgment, since human researchers remain essential for contextual interpretation, validation, and theory-building, and they frame the conclusions as case-specific to this setting.
The authors note the study relies on a single model, lacks a clear ground truth, and used a single human coder per variable, so interpretation of human-LLM agreement warrants caution; future work could explore different models, prompting strategies, and evaluation approaches with more human coders per variable. In addition, Table 1 shows an NA in the theme column for the Accessibility Increase variable because the coder responsible for that variable lacked theme coding, and this gap affects what can be judged about theme-level agreement for that variable.
