Topic modelling and interpretative narrative analysis

Authors

  • Michael Corsten Institut für Sozialwissenschaften, Universität Hildesheim
  • Ulrich Heid Institut für Informationswissenschaft und Sprachtechnologie, Universität Hildesheim
  • Patrick Kahle Bielefeld Graduate School in History and Sociology, Universität Bielefeld
  • Fritz Kliche Institut für Informationswissenschaft und Sprachtechnologie, Universität Hildesheim

DOI:

https://doi.org/10.21241/

Keywords:

Topic modelling, Interpretative Narration Analysis, Automated Processing of Qualitative Data, Computational Linguistics

Abstract

Topic modelling based on Latent Dirichlet Allocation (LDA) (Blei et al. 2003) is used to identify topics and lexical material in text collections and to describe probability distributions of words and topics in a corpus.

Grootendorst (2022) presents BERTopic, a topic modeling approach based on the localization of words, sentences and sections of texts in the embedding space of transformers. The approach is based on the assumption that text fragments are similar if they are close to each other in the embedding space; TF-IDF then allows words that are prominent for a topic to be found. 

As part of work on the interpretative analysis of CV narratives, we investigated which topics are extracted from a collection of CV descriptions when BERTopic is used in a bottom-up approach, and how the topics differ when seeded topic modelling is used, i.e. when certain words expected from a theoretical perspective or from the previous topic analysis are assigned to specific topics in order to control the generation of topics (cf. Kahle and Kliche, 2022).

This opened up the methodological possibility of a threefold comparison of the quality of the results of a human interpretative analysis, bottom-up topic modelling and seeded topic modelling. In the project, which explored text corpora from sociological biographical research, oral history and sociological expert interviews, assessments of the quality of human and machine-generated classifications/interpretations could thus be made in a qualitative (rule-based) and quantitative (stochastic) manner. Cases can be shown in which the results of the topic modelling procedures also provided interpretatively appropriate classifications in rule-based assessments or also showed statistically reliable classifications (e.g. according to Krippendorff's alpha or Cohen's kappa) through human classification.

References

Austin, John L. 1962. How to do Things with Words. Harvard: University Press

Biemann, Chris, Gerhard Heyer und Uwe Quasthoff. 2022. Wissensrohstoff Text. Wiesbaden: Springer.

Blei, David M., Andrew Y. Ng und Michael I. Jordan. 2003. Latent Dirichlet Allocation. The Journal of Machine Learning Research 3:993-1022

Boyd, Ryan L., Kate G. Blackburn und James W. Pennebaker. 2020. The narrative arc: Revealing core narrative structures through text analysis. Science Advances 6(32):eaba2196

Bourdieu, Pierre. 1982. Qu'est-ce que parler veut dire? Paris: Fayard.

Corsten, Michael, und Laura Maleyka. 2024. Die Präsentation der biographischen Visitenkarte: Erzählstimulus und Selbsteinführung am Interviewbeginn. BIOS – Zeitschrift für Biographieforschung, Oral History und Lebensverlaufsanalysen 37:157–181

De Saussure, Ferdinand. 1995 [1916]. Cours de linguistique générale. Paris: Payot, coll.

Eder, Maciej, Jan Rybicki, J. und Mike Kestemont. 2016. Stylometry with R: a package for computational text analysis. R Journal 8:107–121.

Fischer-Rosenthal, Wolfram, und Gabriele Rosenthal. 1997. Narrationsanalyse biographischer Selbstpräsentation. In: Sozialwissenschaftliche Hermeneutik, Hrsg. Ronald Hitzler und Anne Honer, 133–164. Opladen: Leske + Budrich.

Foucault, Michel. 1969. L’archéologie du savoir. Paris: Éditions Gallimard.

Gao, Xin, und Cem Sazara. 2023. Discovering mental health research topics with Topic Modeling. arXiv:2308.13569

Greimas, Algirdas J. 1966. Sémantique structurale. Recherche de méthode. Paris: Larousse.

Grootendorst, Marten. 2022. BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv:2203.05794

Habermas, Tilmann, und Elaine Reese. 2015. Getting a life takes time: The development of the life story in adolescence, its precursors and consequences. Human Development 58:172–201

Harris, Zellig. 1954. Distributional Structure. Word 10(2/3):146–162.

Jafke, Larissa, Laura Maleyka, Holger Herma, Karsten Spindler, Michael Corsten und Kathrin Audehm. 2022. Sich als Subjekt des Sprechens über das eigene Leben einführen. In: Positioning the Subject, Hrsg. Saša Bosančić, Folke Brodersen, Lisa Pfahl, Lena Schürmann, Tina Spies und Boris Traue, 115–138. Wiesbaden: Springer VS.

Kahle, Patrick, und Kliche, Fritz. 2022. Learning inductive and deductive topics in parallel using seeded topic modeling. In: Diversity of Methods and Materials in Digital Human Sciences: Proceedings of the Digital Research Data and Human Sciences DRDHum Conference 2022, December 1–3, Hrsg. Jarmo H. Jantunen, Johanna Kalja-Voima, Matti Laukkarinen, Anna Puupponen, Margareta Salonen, Tuija Saresma, Jenny Tarvainen und Sabine Ylönen, 28–42. Jyväskylä, Finland.

Kallmeyer, Werner, und Fritz Schütze. 1977. Zur Konstitution von Kommunikationsschemata der Sachverhaltsdarstellung. In: Gesprächsanalysen. Vorträge, gehalten anlässlich des 5. Kolloquiums des Instituts für Kommunikationsforschung und Phonetik, Hrsg. Dirk Wegner, 159–274. Bonn.

Küsters, Ivonne. 2021. Narrative Interviews. Wiesbaden: Springer VS.

Labov, William, und Waletzky, Joshua. 1967. Narrative analysis: Oral versions of personal experience. In: Essays on the Verbal and the Visual Arts, Hrsg. June Helm, 3–38. Seattle und London: University of Washington Press.

Lämmert, Eberhard. 1955. Bauformen des Erzählens. Stuttgart: Metzler.

Pennebaker, James W. 2011. The secret life of pronouns: What our words say about us. New York: Bloomsbury.

Schütze, Fritz. 1984. Kognitive Figuren des autobiographischen Stegreiferzählens. In: Biographie und soziale Wirklichkeit. Neue Beiträge und Forschungsperspektiven, Hrsg. Martin Kohli und Günter Robert, 70–117. Stuttgart: Metzler.

Stanzel, Franz K. 1982. Theorie des Erzählens. Göttingen: Vandenhoeck und Ruprecht.

Weber, Dietrich. 1998. Erzählliteratur. Schriftwerk – Kunstwerk – Erzählwerk. Göttingen: Vandenhoeck und Ruprecht.

Wiedemann, Gregor, und Matthias Lemke. 2016. Text Mining für die Analyse qualitativer Daten. Auf dem Weg zu einer Best Practice? In: Text Mining in den Sozialwissenschaften, Hrsg. Matthias Lemke und Gregor Wiedemann, 397–419. Wiesbaden: Springer.

Ziems, Caleb, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, Diyi Yang. 2024. Can Large Language Models Transform Computational Social Science? Computational Linguistics 50:237–291.

Downloads

Published

2026-10-07

Issue

Section

Ad-Hoc: Maschinelles Lernen, Natürliche Sprachverarbeitung und Qualitative Methoden (AdH54)

Most read articles by the same author(s)