#TopicModeling
🔍 How robust are findings based on topic modeling?

Our 🏆 Top Faculty Award article on #TopicModeling + #MultiverseAnalysis is now #openaccess in Communication Methods & Measures.

📄 doi.org/10.1080/19312458.2026.2714769
🔁 gesistsa.github.io/multiverse_tm/

wonderful teamwork @gesis.org 👇
Beyond beyond standardization: studying robustness of empirical claims based on topic modeling through multiverse analysis
While topic modeling is widely used, some scholars have already announced that the application of topic modeling in social science research is impossible to standardize. Meanwhile, researchers have...
doi.org
September 24, 2026 at 8:15 AM
Okay Bluesky, I need your help. I'm looking for some recommendations: What are your favourite tools/examples for #TopicModeling, #Stylometry, and #NetworkAnalysis?

Please comment below or DM me. And feel free to repost.

#DigitalHumanities #NLP #NLProc
July 23, 2025 at 11:24 AM
Do you work with text data? Then our ✨topiclabels✨ #Rstats package may come in handy.

Using open #LLM, it automatically assigns a topic label to a bag of words.

It also works with all popular #TopicModeling packages!

👉 cran.r-project.org/package=topi...
👉 github.com/PetersFritz/...
December 10, 2024 at 10:23 AM
Last update in 2024 🚀 for PsychTopics, our #Rstats #ShinyApp that automatically identifies research topics in psychology from DE, AT, CH & LUX

Data source: PSYNDEX database of @zpid.bsky.social
Method: RollingLDA #TopicModeling (aclanthology.org/2022.sdp-1.2)

👉 abitter.shinyapps.io/psychtopics
December 18, 2024 at 10:40 AM
🚨Exciting research alert! 🚨 Explore the world of topic modeling in communication science with our latest paper. Discover how different validation methods impact model selection and its consequences for theory development. 📚 commsky #TopicModeling www.aup-online.com/content/jour...
Topic Model Validation Methods and their Impact on Model Selection and Evaluation | Amsterdam Univer...
Topic Modeling is currently one of the most widely employed unsupervised text-as-data techniques in the field of communication science. While researchers increasingly recognize the importance of valid...
www.aup-online.com
October 23, 2023 at 1:36 PM
How to read thousands of newspaper articles?

Three researchers at Herder-Institut Marburg used #TopicModeling to analyse over 7,000 articles in Latvian diaspora newspapers - and discovered how displaced communities maintained identity across decades and continents ⬇️
What Algorithms Find in Latvian Exile Newspapers: Songs, Maps, and the Paradoxes of Digital Reading
Simon Donig, Dinara Gagarina and Timur Mitrofanov use computational analysis to gain new insights into how Latvian diaspora newspapers preserved cultural identity.
valuepast.hypotheses.org
April 27, 2026 at 12:23 PM
#CCLS2026: 12th and last talk is "From #LiteraryCriticism to #LiteraryStudies: #TopicModeling Argentine Academic Journals (1982–2024)" by Federico Gabriel Cortés, analyzing the topics of 42 years of academic publishing of literary studies research in Argentina.

Preprint: doi.org/10.26083/tud...
May 29, 2026 at 10:10 AM
Based on an analysis of #DHd conference abstracts, I trace the evolution of #CLS methods from 2014 to 2025: from omnipresent #NetworkAnalysis and #Annotation, to the first appearance of #TopicModeling and #SentimentAnalysis, to #DeepLearning and #GenerativeAI.
May 13, 2025 at 3:20 PM
Interesse an #topicmodeling für historische Fachzeitschriften? Eike Löhden & ich haben 50 Jahrgänge der „Francia“ analysiert.

Wie haben sich Schwerpunkte über die Jahre entwickelt? Welche Unterschiede zeigen sich zwischen deutsch- und französischsprachigen […]

[Original post on fedihum.org]
February 14, 2025 at 9:19 AM
Happy to announce that our #Rstats package ✨topiclabels✨ has been updated on #CRAN 🎉

🤖Using open #LLMs, our package automatically assigns a topic label to a bag of words.
🤝It works with all popular #TopicModeling packages!

Find out more:
👉https://github.com/PetersFritz/topiclabels
October 30, 2024 at 2:15 PM
📢 🧑‍🏫 Das Lehrangebot der DigitalHistory (@humboldtuni.bsky.social) bietet auch im SoSe24 wieder ein reichhaltiges Programm:
Von DataLiteracy & Python über Einführungen in SentimentAnalysis & TopicModeling bis hin zu #auxHist sowie Bibliotheken im Verhältnis zur #digiGW

➡️ hu.berlin/DigHisLehreS...
February 22, 2024 at 8:32 AM
This week's Wednesday webinar (Feb. 26) at Vanderbilt Biostatistics is "Topic Models in Microbiome Analysis," at 1:30 pm CT, by Kris Sankaran @sankaranlab.bsky.social www.vumc.org/biostatistic... #RStats #GISky #GutSky #MedSky #TopicModeling
February 25, 2025 at 8:22 PM
Our paper from #CoroNarrate project is out in European Politics and Society: we investigate with #stm #topicmodeling how institutional actors in Italy & Germany framed #Covid19 economic policies during the #pandemic. Thank you @tillhilmar.bsky.social Patrick Sachweh
doi.org/10.1080/2374...
March 26, 2026 at 9:18 AM
JMIR Formative Res: Improving Suicidal Ideation Detection in Social Media Posts: Topic Modeling and Synthetic Data Augmentation Approach #MentalHealth #SuicidePrevention #SocialMedia #PublicHealth #TopicModeling
Improving Suicidal Ideation Detection in Social Media Posts: Topic Modeling and Synthetic Data Augmentation Approach
Background: In an era dominated by social media conversations, it is pivotal to comprehend how suicide, a critical public health issue, is discussed online. Discussions around suicide often highlight a range of topics, such as mental health challenges, relationship conflicts, and financial distress. However, certain sensitive issues, like those affecting marginalized communities, may be underrepresented in these discussions. This underrepresentation is a critical issue to investigate because it is mainly associated with underserved demographics (eg, racial and sexual minorities), and models trained on such data will underperform on such topics. Objective: The objective of this study was to bridge the gap between established psychology literature on suicidal ideation and social media data by analyzing the topics discussed online. Additionally, by generating synthetic data, we aimed to ensure that datasets used for training classifiers have high coverage of critical risk factors to address and adequately represent underrepresented or misrepresented topics. This approach enhances both the quality and diversity of the data used for detecting suicidal ideation in social media conversations. Methods: We first performed unsupervised topic modeling to analyze suicide-related data from social media and identify the most frequently discussed topics within the dataset. Next, we conducted a scoping review of established psychology literature to identify core risk factors associated with suicide. Using these identified risk factors, we then performed guided topic modeling on the social media dataset to evaluate the presence and coverage of these factors. After identifying topic biases and gaps in the dataset, we explored the use of generative large language models to create topic-diverse synthetic data for augmentation. Finally, the synthetic dataset was evaluated for readability, complexity, topic diversity, and utility in training machine learning classifiers compared to real-world datasets. Results: Our study found that several critical suicide-related topics, particularly those concerning marginalized communities and racism, were significantly underrepresented in the real-world social media data. The introduction of synthetic data, generated using GPT-3.5 Turbo, and the augmented dataset improved topic diversity. The synthetic dataset showed levels of readability and complexity comparable to those of real data. Furthermore, the incorporation of the augmented dataset in fine-tuning classifiers enhanced their ability to detect suicidal ideation, with the F1-score improving from 0.87 to 0.91 on the University of Maryland Reddit Suicidality Dataset test subset and from 0.70 to 0.90 on the synthetic test subset, demonstrating its utility in improving model accuracy for suicidal narrative detection. Conclusions: Our results demonstrate that synthetic datasets can be useful to obtain an enriched understanding of online suicide discussions as well as build more accurate machine learning models for suicidal narrative detection on social media.
dlvr.it
June 11, 2025 at 4:43 PM
JMIR Mental Health: Interpretable Topic Modeling of Spontaneous Speech in #depression Using Large Language Models: Multilingual Four-Cohort Study #Depression #MentalHealth #AIinHealthcare #LanguageModels #TopicModeling
Interpretable Topic Modeling of Spontaneous Speech in #depression Using Large Language Models: Multilingual Four-Cohort Study
Background: #depression is underdiagnosed worldwide, and clinicians rely on interpreting patients’ subjective speech. Qualitative analysis of patient language does not scale, and existing computational #Approaches describe topics with keyword lists that miss clinical nuance. Objective: We evaluated whether clustering spontaneous speech transcripts with large language models (LLMs) yields clusters whose membership is associated with validated clinical scales across multilingual cohorts, and whether LLMs can render those clusters human-readable through fine-grained natural-language descriptions. We further examined which interview questions yield clusters most strongly associated with clinical status, and how sociodemographic factors relate to cluster membership. Methods: We analyzed spontaneous speech transcripts from 4 independent cohorts: a French general population sample (1338 participants) and 3 clinical samples in Italian (n=116), Chinese (n=52), and Spanish (n=90). Responses to open-ended questions were transcribed, embedded with a multilingual language model, dimensionally reduced, and grouped by density-based clustering. An LLM then summarized each cluster into a natural-language description. Cluster membership was tested for association with validated clinical scales (Patient #Health Questionnaire-9, Beck #depression Inventory, Generalized #anxiety Disorder 7-item scale, Athens Insomnia Scale, Multidimensional Fatigue Inventory, and Columbia Suicide Severity Rating Scale), clinician-assigned #depression diagnoses, and sociodemographic factors (age, education, and sex). Results: Unsupervised clustering yielded clusters significantly associated with clinical scores in the French, Italian, and Chinese cohorts, with an exploratory association in the smaller Spanish cohort. In the French general population, Patient #Health Questionnaire-9 #depression scores differed across clusters (η²=0.19, 95% CI 0.17 to 0.24,
dlvr.it
August 26, 2026 at 3:04 PM
Wie kann man Topics verstehen? Erläutert @u_henny anhand einer spanischen Wortwolke, basierend auf einem Korpus hispanoamerikanischer Romane. #TopicModeling #dhd2019
December 10, 2024 at 3:05 PM
Definition von #TopicModeling von @JanHorstmannn: Ein auf Wahrscheinlichkeitsrechnung basierendes Verfahren zur Exploration größerer Textsammlungen. Bietet die Möglichkeit, Textsammlungen thematisch zu explorieren. https://fortext.net/routinen/methoden/topic-modeling #dhd2019
December 10, 2024 at 3:05 PM
What is LDA and Why It’s Not Just for Text

Latent Dirichlet Allocation (LDA). Do you think about its roots in text analysis? #dataclustering #graphdata #imagedataanalysis #LDA #LDAapplications #LDAforimages #LDAinmusic #musicdata #nontextdata #topicmodeling
aicompetence.org/lda-beyond-t...
LDA Beyond Text: Applications In Image, Music, And Graph Data
LDA goes beyond text analysis, uncovering patterns in image, music, and graph data, driving innovative insights across diverse data types.
aicompetence.org
October 11, 2024 at 7:58 PM
.#BSwallow & #EBayer @ #MDurrett et al @ #dayofdh18cc: #topicmodeling of Latin texts = took wks. So for comps / to fill time, they built AWESOME #FOSS app to visualize #topicmodels. http://cs.carleton.edu/cs_comps/1718/latin/final-results/the-app.html #loquela CC @dighall @nolauren...
December 6, 2024 at 9:12 AM
Thrilled to have the chance to talk about my #topicmodeling #research at the coming #python and friends conference @PyGrunn !

And a bit intimidated though, based on the website I think I m the only #female speaker 😬

#phdlife #genderbalance #programming #womenintech
June 13, 2025 at 1:38 PM
This paper uses topic modeling and bias measurement techniques to analyze and determine gender bias in English song lyrics. Our analysis shows the thematic shift in song lyrics over the years

#mathsky #compsky #science #topicmodeling
Beats of Bias: Analyzing Lyrics with Topic Modeling and Gender Bias Measurements
arxiv.org
February 7, 2025 at 5:52 PM