#Quanteda
Caught doing quanteda out in the open...
Der #Demokratie bei der Arbeit zugucken: Forscher*innen der #FUBerlin haben Reden aus allen 16 Landtagen analysiert & in der Datenbank "StateParl“ durchsuchbar gemacht.
Die teils überraschenden Ergebnisse ➡️
www.fu-berlin.de/campusleben/...

Foto: Nicolas Pannetier | Atelier Limo
July 9, 2025 at 4:04 PM
Pour celleux voulant reproduire ce type d'analyse, j'ai publié un package R avec toutes les fonctions permettant de scraper les sous-titres des vidéos YouTube, de préparer les bases de données, les traiter par IRaMuTeQ ou Quanteda.

tdelc.github.io/lexico/index...
January 2, 2026 at 10:37 PM
If you've tried to use the #rstats NLP package {spacyr} in 2023, you may have noticed that the installation was broken. I fixed it and a new version including my changes is now on CRAN and GH 🥳

github.com/quanteda/spa...
GitHub - quanteda/spacyr: R wrapper to spaCy NLP
R wrapper to spaCy NLP. Contribute to quanteda/spacyr development by creating an account on GitHub.
github.com
December 18, 2023 at 10:28 AM
January 14, 2026 at 3:19 PM
Fin de journée d'analyse, avec un petit outil mis en ligne pour m'aider à explorer les textes des vidéos :

tdelc.shinyapps.io/info_continue/

Une application interactive essentiellement basée sur l'excellent package quanteda. Accessible librement.
Exploration lexicométrique
tdelc.shinyapps.io
November 17, 2025 at 12:03 AM
New conceptual review + tutorial on text embeddings out in #APA_Journals w/ @almogsi. Beginner-friendly, but experts will find spicy new takes as well. Tag a colleague who’s still counting words... #RStats #tidyverse #quanteda
1/
June 24, 2025 at 2:10 PM
J'ai rédigé quelques vignettes :
- Scraper des playlists YouTube
- Préparer les données textuelles
- Préparer et exploiter des corpus avec IRaMuTeQ
- Analyser les textes avec Quanteda

tdelc.github.io/lexico/artic...
Scraper des playlists YouTube
tdelc.github.io
January 2, 2026 at 10:37 PM
¡Nueva publicación! Esta vez de parte de uno de nuestros investigadores noveles. Sergio Castro-Cortacero y Nicolas Robinson-Garcia publican para la revista Revista Infonomy como utilizar la herramienta #Quanteda para realizar análisis de textos.

DOI: doi.org/10.3145/info...
January 19, 2026 at 8:00 AM
If you think the number of topics, k, is the only important parameter for topic models, you need to read this post and the research paper. blog.koheiw.net?p=2233 I created a new model to optimize the Dirichlet priors to analyze imbalanced corpus more accurately. #rstats #quanteda
A new topic model for analysis imbalanced corpus
I have been developing and testing a new topic model called model Distributed Asymmetric Allocation (DAA) because latent Dirichlet allocation (LDA) takes a long time to fit to a large corpus but do…
blog.koheiw.net
November 23, 2024 at 10:02 AM
I released the wordvector package v0.5.0. It is rapidly getting better and different from the original Word2vec package. Please read "Align word vectors of multiple Word2vec models" about the new function blog.koheiw.net?p=2299 #rstats #quanteda
Align word vectors of multiple Word2vec models
I have been developing a new R package called wordvector since last year. I started it as a fork of the Word2vec package but made several important changes to make it fully compatible with quanteda…
blog.koheiw.net
May 24, 2025 at 2:50 AM
Also: I thought Ken was playing around with getting Quanteda to do this kind of jam?
September 25, 2026 at 2:40 AM
newspapers get new authors and new guidelines in article writing. And the changes (apart from the en-dash) are too small to really spot with human eyes.

But I do agree with Mario that the 2025 change is... unexpectedly large.

Ok breakfast, then let's see about the quanteda package.

3/3 FOR NOW
November 2, 2025 at 9:15 AM
卒論指導の一環でRのquantedaパッケージを使うと「共起ネットワークがうまく描けない」「KH Coderみたいなわかりやすい図を作りたい」という要望があり、GitHubに公開されているKH Coderのコードを参考にして独自のRパッケージを開発しました

これを使えばquantedaのdfmオブジェクトからKH Coder風の共起ネットワークが簡単に作成できます
(ggplotで作図しているので、図をggplotでカスタマイズすることも可能です)

もし興味のある方がいればぜひお使いください
github.com/namiterashit...
https://github.com/namiterashita/khcnet
December 12, 2025 at 2:20 PM
This week on What's New in R:
✅ Data viz. on domestic terrorism trends by Stephen Ponce
✅ Ted Laderas introduces Databot, Positron’s AI-powered analysis tool
✅ {quanteda} package developed by Kenneth Benoit for analyzing qualitative text data

Read the issue: buff.ly/gUPhm9s

#rstats
January 12, 2026 at 4:03 PM
A few days ago, I received an email from a researcher asking if text analysis is becoming irrelevant because of AI... blog.koheiw.net?p=2254 #text-as-data #quanteda
December 13, 2024 at 12:16 AM
I already blogged today, so don't want to write more. But here is the code
November 24, 2024 at 3:04 PM
Se abordarán estrategias cuantitativas, cualitativas y computacionales para la investigación social, incluyendo:
🔹 IA y modelos de lenguaje
🔹 RRSS (Gephi, UCINET)
🔹 Análisis textual con R (Quanteda)
🔹 Etnografía multisituada
🔹 Participación social y datos digitales
June 5, 2025 at 8:19 PM
Text analysis is gradually moving away from bag-of-words analysis, so I created a new package for Gaussian mixture topic models blog.koheiw.net?p=2437. I did not not do much C++ programing thanks to the Armadillo package. GMTM is extremely fast thanks to multi-threading. #text-as-data #quanteda
September 6, 2026 at 9:05 AM
Hoje estou me sentindo muito garoto de programa fazendo as adaptações do meu código no R utilizando a documentação do pacote que mudou.

Quanteda eu te odeio, mas te amo!
September 9, 2024 at 3:49 PM
Not sure if it's helpful but two R #packages which may be of use are quanteda:: (for quant text analysis) and tidytext:: is useful for text mining. If this helps at all? Otherwise I am clueless on the subject.
November 18, 2024 at 6:21 PM
I've used spacyr a lot. It is totally fine. Currently, there is a bug with the install process, but there is a well-working fix here :https://github.com/quanteda/spacyr/issues/236#issuecomment-1702311825

Reticulate has worked fine when needed. Should probably learn python though. Just, time...
October 22, 2023 at 6:33 PM
✨ New to NLP? We link a post that gives you a general understanding of text mining and the prominent bag of words approach in #rstats

https://bit.ly/text-mining-quanteda
Methods Bites
Blog of the MZES Social Science Data Lab
bit.ly
April 1, 2023 at 4:19 PM
CRAN updates: gson ibdsim2 quanteda roxygen2 uGMAR #rstats
August 4, 2026 at 8:02 PM
Updates on CRAN: enrichit (0.2.1), gson (0.2.1), hcinfer (0.2.0), ibdsim2 (2.3.3), quanteda (4.5.0), roxygen2 (8.1.0), uGMAR (3.6.1)
August 4, 2026 at 10:06 PM