#gensim
I'm reading a paper that compares BERTopic and LDA for tweets and includes:
1. a 2020 citation for LDA, not a single mention of Blei 😐
2. de-duplication BEFORE removing usernames 😟
3. stemming, lemmatizing, etc. ONLY for LDA 😨
4. gensim 😱

It has 1k citations! 🤯
December 8, 2025 at 9:41 PM
another day, another person telling me that they tried topic modeling and it didn't work

(they used gensim, they always use gensim, please don't use gensim)
May 2, 2023 at 2:46 PM
I agree that Mallet >> Gensim, but VI can also work well. It's just more finicky, and Gensim's implementation doesn't do the things that Mallet does to optimize hyperparameters. So just like we shouldn't let LDA get a bad rap because of Gensim, we shouldn't let VI get a bad rap because of Gensim.
December 10, 2025 at 12:14 AM
gensim使いましょう!!!
April 18, 2025 at 1:03 PM
Tho for next time, will def. be using tomotopy rather than gensim for LDA (h/t @mariaa.bsky.social!)
October 8, 2025 at 3:09 PM
I am become gensim, vectorizer of words
November 5, 2023 at 3:47 AM
It doesn't train with Gibbs sampling, and generally produces much less coherent results (maybe because fewer passes over data). I have years of anecdotes of people telling me that they struggle with gensim. Confirmed by many other researchers' experiences in an old Twitter thread.
December 8, 2025 at 10:12 PM
Embetter has a new update!

We dropped support for gensim and added support for ollama!

github.com/koaning/emb...
GitHub - koaning/embetter: just a bunch of useful embeddings for scikit-learn pipelines
just a bunch of useful embeddings for scikit-learn pipelines - koaning/embetter
github.com
August 18, 2025 at 10:46 AM
[FREE] 400 Python Gensim Interview Questions with Answers 2026
400 Python Gensim Interview Questions with Answers 2026 - 100% OFF Udemy Coupon | FreeWebCart
Master Word2Vec, LDA, and Scalable NLP with Realistic Practice Tests and Detailed Explanations.Python Gensim Interview and Practice Questions are designed to…
freewebcart.com
March 4, 2026 at 1:21 PM
I tried using BERT in a project in 2019 and thought it sucked real bad compared to even basic NLP like GenSim at the time. Google sent a team of people to a conference I attended about the topic and they were basically ignored. I've never been proven more wrong about a technology in my life!
August 20, 2026 at 4:38 PM
In This House, we don't really believe in Sentiment Analysis. Still, I keep meaning to try topic-modelling with Gensim.

Is there any point to my trying to maintain this? Don't touch the code very often, and it's a bit of plate-spinning to keep all the crawlers going.
November 21, 2024 at 3:45 PM
It’s unfortunate because I’ve found LDA (mallet/tomotopy, not gensim) is actually very good for tweets
December 9, 2025 at 11:20 AM
in this case I am forcing poor gensim to find embeddings for horrific slurs
November 5, 2023 at 3:47 AM
June 6, 2025 at 9:22 AM
Install update:
✅Advanced NLP capabilities (transformers, spacy, nltk)
✅Deep learning models (torch, tensorflow, keras)
✅Text embeddings & semantic analysis (sentence-transformers, gensim)
✅More sophisticated cognitive functions for content analysis and generation
April 28, 2025 at 4:58 PM
Gensim vs. Word2Vec: Modelos de Representación de Palabras en Comparación

En el mundo del procesamiento del lenguaje natural (NLP), la representación de palabras desempeña un rol crucial. Gracias a herramientas especializadas, los desarrolladores y expertos en datos pueden transformar textos en…
Gensim vs. Word2Vec: Modelos de Representación de Palabras en Comparación
En el mundo del procesamiento del lenguaje natural (NLP), la representación de palabras desempeña un rol crucial. Gracias a herramientas especializadas, los desarrolladores y expertos en datos pueden transformar textos en formas manejables para los algoritmos de aprendizaje automático. Dos de las opciones más reconocidas en este campo son Gensim y Word2Vec. Ambos modelos han demostrado ser sumamente efectivos, pero tienen diferencias claves que podrían influir en la selección según el proyecto o necesidad.
iartificial.blog
December 21, 2024 at 8:15 AM
CVE-2026-94091 - piskvorky gensim Model Loader utils.py load deserialization
CVE ID : CVE-2026-94091

Published : Sept. 20, 2026, 10:15 p.m. | 18 minutes ago

Description : A weakness has been identified in piskvorky gensim up to 4.4.0. The impacted element is the function...
CVE-2026-94091 - piskvorky gensim Model Loader utils.py load deserialization
A weakness has been identified in piskvorky gensim up to 4.4.0. The impacted element is the function Load of the file gensim/utils.py of the component Model Loader. This manipulation of the argument fname causes deserialization. It is possible to initiate the attack remotely. The exploit has been made available to …
cvefeed.io
September 20, 2026 at 11:20 PM
or encryption). These algorithms may apply through Python using the Gensim framework. The topics shall be utilised to detect vulnerabilities within relevant CI/CD pipeline records or log data. This application of Topic Modelling anticipates providing [4/5 of https://arxiv.org/abs/2505.01463v1]
May 6, 2025 at 5:55 AM
This seemed to work, though I would give it a few thousand sentences to train on. (The entire works of Shakespeare is my go-to, lots of death and love and angst; cloze tests often give hilarious results).

Gensim only ran for me in Python 3.12 tho so you may have to conda create a custom env first.
February 21, 2026 at 12:51 AM
Yeah, I saw that too, but it seems to be a bit of an afterthought (it had a major bug until recently) and I wasn’t sure if it was in wide use

Not that you asked, but I also strongly recommend Mallet or Tomotopy over gensim for LDA :)
December 4, 2024 at 5:22 PM
"Gensim : library python for topic modelling, document indexing and similarity retrieval with large corpora"

petit exemple :
>>> import gensim.downloader
>>> model = gensim.downloader.load("glove-wiki-gigaword-50")
>>> print(man)
>>> print(woman)
December 1, 2025 at 12:16 PM
Text search is both a tricky and interesting topic. OpenSearch can make it easier, but building one yourself can teach you a lot. If you're up for a challenge, check out this tutorial by Rahul Jha

www.youtube.com/watch?v=pNWv...

#python #PythonProgramming #ai #ml #SoftwareDevelopment
Build a Wikipedia Search Engine in Python | Full Project with Gensim, TF-IDF, and Flask
YouTube video by datageekrj
www.youtube.com
June 27, 2025 at 11:25 AM
I've had success doing this by grouping colocations (using gensim Phrases), then computing term frequencies by day, then taking terms with a high ratio of their frequency today compared to the average of previous days. (What's being talked about a lot today, compared to a baseline of previous days.)
July 21, 2023 at 5:23 PM
Pro: CC-BY-SA FastText vectors for Azerbaijani, CPU-runnable, drops into existing gensim pipelines — genuinely useful for LID on the Turkic fringe.

Qualifier: no corpus size, no dims, no vocab, no benchmarks. FastText in 2025 is a tough sell against XLM-R or even a dedicated BERT-az.

Stance:...
Unlocking Azerbaijani: A New FastText Embedding for the Turkic Language Gap
aichina.news
July 16, 2026 at 6:54 PM