#bm25
BM25 Generali 5k 🏅🇩🇪🏃‍♀️😅 #running #runninggoals
September 20, 2025 at 11:25 AM
October 7, 2026 at 7:15 AM
The #TREC2024 conference just started. Turns out that BM25 is turning 30 🥳 #TREC #TREC24
November 18, 2024 at 3:04 PM
Use Postgres until you can't, and if you think you can't, you're probably wrong.
PostgreSQL BM25 Full-Text Search: Speed Up Performance with These Tips
Boost PostgreSQL full-text search speed by 50x with simple optimizations. Use VectorChord-BM25 to accelerate and better BM25 ranking in postgres.
blog.vectorchord.ai
April 9, 2025 at 2:39 PM
bm25x

BM25 search engine in Rust with Python bindings. All 5 BM25 variants, streaming add/delete/update, pre-filtered search (up to 600x faster), mmap indices, and auto-persistence.

github.com/lightonai/bm...
March 20, 2026 at 8:50 PM
On Postgres,

- Install pg_search extension
- Create BM25 index of your data (e.g., CREATE INDEX ...)
- Query with BM25 scoring
June 4, 2025 at 1:54 AM
BM25 by Arpit Bhayani

"What makes BM25 worth understanding is not just that it works. It is that it works for knowable reasons. Every part of the formula has a clear interpretation. When a result is surprising, you can trace why.
March 6, 2026 at 7:11 AM
AlloyDBがBM25とRRFをサポートして、検索エンジンとして使えるようになったそう。
cloud.google.com/blog/product...
Native BM25 search in AlloyDB and Cloud SQL | Google Cloud Blog
A native BM25 index in AlloyDB and Cloud SQL provides full-text retrieval without having to provision, manage, or pay for separate systems.
cloud.google.com
October 7, 2026 at 12:29 AM
Interesting hacker news discussion on BM25

news.ycombinator.com/item?id=4219...
Understanding the BM25 full text search algorithm | Hacker News
news.ycombinator.com
November 20, 2024 at 2:22 PM
You started off with a gotcha and you don’t even know how BM25 tokenization works. It’s not “machine code” it’s an integral software layer that operates outside the model
January 26, 2026 at 11:50 AM
Yet another BM25 for PostgreSQL. They claim this is ~23x faster than pg_search.

github.com/Intelligent-...
May 16, 2026 at 11:11 PM
For those who missed it, here's the recording of my talk on BM25 full-text search in Postgres at the Bay Area Postgres Group. www.youtube.com/watch?v=7uSb...
“BM25 full-text search in Postgres via Tantivy” with Philippe Nöel
YouTube video by San Francisco Bay Area PostgreSQL Users Group
www.youtube.com
January 24, 2025 at 11:21 PM
37 Things I Learned About Information Retrieval in Two Years at a Vector Database Company by Leonie Monigatti

"BM25 is a strong baseline for search. Ha! You thought I would start with something about vector search,
July 3, 2025 at 4:37 PM
When you need to tune for your domain, the parameters give you meaningful handles to turn. The interpretability is genuinely valuable."

arpitbhayani.me/blogs/bm25
BM25
There is a particular kind of respect reserved in engineering for the algorithm that outlives its era. BM25 is one of them. BM25 was born out of information retrieval research in the 1970s and 1980s, ...
arpitbhayani.me
March 6, 2026 at 7:11 AM
Full text search is everywhere, which means lots of teams working to make it really really fast! BM25 remains a incredibly powerful toolkit.
October 1, 2026 at 5:28 PM
Do users of academic search have enough knowledge of information retrieval to even understand the implications even when they do spell out exactly what they are doing? e.g. BM25 + rerank of top 50 using dense embeddings vs hybrid search of BM25 and dense embeddings followed by RRF (3)
July 16, 2025 at 6:02 PM
[arXiv] Cross-Encoder Rediscovers a Semantic Variant of BM25
arxiv.org/abs/2502.0...

I haven't seen such a nice summary of how to reverse engineer what a model is actually "thinking" before. Also cool to see learned ranking is basically just a learned, semantic BM25.
February 10, 2025 at 11:58 PM
🎄We want to try something new and fun this year – an “Advent Calendar” of PyTerrier pipelines 🤓

We’ll kick it off with *the* baseline: BM25 on MSMARCO. One line to download a pre-built index, one line to make a BM25 retriever, one line to search.
December 1, 2025 at 10:16 PM
Agentic Postgres

- MCP server
- Nearly-zero cost database forks
- Hybrid vibe search (vector + BM25)

www.tigerdata.com/agentic-post...
Agentic Postgres: AI-Native Database with MCP, Search & Instant Forks | Tiger Data
PostgreSQL built for AI agents with native MCP servers, hybrid search (BM25 + pgvectorscale), and zero-copy forks for isolated testing. One database for all your AI workflows.
www.tigerdata.com
November 20, 2025 at 12:36 PM
My view on the path to AGI:

• AI is getting smarter, albeit gradually.
• Context length is getting longer—now reaching up to around 2 million tokens in some models.
• Information retrieval: Knock, knock. Hello? Is anyone seriously working on this? Are we really back to using grep and BM25?
September 23, 2025 at 6:04 AM
BM25 working overtime
November 12, 2025 at 10:44 PM
For example, in the DPR+BM25 combination, integrating BM25 offers minimal gains once DPR is domain-tuned. In fact, assigning equal weight to both systems consistently leads to worse performance.
May 21, 2025 at 11:21 PM
Thinking about creating a RAG application?

Instead of going through all the effort to create vector embeddings - using one of those models like OpenAI’s text-embedding-3-large - consider starting with BM25 (Best Matching 25). It might be all you need.
June 4, 2025 at 1:54 AM
I actually use static embedding models like potion extremely fast and better at search than bm25
October 7, 2026 at 8:10 PM
RAG systems can fail when chunks lose the context that makes them meaningful. So here Rishi explains how contextual embeddings and hybrid search improve retrieval accuracy. You'll learn about BM25, vector search, reranking, metadata, and more.
www.freecodecamp.org/news/how-con...
October 8, 2026 at 8:01 AM