https://www.postgresql.org/about/news/pgx-bm25-10-bm25-ranked-full-text-search-as-a-native-postgresql-index-3396/
#postgresql
https://www.postgresql.org/about/news/pgx-bm25-10-bm25-ranked-full-text-search-as-a-native-postgresql-index-3396/
#postgresql
BM25 search engine in Rust with Python bindings. All 5 BM25 variants, streaming add/delete/update, pre-filtered search (up to 600x faster), mmap indices, and auto-persistence.
github.com/lightonai/bm...
BM25 search engine in Rust with Python bindings. All 5 BM25 variants, streaming add/delete/update, pre-filtered search (up to 600x faster), mmap indices, and auto-persistence.
github.com/lightonai/bm...
- Install pg_search extension
- Create BM25 index of your data (e.g., CREATE INDEX ...)
- Query with BM25 scoring
- Install pg_search extension
- Create BM25 index of your data (e.g., CREATE INDEX ...)
- Query with BM25 scoring
"What makes BM25 worth understanding is not just that it works. It is that it works for knowable reasons. Every part of the formula has a clear interpretation. When a result is surprising, you can trace why.
"What makes BM25 worth understanding is not just that it works. It is that it works for knowable reasons. Every part of the formula has a clear interpretation. When a result is surprising, you can trace why.
cloud.google.com/blog/product...
cloud.google.com/blog/product...
github.com/Intelligent-...
github.com/Intelligent-...
"BM25 is a strong baseline for search. Ha! You thought I would start with something about vector search,
"BM25 is a strong baseline for search. Ha! You thought I would start with something about vector search,
arpitbhayani.me/blogs/bm25
arpitbhayani.me/blogs/bm25
Today, we are faster than TIN.
www.paradedb.com/blog/opening...
arxiv.org/abs/2502.0...
I haven't seen such a nice summary of how to reverse engineer what a model is actually "thinking" before. Also cool to see learned ranking is basically just a learned, semantic BM25.
arxiv.org/abs/2502.0...
I haven't seen such a nice summary of how to reverse engineer what a model is actually "thinking" before. Also cool to see learned ranking is basically just a learned, semantic BM25.
We’ll kick it off with *the* baseline: BM25 on MSMARCO. One line to download a pre-built index, one line to make a BM25 retriever, one line to search.
We’ll kick it off with *the* baseline: BM25 on MSMARCO. One line to download a pre-built index, one line to make a BM25 retriever, one line to search.
- MCP server
- Nearly-zero cost database forks
- Hybrid vibe search (vector + BM25)
www.tigerdata.com/agentic-post...
- MCP server
- Nearly-zero cost database forks
- Hybrid vibe search (vector + BM25)
www.tigerdata.com/agentic-post...
• AI is getting smarter, albeit gradually.
• Context length is getting longer—now reaching up to around 2 million tokens in some models.
• Information retrieval: Knock, knock. Hello? Is anyone seriously working on this? Are we really back to using grep and BM25?
• AI is getting smarter, albeit gradually.
• Context length is getting longer—now reaching up to around 2 million tokens in some models.
• Information retrieval: Knock, knock. Hello? Is anyone seriously working on this? Are we really back to using grep and BM25?
Instead of going through all the effort to create vector embeddings - using one of those models like OpenAI’s text-embedding-3-large - consider starting with BM25 (Best Matching 25). It might be all you need.
Instead of going through all the effort to create vector embeddings - using one of those models like OpenAI’s text-embedding-3-large - consider starting with BM25 (Best Matching 25). It might be all you need.
www.freecodecamp.org/news/how-con...
www.freecodecamp.org/news/how-con...