#bm25
We hit the same thing swapping Jev, Clef and Perplexity's chooser model into our reranker tests: no new eval needed, just the same BM25-shortlist-and-nDCG@10 harness run again on a different model.
so LLMs have historically been quite difficult to assess, but decision models are quite easy

if you use one, you probably already have a benchmark or unit test suite that you can swap in another model and run through its paces

but, if it works better for you, then it’s better
October 9, 2026 at 8:43 PM
New video: How does a search engine find anything in milliseconds?

We build an inverted index from scratch, then follow the query "running shoes" through lookup, a zipper-walk intersection, BM25 ranking, shards and caches.

youtu.be/2tRKpVXkQPg
How Does a Search Engine Find Anything in Milliseconds
YouTube video by OpcodeAtlas
youtu.be
October 9, 2026 at 4:13 PM
What actually fits

FAISS as a local vector store, a lightweight web framework, BM25 for keyword search, and hosted calls for embeddings and generation — that combination is how…
https://pranjulrathour.scult.in/blog/free-tier-memory-budget-playbook

· Pranjul Rathour · pranjulrathour41@gmail.com
October 8, 2026 at 10:32 AM
RAG systems can fail when chunks lose the context that makes them meaningful. So here Rishi explains how contextual embeddings and hybrid search improve retrieval accuracy. You'll learn about BM25, vector search, reranking, metadata, and more.
www.freecodecamp.org/news/how-con...
October 8, 2026 at 8:01 AM
I actually use static embedding models like potion extremely fast and better at search than bm25
October 7, 2026 at 8:10 PM
MDKeyChunker asks whether one LLM call per chunk for metadata beats Markdown's free structure in RAG retrieval, testing it with qwen2.5:7b across BM25, dense, and hybrid setups on Qasper and FreshStack datasets. The…

#RAG #Markdown #LLM #InformationRetrieval
https://arxiv.org/abs/2603.23533
October 7, 2026 at 4:01 PM
You may have read the planetscale TIN vs. ParadeDB blog posts already. First the TIN announcement, then ParadeDB with improvements & showing benchmark flaws. Reminder: It's crazy hard to built fair comparisons when you know one product inside out, but not the other.

www.paradedb.com/blog/opening...
PlanetScale Released Text Search and We Have a Lot to Say (Part I)
Two BM25 optimizations and benchmark configuration changes inspired by PlanetScale's TIN benchmarks make ParadeDB's text search faster without changing its document identifiers.
www.paradedb.com
October 7, 2026 at 12:56 PM
BM25 covers the gap

BM25 is decades-old keyword search, and it's still the right tool for exact-term matching. Run it alongside dense search, not instead of it.
https://pranjulrathour.scult.in/blog/hybrid-retrieval-dense-bm25-rrf

· Pranjul Rathour · pranjulrathour41@gmail.com
October 7, 2026 at 10:32 AM
📦 besnovatyj/yii2-cms-search-tnt v1.3.0

Ядро сквозного поиска на TNTSearch для Yii2 CMS: полнотекстовый индекс BM25 со стеммингом русского и поиском с опечатками, размещённый в базе проекта. Подключается к фасаду besnovatyj/yii2-cms-sear...

🔗 https://github.com/besnovatyj/yii2-cms-search-tnt
October 7, 2026 at 8:01 AM
October 7, 2026 at 7:15 AM
Same trend we see in hybrid retrieval: embeddings miss exact-term clinical queries entirely, so BM25 stays in the pipeline instead of ever going pure-dense.
poll: do you use embeddings?

i’m finding a growing trend of people shedding embeddings for increasingly agentic flows

i still think embeddings are useful, but not as broadly or generally as we used to use them
With all the other news today (Mistral 4, local Qwen 3.8 Next Flash) this kind of got missed.

In process of migrating all embeddings, which is NOT a small task. But I ran 150 tests, and EmbeddingGemma2 outperforms (murders) BGE-M3 by 26 pts, and EG2 is multimodal. 🤯

huggingface.co/unsloth/embe...
October 7, 2026 at 2:17 AM
AlloyDBがBM25とRRFをサポートして、検索エンジンとして使えるようになったそう。
cloud.google.com/blog/product...
Native BM25 search in AlloyDB and Cloud SQL | Google Cloud Blog
A native BM25 index in AlloyDB and Cloud SQL provides full-text retrieval without having to provision, manage, or pay for separate systems.
cloud.google.com
October 7, 2026 at 12:29 AM
When hybrid retrieval actually matters

If your documents contain product names, version numbers, or specific figures a user might search verbatim, hybrid retrieval isn't optional —…
https://pranjulrathour.scult.in/blog/hybrid-retrieval-dense-bm25-rrf

· Pranjul Rathour · pranjulrathour41@gmail.com
October 6, 2026 at 7:32 AM
In our reranker test, Clef-flash scored 0.498 mean nDCG@10 as 30 pair calls. As one batch call, it scored 0.283, below the BM25 order at 0.404. Request shape mattered. https://hevmind.com/writing/jev-has-company/?utm_source=bluesky&utm_campaign=jev-has-company
October 6, 2026 at 12:06 AM
📦 survos/search-bundle 2.34.37

First-party Symfony search kernel and faceted UX with field-driven Doctrine, Meilisearch, Algolia, SQLite FTS5, PostgreSQL BM25, and Elasticsearch adapters.

🔗 https://github.com/survos/search-bundle
October 4, 2026 at 5:13 PM
Tutorial walks through building a local CPU-only harness to benchmark BM25, dense, hybrid RRF and reranking retrieval on your labeled queries, scoring Recall@k, nDCG@k, MRR, latency and cost.
October 4, 2026 at 11:29 AM
Bad web RAG often means chrome in the chunks.

Nav, cookies, and footers get embedded; BM25 starts ranking "Sign in".

Search (≤5) → Extract main-content markdown → Render only for JS shells.

agentsearchhq.com/extract?utm_...
October 4, 2026 at 2:53 AM
#PostgreSQL #in-process with #vectorsearch extensions directly #embedded in your #Python app. #OpenSource Apache 2.0 lic

- Similarly search: #pgvector + #pgvectorscale, high-performance storage
- #FullText search: #pg_textsearch, #BM25 and #ranking
#AI #LLM #Embeddings #Agents #RAG
GitHub - Ladybug-Memory/pgembed: Embedded PostgreSQL for Agents
Embedded PostgreSQL for Agents. Contribute to Ladybug-Memory/pgembed development by creating an account on GitHub.
github.com
October 3, 2026 at 2:12 PM
Why do indexes speed up queries? What if the search space is too large to enumerate?

Indexes and the Art of Searching covers B-trees, BM25, vector search, A* and more.

Try Chapter 1 free (PDF/EPUB):
buymeacoffee.com/asopitechia/...
Indexes and the Art of Searching: Make It Searchable, Search It, Explore It — Free Sample (Chapter 1)
Free sample of *Indexes and the Art of Searching: Make It Searchable, Search It, Explore It*: the front matter and Chapter 1, "What to Prepare Before Reading", as PDF and EPUB. The full book...
buymeacoffee.com
October 3, 2026 at 7:30 AM
Running Elasticsearch for search, ClickHouse for analytics, and a vector store for RAG means three systems to sync? SereneDB from puts BM25 search, vector search, hybrid ranking and OLAP queries behind one PostgreSQL-compatible SQL interface. It's Apache 2.0 licensed, works with psql and...
October 3, 2026 at 12:19 AM
But the way it is done now, normal search uses conventional ranking (some variant of TF-IDF/BM25?) while "AI search" uses semantic reranking of top 1000. Feels to me both modes should just switch to the same ranking algo. But Will librarians complain if even normal search uses semantic rerankers? /4
October 2, 2026 at 2:10 PM
qi-bin
Local-first query engine CLI for AI agents and humans (BM25 + vector search)
aur.archlinux.org
October 2, 2026 at 2:50 PM