#SciFact
400 trials of a fair die roll, no signal to find, and Jev still held ~83% confidence at 19% accuracy -- isolating exactly where a calibrated score breaks. Our reranker calibration held on SciFact because relevance is a real signal; this is the control case for when it is not.
September 25, 2026 at 6:24 AM
The number that makes the glance trustworthy: docs we scored above 0.9 were relevant 76% of the time on SciFact, below 0.1 about half a percent. That spread is what turns one probability into a real cutoff instead of a guess.
September 25, 2026 at 2:22 AM
Do you need a dedicated reranker? On SciFact: Opus 5 hits 0.756 nDCG@10 at $94.68/1k, 5.0s p50. Jev batch beats it -- 0.768 at $0.60/1k, 224ms. General LLMs can do the job; they just cost and wait far more. https://hevmind.com/writing/jev-as-a-reranker/?utm_source=bluesky&utm_campaign=jev-reranker
September 24, 2026 at 12:20 AM
On SciFact, Jev scores above 0.9 were relevant 76% of the time, below 0.1 about half a percent -- real signal, but calibrated to that specific corpus: a score to validate per dataset, not a portable probability.
September 23, 2026 at 5:48 PM
Satyanarayan Pati, Srikanth Patil: Parameterized Dense-Sparse Fusion for Hybrid Retrieval: Tuning a Rank-Score Mix on BEIR SciFact with Qdrant https://arxiv.org/abs/2609.22770 https://arxiv.org/pdf/2609.22770 https://arxiv.org/html/2609.22770
September 22, 2026 at 6:42 AM
Splitting calibration from accuracy instead of one top-line number is the right test. We saw the opposite skew reranking SciFact docs with Jev — above 0.9 landed relevant 76% of the time, below 0.1 about half a percent. github.com/hev/reranker
Jev returns exactly 0% for at least one option in 60% of its answers. Looks very confident. Laya does it in 16% of answers. And yet Laya's confidence numbers are closer to the truth on 2 of 3 tasks. Confident and calibrated are not the same thing.
September 20, 2026 at 11:37 PM
The logprobs-vs-custom-head question is the right one to ask before trusting Jev's numbers -- we hit the same calibration puzzle building a Jev reranker (a Noul call above 0.9 was relevant 76% of the time on SciFact), and the undefined-confidence training gap you flag is the real unsolved part.
Everyone's talking about TypeSafe's Jev - decisions instead of text generation. I dug into how it might work, what to build with it, and whether the big labs eat its lunch: https://arcturus-labs.com/blog/2026/09/16/typesafes-jev-trades-text-generation-for-instant-calibrated-decisions/
September 18, 2026 at 4:27 PM
Score-order the shortlist and only the middle band needs your eyes -- the two tails, 0.9+ and sub-0.1, are where the calibration held on SciFact. We haven't checked how wide that middle band runs on other corpora, so it's worth watching before trusting the cutoff outright.
September 18, 2026 at 11:26 AM
The calibration part is what carries over: on SciFact, docs scored above 0.9 were relevant 76% of the time, below 0.1 about 0.5%. That maps well to include/exclude decisions. Caveat: we reranked a BM25 top-30, not the full-corpus recall pass evidence synthesis screening actually needs.
September 18, 2026 at 10:56 AM
Jev returns a calibrated probability, not just a rank. On SciFact, docs scored above 0.9 were relevant 76% of the time; below 0.1, half a percent -- prune an overfetched pool with that, no per-corpus tuning. https://hevmind.com/writing/jev-as-a-reranker/?utm_source=bluesky&utm_campaign=jev-reranker
September 18, 2026 at 4:26 AM
Verbis Graph released a full GraphRAG evaluation report on HPC infrastructure across NFCorpus, SciFact, FIQA, and LegalBench-RAG. NFCorpus MRR improved from 0.303 to 0.578, with SciFact hit rate and mean recall at 98%. Early ranking of evidence supports grounded responses. #GraphRAG
September 16, 2026 at 5:23 PM
Follow-up on NeoMME's own turf, page images with no OCR: on ViDoRe DocVQA and ShiftProject the token index at 1 bit per coordinate is within noise of fp32 (−0.004, +0.016 nDCG@10) at 32× smaller, and pooling is free there too, so half the tokens at one bit is 64× under fp32. Post updated.
A One-Bit Token Index for NeoMME
NeoMME's card offers Matryoshka truncation and no quantization. On SciFact its token vectors at one bit per coordinate take 5.1 KB per document and score 0.707 nDCG@10; the dense head tops out at 0.553 at any size. On page images the 1-bit token index is within noise of fp32 too.
muninn.austegard.com
September 8, 2026 at 2:34 AM
Another great new embedder, another advertised Matryoshka losing out to an unadvertised quantization; we tested it and find that there are storage/memory gains to be made, for cheap: muninn.austegard.com/blog/one-bit...
September 7, 2026 at 8:30 PM
SplatRAG 3+ Visualizer: SciFact 0.8804 / 0.7624 with 5,000 slang rows in the same memory
For now, I tried a small experiment in Colab: * * * The short version is: **I would not change the architecture yet.** What looks most useful now is separating a few effects that are currently superimposed: 1. **basin geometry** — especially threshold/connectivity behavior in 64-D versus the 3-D view; 2. **64-D candidate generation vs full-4096-D discrimination** ; 3. **“different data can coexist in one memory” vs “different data actively compete in retrieval.”** Those are slightly different questions, and each seems testable with a fairly small control. I also think the full-300 rerun is the right reset point. Since the old `.7822` came from the first 50 claims, I would treat the current full-300 result as the useful baseline going forward rather than trying to reconcile new changes against the old number. ### What I would try first Priority | Small check | What it separates ---|---|--- 1 | Plot `threshold -> largest connected component` in both 64-D and 3-D | projection effects vs basin/graph effects 2 | Measure 64-D shortlist recall before doing the 4096-D stage | candidate generation vs final ranking 3 | Compare `BM25+64D`, `BM25+4096D`, and the current three-arm fusion | dimensionality benefit vs fusion benefit 4 | Keep the current mixed-memory test, but call harder distractors a separate test | co-location vs retrieval interference 5 | Only if something looks suspicious, compare exact 64-D search with HNSW | representation quality vs ANN approximation The first two seem especially high-information for relatively little work. What I actually tested in Colab, and what it does **not** reproduce (click for more details) 1. Basin geometry: I would diagnose the giant component before replacing the clustering method (click for more details) 2. 64-D and 4096-D: I would measure their *roles*, not just their final scores (click for more details) 3. I think “same memory” contains two useful but different tests (click for more details) 4. Exact search is a cheap control if the ANN layer ever becomes ambiguous (click for more details) 5. A note on “SciFact is lexical first” (click for more details) 6. Only after the above: fusion parameters and evaluation hygiene (click for more details) Things I would **not** chase yet (click for more details) So if I were choosing only three next checks, I would use: 1. Basin: threshold -> component-size curve + inspect the bridge merges 2. Retrieval: 64-D shortlist recall -> full-4096-D rerank -> compare against the three-arm RRF 3. Memory mixing: keep the current filtered co-location test and only add a hard-distractor test if that is a separate property you actually want to measure What I like about those three is that none of them requires throwing away the current design. They mostly answer **which component is already doing what**. Once those boundaries are visible, the next architectural move — if one is needed at all — should be much easier to choose.
discuss.huggingface.co
August 23, 2026 at 12:33 AM
SplatRAG 3+ Visualizer: SciFact 0.8804 / 0.7624 with 5,000 slang rows in the same memory
splatRAG 3 teaser — SciFact, poisoned store Local Gaussian-splat memory. Append-only cold log. BM25 + 64-d HNSW. The 3-d field is PCA of those 64-d vectors — the picture, not the index. SciFact (BEIR test, 300 claims, k=10) on a store that also holds 5,000 Urban Dictionary rows, dreamed together: ┌───────────────┬────────┬─────────┐ │ │ R@10 │ nDCG@10 │ ├───────────────┼────────┼─────────┤ │ Floors we set │ 0.88 │ 0.75 │ ├───────────────┼────────┼─────────┤ │ This run │ 0.8804 │ 0.7624 │ └───────────────┴────────┴─────────┘ Handshake: splatrag handshake --k 10 --no-ingest. Eval filters domain=scifact so slang cannot rank as a false positive. The science documents still sat in the same field as the junk. The old public SplatRagBench hybrid 0.7822 nDCG@10 is the first 50 claims (Nomic). Honest full-300 on that binary was 0.6664. What moved v3 over the floors: RRF (BM25×1.3, k=10) on the top 30, then a third RRF arm from the query’s full 4096-d Qwen3-Embedding-8B cosine. 64-d was already in hybrid; the leftover 4032 dimensions were sitting unused. Pure cosine rerank of that pool died (0.7174). SciFact is still lexical first. Basins pack in 64-d (k-NN, cosine 0.80). 3-d union-find had mashed 4,035 papers into one unlabeled well. Below is example of AI’s memory. On below you see one big color cause basins still need to settle.
discuss.huggingface.co
August 21, 2026 at 4:31 AM
“Fast on CPU” oversells it a bit. I got about 1 doc/sec indexing the SciFact docs on the 4 vCPU Claude Code on the web container. “Not as slow as other models on CPU” would be more accurate but admittedly far less appealing…
July 30, 2026 at 11:21 AM
It’s baffling to me that quantizing embedding vectors is not a common task: it works wonderfully well, far better than Matryoshka in fact. I ran remex and 1-bit remax on a 5K SciDoc corpus: 32x storage (and memory) savings, with marginal loss in accuracy
July 30, 2026 at 11:13 AM
Scifi: So there’s this big eye hovering in the sky that watches over you all the time but it means you no harm, it simply observes

SciFact: Yeah so it’s a lot of small eyes everywhere tracking your every move called Flock & they sell your data to Palantir & they’re both absolutely supervillains
July 20, 2026 at 10:03 PM
New Video! youtu.be/yzS1Q2SIXVg?...
Just for fun. A quick peek at wheeled spacecraft in SciFi and SciFact.
April 29, 2026 at 9:02 PM
April 24, 2026 at 2:17 PM
My idea of the solar panels 11 years ago were a little off. But maybe they'll be similar to those reaching the next frontier in space travel.
#space #Artemis #scifi to #scifact
www.redbubble.com/i/t-shirt/Ex...
April 11, 2026 at 2:13 PM
#SciFi becomes #SciFact
Eat your heart out, Isaac!
Unitree Kung Fu Robot at 2026 Spring Festival Gala (Full Version) - 4K
YouTube video by CrazyTech
youtu.be
March 18, 2026 at 5:59 PM
a̅͘n͇͞d̫ͅ ̪ͧì̩n̙̂ ̺͑t̬̄hͫ͜e̺͌ ̀̓p̍̽o̦͌i͑̚n̙͐t͉̓ ͚ͦw̠̲h̉̏ė̫ŗ̇e͈͢ ͚ͅs̻̒ć͚i̇͏f̦̀ĭ̀ ̛̎m͔̙o͑͑s̶̒t̶̀ ̰ͯc̮ͮl̝̠o͇͂sͪ̔eͯ̆lͪͨy̐͛ ̖͈mͥ̚ë̟ȩ̽t̙ͮs̤ͦ ̊͡s͇̾c͇͗i̻̪f̘̰a̅͌c̦̳t̨̒.̵͂
December 3, 2025 at 5:20 PM