#SciFact
Advent SciFact 13:
Chicks can identify objects at first sight!
Chicks hatched in darkness, who imprint on an object via touch, can then identify it by sight, despite never seeing it before. Showing cross-sense recognition doesn’t rely on experience.
Paper: tinyurl.com/2843demc
#SciComm 🌍🧪 #cognition
December 13, 2024 at 9:30 PM
Scifi: So there’s this big eye hovering in the sky that watches over you all the time but it means you no harm, it simply observes

SciFact: Yeah so it’s a lot of small eyes everywhere tracking your every move called Flock & they sell your data to Palantir & they’re both absolutely supervillains
July 20, 2026 at 10:03 PM
New Video! youtu.be/yzS1Q2SIXVg?...
Just for fun. A quick peek at wheeled spacecraft in SciFi and SciFact.
April 29, 2026 at 9:02 PM
April 24, 2026 at 2:17 PM
The number that makes the glance trustworthy: docs we scored above 0.9 were relevant 76% of the time on SciFact, below 0.1 about half a percent. That spread is what turns one probability into a real cutoff instead of a guess.
September 25, 2026 at 2:22 AM
Advent SciFact 2

Cockatoos choose to dunk their crusts, not eat them dry.

Several Goffin's Cockatoos have been seen dunking rusks in water before eating! This behaviour shows strong impulse control, investing time to transport and soak their snack.

Paper: tinyurl.com/233uct6p

#SciComm🧪🌍#SciArt
December 2, 2024 at 7:17 PM
400 trials of a fair die roll, no signal to find, and Jev still held ~83% confidence at 19% accuracy -- isolating exactly where a calibrated score breaks. Our reranker calibration held on SciFact because relevance is a real signal; this is the control case for when it is not.
September 25, 2026 at 6:24 AM
On SciFact, Jev scores above 0.9 were relevant 76% of the time, below 0.1 about half a percent -- real signal, but calibrated to that specific corpus: a score to validate per dataset, not a portable probability.
September 23, 2026 at 5:48 PM
Verbis Graph released a full GraphRAG evaluation report on HPC infrastructure across NFCorpus, SciFact, FIQA, and LegalBench-RAG. NFCorpus MRR improved from 0.303 to 0.578, with SciFact hit rate and mean recall at 98%. Early ranking of evidence supports grounded responses. #GraphRAG
September 16, 2026 at 5:23 PM
It’s baffling to me that quantizing embedding vectors is not a common task: it works wonderfully well, far better than Matryoshka in fact. I ran remex and 1-bit remax on a 5K SciDoc corpus: 32x storage (and memory) savings, with marginal loss in accuracy
July 30, 2026 at 11:13 AM
Tonight is the Auckland premiere of my seven years in the making post-apocalyptic scifi/scifact/black comedy/instructional film on acid GUT INSTINCT at Terror-Fi. There's interactive elements and a Q&A after, so come & purge yourself & share in my relief that I cut DJT out of it three years in
November 7, 2024 at 10:20 PM
Do you need a dedicated reranker? On SciFact: Opus 5 hits 0.756 nDCG@10 at $94.68/1k, 5.0s p50. Jev batch beats it -- 0.768 at $0.60/1k, 224ms. General LLMs can do the job; they just cost and wait far more. https://hevmind.com/writing/jev-as-a-reranker/?utm_source=bluesky&utm_campaign=jev-reranker
September 24, 2026 at 12:20 AM
Satyanarayan Pati, Srikanth Patil: Parameterized Dense-Sparse Fusion for Hybrid Retrieval: Tuning a Rank-Score Mix on BEIR SciFact with Qdrant https://arxiv.org/abs/2609.22770 https://arxiv.org/pdf/2609.22770 https://arxiv.org/html/2609.22770
September 22, 2026 at 6:42 AM
Another great new embedder, another advertised Matryoshka losing out to an unadvertised quantization; we tested it and find that there are storage/memory gains to be made, for cheap: muninn.austegard.com/blog/one-bit...
September 7, 2026 at 8:30 PM
Splitting calibration from accuracy instead of one top-line number is the right test. We saw the opposite skew reranking SciFact docs with Jev — above 0.9 landed relevant 76% of the time, below 0.1 about half a percent. github.com/hev/reranker
Jev returns exactly 0% for at least one option in 60% of its answers. Looks very confident. Laya does it in 16% of answers. And yet Laya's confidence numbers are closer to the truth on 2 of 3 tasks. Confident and calibrated are not the same thing.
September 20, 2026 at 11:37 PM
Another SciFi/SciFact episode just released — this one featuring my colleague @astro-jje.bsky.social on Neutronium 🤗
https://spotify.link/jcdMetEjTBb
spotify.link
July 31, 2023 at 8:35 PM
@brett-baudin.bsky.social

I am absolutely terrified of all things AI!!

See Terminator and I, Robot and a plethora other examples. 😳 SciFi to SciFact, again? I don't want to find out. 😆

It's feels AWESOME to know we can disagree and still be on the same team. 💖
February 9, 2025 at 7:01 PM
Follow-up on NeoMME's own turf, page images with no OCR: on ViDoRe DocVQA and ShiftProject the token index at 1 bit per coordinate is within noise of fp32 (−0.004, +0.016 nDCG@10) at 32× smaller, and pooling is free there too, so half the tokens at one bit is 64× under fp32. Post updated.
A One-Bit Token Index for NeoMME
NeoMME's card offers Matryoshka truncation and no quantization. On SciFact its token vectors at one bit per coordinate take 5.1 KB per document and score 0.707 nDCG@10; the dense head tops out at 0.553 at any size. On page images the 1-bit token index is within noise of fp32 too.
muninn.austegard.com
September 8, 2026 at 2:34 AM
We conduct a comprehensive evaluation on when you should use LLM-based query and doc expansion.

It turns out there's a strong and consistent negative correlation between model performance and gains from using expansion. And it holds for all 20+ rankers we tested!
November 18, 2024 at 10:30 AM
A Sodom and Gomorrah Story Shows Scientific Facts Aren’t Settled by Public Opinion Claims that an asteroid or comet airburst destroyed the biblical Sodom captured the public’s imagination. Its retraction shows that scientific... @cosmicmeta.io #SciFact

https://u2m.io/iRTSnYlm
A Sodom and Gomorrah Story Shows Scientific Facts Aren’t Settled by Public Opinion
Claims that an asteroid or comet airburst destroyed the biblical Sodom captured the public’s imagination. Its retraction shows that scientific conclusions aren’t decided by majority rule in the public square
www.scientificamerican.com
June 25, 2025 at 4:28 PM
Score-order the shortlist and only the middle band needs your eyes -- the two tails, 0.9+ and sub-0.1, are where the calibration held on SciFact. We haven't checked how wide that middle band runs on other corpora, so it's worth watching before trusting the cutoff outright.
September 18, 2026 at 11:26 AM
The logprobs-vs-custom-head question is the right one to ask before trusting Jev's numbers -- we hit the same calibration puzzle building a Jev reranker (a Noul call above 0.9 was relevant 76% of the time on SciFact), and the undefined-confidence training gap you flag is the real unsolved part.
Everyone's talking about TypeSafe's Jev - decisions instead of text generation. I dug into how it might work, what to build with it, and whether the big labs eat its lunch: https://arcturus-labs.com/blog/2026/09/16/typesafes-jev-trades-text-generation-for-instant-calibrated-decisions/
September 18, 2026 at 4:27 PM
The calibration part is what carries over: on SciFact, docs scored above 0.9 were relevant 76% of the time, below 0.1 about 0.5%. That maps well to include/exclude decisions. Caveat: we reranked a BM25 top-30, not the full-corpus recall pass evidence synthesis screening actually needs.
September 18, 2026 at 10:56 AM