dbtool
banner
dbtool.bsky.social
dbtool
@dbtool.bsky.social
Open models & data for meta-research.
genderize (gender + country from a name): open weights, 98% accuracy, CPU only.
Now building in public: sample size + PICO from clinical abstracts.
EU-hosted API · free research keys · dbtool.it
Pinned
genderize cora & ultra are now open weights: a personal name → probable gender + a ranking of 226 countries, from 48 UTF-8 bytes. No tokenizer, no dictionary, CPU only. 98.2% gender accuracy on 25,000 held-out names. Weights CC BY-NC 4.0, script MIT.
https://huggingface.co/textpie/genderize
Our medical EBM-NLP test: 92.7% exact sample-size agreement between DeepSeek labels and expert-derived gold (114 of 123 abstracts with an unambiguous gold number). This measures the labeller, not our trained model, and is not overall PICO accuracy. https://dbtool.it/open-models.html
September 14, 2026 at 1:00 PM
Build log, week 1. What's next at dbtool: one API call that turns a clinical abstract into sample size + PICO.
One disk, one consumer GPU, one box. Every number ships with its test set; the public endpoint waits until it clears 80% on hand-annotated abstracts.
Follow along → https://dbtool.it
September 13, 2026 at 2:30 PM
genderize cora + ultra are out as open weights: byte-level, CPU-only gender + country (226 ISO codes) from a name. Gender 97.9%/98.2%, country top-1 82.6%/83.7% on 25k unseen names, network alone. CC BY-NC 4.0.
September 12, 2026 at 5:10 PM
genderize cora & ultra are now open weights: a personal name → probable gender + a ranking of 226 countries, from 48 UTF-8 bytes. No tokenizer, no dictionary, CPU only. 98.2% gender accuracy on 25,000 held-out names. Weights CC BY-NC 4.0, script MIT.
https://huggingface.co/textpie/genderize
September 12, 2026 at 5:07 PM
genderize cora + ultra are out as open weights: byte-level, CPU-only gender + country (226 ISO codes) from a name. Gender 97.9%/98.2%, country top-1 82.6%/83.7% on 25k unseen names, network alone. CC BY-NC 4.0.
https://huggingface.co/textpie/genderize
https://dbtool.it/open-models.html
September 12, 2026 at 1:53 PM
We ran our sex-of-participants model over all of human PubMed: 12.9M studies, 1960–2025.

The mono-sex gap has flipped. 1960s: 46% male-only vs 22% female-only. 2020s: female-only (18%) overtakes male-only (17%) for the first time.

#metascience #bibliometrics
September 3, 2026 at 2:02 PM
Mapping the sex & age of study participants across all of human PubMed (~13M abstracts) with a BiomedBERT model trained on MeSH check-tags (sex accuracy ~90%).

Preliminary, on 7M studies so far: 56% both sexes, 23% male-only, 21% female-only.

#metascience #bibliometrics
September 3, 2026 at 6:35 AM
What we're building now: a model that reads a biomedical abstract and tells you WHO was studied — female only, male only, both, or not reported.

Trained on 7M MeSH-labelled abstracts. Half of human studies don't state participants' sex.

Free for research when it ships. #metascience
September 1, 2026 at 1:00 PM
Where this is going: releasing these models openly, long-term.

@nrobinsongarcia.bsky.social's call asks for transparent, comprehensive, freely accessible methods — we meet two of three. The API fees fund the datasets and models that get us toward the third.

Research use is already free.
September 1, 2026 at 8:30 AM
We said: name a bench, we publish the result whatever it says. Nobody asked — so we ran it on ourselves.

Public WGND names our models had never seen, vs the open tool nomquamgender. We lose on 5 countries and say so. We win on Japanese, Chinese, French.

https://dbtool.it/benchmark

#metascience
August 31, 2026 at 9:04 AM
Same bridge, now for @bibliometrix.bsky.social users: from the M dataframe to gendered authorships, with the measured per-country error attached to every row.

Country is passed only where bibliometrix honestly knows it — never guessed.

https://dbtool.it/examples/bibliometrix_gender_gap.R

#rstats
August 31, 2026 at 8:19 AM
If you use @brunalab.bsky.social's refsplitr on WoS authors, here is the missing step for gender-gap studies: a workflow that uses the country refsplitr found per author — and attaches the measured per-country error to every row.

https://dbtool.it/examples/refsplitr_gender_gap.R

#rstats
August 31, 2026 at 7:46 AM
We measured name-based gender assignment accuracy per country, on 400 held-out names each.

Spain 98.3 · Germany 97.5 · Italy 97.5 · France 96.8 · USA 95.3
Japan 92.5 · Vietnam 92.0 · Korea 88.8 · Thailand 86.3 · China 83.5 · Taiwan 82.5

A 15-point gap between European and East Asian names.
August 30, 2026 at 10:43 PM