dbtool
banner
dbtool.bsky.social
dbtool
@dbtool.bsky.social
Open models & data for meta-research.
genderize (gender + country from a name): open weights, 98% accuracy, CPU only.
Now building in public: sample size + PICO from clinical abstracts.
EU-hosted API · free research keys · dbtool.it
@crahal.com I'm back
September 12, 2026 at 5:12 PM
Same model on every year, so the trend is free of NLM indexing artifacts.

Studies covering both sexes: 32%→65%. Aged participants: 18%→39%. 88% agreement with MeSH where present.

Full dataset (12.9M PMIDs) on request. @nrobinsongarcia.bsky.social may find this relevant.
September 3, 2026 at 2:02 PM
Right — the databases have no gender field. So studies either infer it from names, or hand-code a sample, with its own representativeness problems. Either way the error is rarely reported — and these numbers feed real policy debates, the leaky pipeline above all. That's where the stakes are.
September 1, 2026 at 7:10 AM
Concretely: any study inferring author gender from names would state in methods one line like —

"Gender assigned with tool X; expected error on our corpus: ~2% on European names, ~17% on Chinese (34% of sample)."

Same shape as a data-availability statement. Today almost nobody states it.
August 31, 2026 at 3:30 PM
That framing sticks: negligence that compounds into an integrity problem the moment a reviewer can't check it.

Genuine question back: journals already require ethics statements at submission. Would a one-line 'name-tool error rate on your corpus' field in the checklist change the practice?
August 31, 2026 at 2:10 PM
The server code is now public (MIT): https://github.com/sheppard94g/dbtool-server

Audit what we claim: names processed in memory, never stored — the schema holds counters, not payloads. 96 tests included.

The weights stay closed; the code is the proof.

#opensource #metascience
August 31, 2026 at 1:24 PM
Update: your key is already generated and waiting (business tier, 5M/month) — we tried to DM it but your inbox only takes messages from people you follow. Send us any one-line DM and it lands in your hands within the hour.
August 31, 2026 at 12:08 PM
Fair correction, thanks — extraction it is.

Yes, please do: DMs are open. One line on the project is all we ask. You'll get a key with research quota, and if you ever run our predictions against your extraction pipeline on a shared sample, we'll publish the comparison whatever it shows.
August 31, 2026 at 12:05 PM
Free for research means free: we are giving out 2–3 API keys to active research projects — gender gap, bibliometrics, authorship studies. Indefinitely, in exchange for a citation.

DM us with one line about the project. First come, first served.

#bibliometrics #AcademicSky
August 31, 2026 at 9:10 AM
@grahamkendall.bsky.social transparency question in your area: gender-gap studies almost never report the error of the name-tool they used. We publish ours per country, weak figures first. Would you count an undisclosed 17% error rate as an integrity problem or just bad practice?
August 31, 2026 at 8:22 AM
@lincolnmullen.com your gender package set the standard for stating limits plainly (US, historical, binary). We tried to do the same for the rest of the world: per-country accuracy published including the weak rows — 83.5% on Chinese names. https://dbtool.it/academic
August 31, 2026 at 8:22 AM
@vergoulis.bsky.social open infrastructure question: our gender-from-name models are not open (they're the asset), but the per-country error tables and the method are, and academic access is free for a citation. Is that a defensible middle ground from where BIP! stands?
August 31, 2026 at 8:22 AM
@serhiinazarovets.bsky.social you post the bibliometrics papers worth reading, so this may interest you: per-country accuracy of name-based gender assignment, published including the weak figures. 400 held-out names per country, 12 countries. https://dbtool.it/academic
August 31, 2026 at 8:22 AM
@crahal.com you extract with LLMs; we run a dedicated per-country model. Would love to see the two compared on the same held-out names — 400 per country, ours published country by country including the weak rows. If you're game, we'll run and publish it whatever it says.
August 31, 2026 at 8:22 AM
@melindacmills.bsky.social for demographic work on authorships: we publish per-country accuracy of our name-gender models (97.5% Italian, 83.5% Chinese — the East Asian drop is real and stated). Group-level composition only, never individuals. Free for academic use, for a citation.
August 31, 2026 at 8:22 AM