genderize (gender + country from a name): open weights, 98% accuracy, CPU only.
Now building in public: sample size + PICO from clinical abstracts.
EU-hosted API · free research keys · dbtool.it
Studies covering both sexes: 32%→65%. Aged participants: 18%→39%. 88% agreement with MeSH where present.
Full dataset (12.9M PMIDs) on request. @nrobinsongarcia.bsky.social may find this relevant.
Studies covering both sexes: 32%→65%. Aged participants: 18%→39%. 88% agreement with MeSH where present.
Full dataset (12.9M PMIDs) on request. @nrobinsongarcia.bsky.social may find this relevant.
"Gender assigned with tool X; expected error on our corpus: ~2% on European names, ~17% on Chinese (34% of sample)."
Same shape as a data-availability statement. Today almost nobody states it.
"Gender assigned with tool X; expected error on our corpus: ~2% on European names, ~17% on Chinese (34% of sample)."
Same shape as a data-availability statement. Today almost nobody states it.
Genuine question back: journals already require ethics statements at submission. Would a one-line 'name-tool error rate on your corpus' field in the checklist change the practice?
Genuine question back: journals already require ethics statements at submission. Would a one-line 'name-tool error rate on your corpus' field in the checklist change the practice?
Audit what we claim: names processed in memory, never stored — the schema holds counters, not payloads. 96 tests included.
The weights stay closed; the code is the proof.
#opensource #metascience
Audit what we claim: names processed in memory, never stored — the schema holds counters, not payloads. 96 tests included.
The weights stay closed; the code is the proof.
#opensource #metascience
Yes, please do: DMs are open. One line on the project is all we ask. You'll get a key with research quota, and if you ever run our predictions against your extraction pipeline on a shared sample, we'll publish the comparison whatever it shows.
Yes, please do: DMs are open. One line on the project is all we ask. You'll get a key with research quota, and if you ever run our predictions against your extraction pipeline on a shared sample, we'll publish the comparison whatever it shows.
DM us with one line about the project. First come, first served.
#bibliometrics #AcademicSky
DM us with one line about the project. First come, first served.
#bibliometrics #AcademicSky