#codebooks
wow Claude has gotten really good at navigating survey codebooks. kinda a workflow game changer in terms of time saved.
January 3, 2025 at 10:05 PM
oh dear. they used to reserve "low IQ" for racial tagging. are they burning their codebooks?
March 31, 2026 at 12:58 PM
First day back in the campus office. I cleaned up two research articles worth of data in excel, created two new codebooks for the variables, uploaded one article's data into SPSS, cleaned it up within that software program, and ran three tests, interpreted them in a separate notebook
August 22, 2025 at 9:54 PM
this little function that @jordannafa.bsky.social made for searching surveys with shitty codebooks has become something that I use very frequently
January 28, 2025 at 8:37 PM
The theft of vital codebooks and murder will lead Lissa and Miles back to London to ferret out their enemies and put an end to the smuggler’s consortium once and for all.

mybook.to/Betrayal-Boo...

#HistoricalFiction
September 29, 2026 at 9:26 PM
NEW ARTICLE: Want to use LLMs to extract information at scale in sociology? @eollion.bsky.social @oms279.bsky.social and C. Ton have you covered. Read "From Codebooks to Promptbooks," now in Sociological Methods & Research

doi.org/10.1177/0049...

(Preprint @socarxiv.bsky.social: osf.io/wjvfq_v1)
May 7, 2025 at 12:59 PM
This is, in my view, one of the risks of relying on secondary data. Students who participate in data collection (quant or qual) see that data are socially constructed and that careful interpretation is required. For those who only use pre-made data/codebooks, those are harder conclusions to draw.
February 22, 2025 at 2:39 PM
Another cool job was the guy who boarded captured German ships and rummaged around trying to find codebooks he could bring back to Bletchley.
July 19, 2026 at 1:39 PM
If you want to run MERFISH spatial transcriptomics and you search for the optimized probe codebooks, here they are: www.science.org/doi/10.1126/...
Boosting multiplexing capabilities for error-robust spatial transcriptomic methods using a set exchange approach
Using an algorithm, highly optimized error-robust codebooks for spatial transcriptomic methods can systematically be produced.
www.science.org
May 20, 2025 at 5:14 PM
The FBI was “frigging stumped,” said the same former bureau official. Until they realized that the Czech “illegal” was just an avid knitter. The codebooks weren’t codebooks, some “brilliant” new secret Soviet bloc commo device. They were leftover knitting patterns.

@bbl-astrophyscs.bsky.social
March 8, 2026 at 2:53 AM
Slightly paraphrasing @oms279.bsky.social during his talk at #COMPTEXT2025:

"The single most important use case for LLMs in sociology is turning unstructured data into structured data."

Discussing his recent work on codebooks, prompts, and information extraction: osf.io/preprints/so...
April 25, 2025 at 2:16 PM
I'm glad I have data and codebooks for 2011-2023 stashed. I'll have to upload them later.
January 31, 2025 at 11:13 PM
gonna once again shoutout this fantastic little function from @ajordannafa.com that I use basically all the time when navigating surveys with shitty or non-existent codebooks. returns a searchable dataframe of variables and labels.
November 25, 2025 at 8:18 PM
The thing is, which is somewhat suspicious to me, like this (alleged) NSA guy should know, is that one can use physical codebooks and other forms of cryptography and counterintrl *layered* with signal.
January 26, 2026 at 5:43 PM
In political science, it is rare to have good codebooks in replication materials. Figuring out what variables are is like excavating ancient ruins, sifting through crumbling papyrus to understand what "pg_party_id_2" is.
March 16, 2026 at 3:32 PM
i love when old players guides or codebooks used the fact they weren't authorized as tho it were some kind of selling point
February 24, 2026 at 5:47 PM
My amazing RA, Iro Eleftheriadou, started a gallery of codebooks for datasets that are publicly accessible but hard to find on the OSF
https://rubenarslan.github.io/codebook_gallery/

Let me know if we should add your data to the gallery or just send a pull...
Codebooks
rubenarslan.github.io
November 3, 2024 at 3:46 AM
I thought this was a clever and useful paper from Xiong, ... Hovy, El-Assady, Ash "Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification." Using LLMs to help humans refine their codebooks (before codebooks are fixed for the true annotation stage) arxiv.org/pdf/2507.05010
arxiv.org
July 23, 2025 at 1:00 PM
The theft of vital codebooks and murder will lead Lissa and Miles back to London to ferret out their enemies and put an end to the smuggler’s consortium once and for all.

mybook.to/Betrayal-Boo...

#HistoricalFiction
September 18, 2026 at 3:03 PM
Which incantation are you murmuring? They change out the codebooks every Turning
February 24, 2026 at 3:16 AM
Regina George may be a Mean Girl, but even she knows the truth: data need documentation to be reusable. This #MemeMonday we celebrate the importance of readmes, codebooks, methodologies, metadata -- whatever you call it, it's essential for reusing data of all types.
October 5, 2026 at 5:45 PM
Very excited that my paper with @katakeith.bsky.social is now out in @polanalysis.bsky.social. We investigate whether LLMs actually follow the instructions/definitions provided in codebooks, propose some diagnostics, and release a new evaluation dataset.
www.cambridge.org/core/journal...
Codebook LLMs: Evaluating LLMs as Measurement Tools for Political Science Concepts | Political Analysis | Cambridge Core
Codebook LLMs: Evaluating LLMs as Measurement Tools for Political Science Concepts
www.cambridge.org
September 19, 2025 at 1:45 PM
Pictures and story of #RN #RoyalNavy HMS Tartar #HMSTartar including 1941 capture of #Enigma & codebooks onboard #DE #Germany #weathership #Lauenburg #WBS3

laststandonzombieisland.com/2026/03/04/w...

#Navalhistory
May 18, 2026 at 3:16 PM
Alan Turing and the Bletchely Park team had broken the code used by these machines and the Allies didn’t want the Germans to find that out. Capturing a submarine with the machine and codebooks intact in 1944 was therefore an embarrassment.
July 24, 2026 at 10:43 AM
This came up in my class yesterday. We read @oms279.bsky.social’s wonderful paper “From Codebooks to Promptbooks” (highly recommend for teaching), but got stuck on the issue of non-random errors, which is a not a unique issue to LMs. Traditional methods…

Paper: journals.sagepub.com/doi/10.1177/...
And isn't the (limited but often only feasible) solution to both human/machine annnotations to just use expert test sets, assess error &take it into account in stats? Like what else is there to do? You seem to criticize DSL, but what's the better way? (besides just... not doing large scale research)
September 17, 2025 at 2:07 PM