Gordon Forbes
banner
gforb.bsky.social
Gordon Forbes
@gforb.bsky.social
I love Netflix for their data science blog and The BBC for their ggplot2 resources.
We've pre-printed an attempt at doing this in an academic CTU (see bsky.app/profile/gfor...)

The models were good for the descriptive bits, that largely summarise the protocol, but fell short at more statistics tasks.

Basicly it can save you time that you can then spend thinking about stats.
bsky.app
March 20, 2026 at 3:31 PM
tldr: Cool AI can write SAPs, but it makes mistakes, but it will save time.

Statisticians are still needed but we can cut out some of the boring bits to give more time for the interesting parts.

This is super scalable, get it touch if you want to try it out. We’ve got much more coming.
March 20, 2026 at 3:26 PM
Some intial user feedback sums it up:

"It allowed me to spend more time reflecting on the analysis itself rather than copying / pasting and doing other boring stuff etc... The main analysis model wasn't really there but I can do that."
March 20, 2026 at 3:25 PM
The good: Approx. 80% accuracy with very good accuracy for areas describing aspects of the trial design.

The bad: Accuracy was worse where statistical reasoning was required. Especially for sensitivity analysis where the AI proposes reasonable sounding analysis which are not what you’d want to do.
March 20, 2026 at 3:23 PM
Tool based on structured prompts, validated by expert statisticians against published SAP content guidelines.
March 20, 2026 at 3:22 PM
How is AI going to make this all easier?

And relatedly, how can we stop the increase in productivity form AI leading to an overwhelming ammount of methedologicaly questionable research.
July 30, 2025 at 2:08 PM
To me this is what age as catagorical means - eg. age 0-18, 19 - 65, 65+.

I would describe using individual integer ages as (ie. 1,2,3,4,5,6,...) as descrete age.

The 'big catagory' approach is used worryingly often - sometimes due to restrictions on the data.
July 8, 2025 at 2:40 PM
Hard disagree! It is acceptable and probably preferable for a group to write a paper without understanding the details of each other's work.

eg. I want to be able to write "the model was estimated with restricted maximum likelihood" in a stats section without explaining REML to my collaborators.
March 28, 2025 at 12:47 PM
I totally agree that, theoretically, it makes no sense to shrink parameters to zero.

In practice, if it means a model can be applied without collecting mostly irrelevant data, this can be a huge win. Especially in health when that extra data can involve invasive or expensive tests.
March 13, 2025 at 9:14 AM