Benjamin Meindl
bmeindl.bsky.social
Benjamin Meindl
@bmeindl.bsky.social
Head of Agentic AI Platform at IU International University, building AI tutoring at university scale. Notes on what the evidence on AI in education actually says, and on running agents for real work. Views my own.
Opus 5.5 update: no regrets. For my agent work, it’s crazy good and reasonably fast. And my Claude Max plan now lasts longer: I hit the usage limits less often. So I’m back on Claude, and I really like it again.
September 29, 2026 at 1:57 PM
Simplicity first. The most interesting AI model I tried this month nails it.

Jev writes no text. It only answers closed questions (yes/no, pick one, score), for a tiny fraction of a cent each. Not smarter than an LLM, but it forces you to split fuzzy decisions into small, clear ones.

More ↓
A Model That Doesn't Write
Notes on simplicity first, from testing TypeSafe's Jev in my own agent tooling
medium.com
September 29, 2026 at 8:05 AM
My agent work had drifted toward Codex. Claude Opus 5.5 is pulling me back to Claude for real work. Reminder to self: keep the setup portable. Switching tools should be a small task, not a rebuild.
September 28, 2026 at 4:23 PM
AI writes a practice question and labels it "hard". In an observational Drexel study (7,888 student answers), the labels tracked the AI's own Bloom labels (ρ=.90), barely the difficulty measured from answers (ρ=.06). Many items hit a ceiling, but still:

A label isn't a measurement.

More ↓
Against What, Measured When — Weekly Notes for EdTech Platform Builders, September 28, 2026
A tutor tested against students with consumer AI, agent incidents disclosed on a schedule, difficulty labels that fit the AI, not students
medium.com
September 28, 2026 at 4:00 PM
An approval button is only useful if the agent does what you approved.

The Loopjacking study showed how, in tested agent workflows, the action could be swapped after the click. For EdTech builders: bind approval to the exact action.

More in this week’s EdTech AI notes ↓
The Claim That Survives the Source
Weekly Notes for EdTech Platform Builders · September 21, 2026
medium.com
September 22, 2026 at 12:25 PM
55 CS students, three C tasks, ChatGPT vs web search. Coding scores 89% vs 69%. Cued recall 41% vs 53% right away, 39% vs 52% at 48h: the gap was there from the start, it didn't widen. Self-attributed ownership 45% vs 81%.

Small study. It shows the deliverable and what stuck moving apart.
Your Programming Students' Cognition with ChatGPT: Higher Performance, Lower Retention, and Reduced Ownership
Bergh, Tag, Vassar, Renzella (arXiv, Sep 2026). Between-subjects experiment, 59 enrolled / 55 analysed, pseudo-random assignment, three introductory C tasks.
arxiv.org
September 22, 2026 at 8:43 AM
Two randomised AI-tutoring trials in two weeks. Vendor-funded: g = 0.33 on a four-week post-test. Independently funded: a suggestive +3.1 pts on the practiced question a week later (mastery group), near zero on the unpracticed one.

"AI tutoring works" is not the sentence that survives the papers.
The Claim That Survives the Source — Weekly Notes for EdTech Platform Builders
Two randomised tutoring trials, a 48-hour recall gap, an approval spent on the wrong action, and three incompatible meanings of "open"
medium.com
September 21, 2026 at 8:21 PM