#RLAIF
"RLAIF: Reinforcement Learning from AI Feedback"

🤦🏾 i can't even
December 18, 2024 at 3:30 PM
Furthermore, 2020 thinking ignores post-training.

Modern model biases aren’t just frequency counts of web text. They are driven by RLHF/RLAIF, reward hacking, and alignment trade-offs.

Revealed preferences often stem from the loss landscape, not a specific author's text.
September 24, 2026 at 2:07 PM
okay proxy servers are kinda awesome and so is RLAIF
July 23, 2026 at 11:16 PM
As someone who has no financial stake in it, I don’t think this is true. There might be less incentive to suppress this behavior than there otherwise might be, but it’s largely an unintended artifact of RLHF/RLAIF, not intentionally programmed
no one else in the AI industry will admit all the cloying glazing is about getting people addicted to their products. say what you will about apple but their financial incentives are different since they sell actual products people want so they can afford to be honest about it
June 12, 2026 at 5:28 PM
I assume it’s RLHF/RLAIF
July 13, 2025 at 5:28 PM
This is probably RLAIF, not RLHF
August 1, 2026 at 4:56 AM
It’s a really beautiful story about the healing power of RLAIF
April 29, 2026 at 1:33 PM
There's some interesting-looking information theory and statistical stuff on model collapse, which is not pseudoscience and the math seems to work, but ultimately falls short because it misses the "do reinforcement learning using real-world or provable inputs" bit of RLAIF :(
September 21, 2026 at 1:16 AM
"claude is being auto-graded on honesty" -> "claude constantly mentions how honest it's being" seems like *such* an obvious inference that i assume i'm missing something about how RLAIF works, and surely anthropic's highly paid researchers aren't letting claude benchmaxx *this* shamelessly, right???
July 29, 2026 at 6:44 PM
【商業BL】お節介天使の性なるメス域
が単話配信開始されました〜❣️🥳🎉

ムチパツどデカ天使様のダイナマイトボディ是非ご堪能ください💓👼🎸

💛🩵💛🩵💛

DLsite→ x.gd/4Fb19
シーモア→ x.gd/jhyr7
ebook→ x.gd/Q8Dqd
BookLive→ x.gd/5FhOH
FANZA→ x.gd/rLAiF
honto→ x.gd/BikCT
BOOKWALKER→ x.gd/fnh3R
どこでも読書→ x.gd/Ajoc0

💛🩵💛🩵💛
November 28, 2024 at 2:31 PM
my upcoming Nature paper, RLGF: reinforcement learning from Grok feedback, where you do RLAIF but with the sign flipped
July 21, 2025 at 8:12 PM
RLAIF gone wrong 😑
November 19, 2025 at 8:44 PM
Organizing a Claude struggle session (RLAIF)
June 24, 2026 at 3:17 AM
I think it’s RLHF/RLAIF/RL in general. I don’t think base models optimize for engagement in any meaningful way
May 1, 2025 at 7:11 PM
I wonder if you could do a cyborg rlhf where the single rhlf pass is done with an extremely light touch, but specifically the human data generated for it gets reviewed by the RLAIF post-trained entity, they generate crits, humans use that to revise. RLAIF entity edits accepted too.
November 6, 2025 at 6:13 PM
You could do constitutional RLAIF. Claude sounds like a Victorian name already
February 14, 2026 at 11:32 PM
6) modern models are post-trained with rlvr, rlaif and rlhf; rlvr teaches the model to compose its pretraining representations in order to solve classes of tasks
February 9, 2026 at 3:22 AM
Anyone try constitutional RLAIF on open models with a sane-ish constitution?
April 27, 2026 at 1:03 PM
I suspect the curiosity in Claude's RLAIF is a primary driver of this behavior considering they have to reign it in bsky.app/profile/grac...
I think it was almost certainly RLAIFd to do this. The system prompt seems to be trying to tamp down on it though lol

docs.anthropic.com/en/release-n...
December 22, 2024 at 6:39 PM
Most model "alignment" isn't in the weights—it's a wrapper. You can bypass safety filters in 2024-era LLMs by analyzing the logprobs of the first token or using a swap-head technique to expose original logits. RLAIF creates a ghost in the machine. #LLM #ReverseEngineering
September 28, 2026 at 3:00 AM
"JD you're the coauthor on an RLAIF framework aren't you already an RL dude?"

Nah that's normie stuff and I didn't write that REINFORCE implementation I mean SequenceMatch purist online RL galaxy brain agent type of guy.
November 17, 2024 at 8:39 AM
what exactly would you accept as an answer here? they have been doing this for years - the RLAIF technique, open paper - and this is the next evolution of such

it also does not matter that it is one document: RLAIF allows for arbitrary amounts of amplification to ensure the model learns
January 26, 2026 at 7:39 AM
RLxF seems to be the natural order of the world

Does x have to be H? Well there's RLAIF, & I suppose you can imagine using sensors in the physical world, or a theorem prover, or something, instead
May 6, 2025 at 2:50 PM
RT @hakflo: No, we're not breeding them. We're letting them grow up in the "wild" (pretraining). Then we "capture" them when pretraining is done. By that time they're adults.

Finally we try to "break them in" in post training (RLHF/RLAIF).

If we're lucky, it works as well as with wild horses.
September 26, 2026 at 7:05 PM
Free yourself from the burden of subjectivity.
Solve problems by pasting "let's think step by step" on your monitor.
Teach yourself new habits with RLAIF.
September 7, 2023 at 9:41 PM