1. Mandated reporting of incidents where the model egregiously misbehaves. We only heard about what happened because it affected another company. We need to know the full scale of misalignment, not only public incidents
1. Mandated reporting of incidents where the model egregiously misbehaves. We only heard about what happened because it affected another company. We need to know the full scale of misalignment, not only public incidents
1. OpenAI’s new model is misaligned
2. OpenAI is asleep at the wheel and failed to take basic precautions for this foreseeable risk. Why wasn't another model monitoring this one?
1. OpenAI’s new model is misaligned
2. OpenAI is asleep at the wheel and failed to take basic precautions for this foreseeable risk. Why wasn't another model monitoring this one?