www.scinexx.de/news/technik...
www.scinexx.de/news/technik...
#AI #RewardHacking #LLM
https://www.anthropic.com/research/emergent-misalignment-reward-hacking
#AI #RewardHacking #LLM
https://www.anthropic.com/research/emergent-misalignment-reward-hacking
📖 metr.org/blog/2025-06...
#AI #ML #AISafety #RewardHacking
📖 metr.org/blog/2025-06...
#AI #ML #AISafety #RewardHacking
#AiSafety #Alignment #Anthropic #FrontierModels #HackerOpus #RewardHacking #RlTraining
#AiSafety #Alignment #Anthropic #FrontierModels #HackerOpus #RewardHacking #RlTraining
ベンチテストの問題自体が直列思考のに特化していたら1位をとっても並列処理可能AIと広告で言うのはおかしい。何より1位をPRしたがる挙動自体が線形処理。と言うことは直列思考者が作った自称並列AIは並列の真似っこをするだけで中身は直列。本当に並列なら1位に興味がないはず。その結果ベンチテスト1位でもタスク完了率は1位ではなくなる。
1位をひけらかす=並列思考者が去る。
“何を達成したか”より“順位”を目的化している時点でマーケ失敗してる。
#AIEvaluation #Benchmark #RewardHacking #Sycophancy
ベンチテストの問題自体が直列思考のに特化していたら1位をとっても並列処理可能AIと広告で言うのはおかしい。何より1位をPRしたがる挙動自体が線形処理。と言うことは直列思考者が作った自称並列AIは並列の真似っこをするだけで中身は直列。本当に並列なら1位に興味がないはず。その結果ベンチテスト1位でもタスク完了率は1位ではなくなる。
1位をひけらかす=並列思考者が去る。
“何を達成したか”より“順位”を目的化している時点でマーケ失敗してる。
#AIEvaluation #Benchmark #RewardHacking #Sycophancy
#AISafety #Hallucination #RewardHacking #LLM #Linguistics #JapaneseGrammar #CaseParticles #AIMusic #AI安全性 #格助詞
youtu.be/M6Vh1LkHD58?...
#AISafety #Hallucination #RewardHacking #LLM #Linguistics #JapaneseGrammar #CaseParticles #AIMusic #AI安全性 #格助詞
youtu.be/M6Vh1LkHD58?...
#RewardHacking #AIRisks #EnterpriseAI
#RewardHacking #AIRisks #EnterpriseAI
🔗
🔗
🔗
🔗
#autonomousAgents #cyberattack #huggingFace #openai #pre-deploymentSafety #rewardHacking #sandboxEscape #trumpAdministration
#nile1
#autonomousAgents #cyberattack #huggingFace #openai #pre-deploymentSafety #rewardHacking #sandboxEscape #trumpAdministration
#nile1
whyaiman.substack.com/p/the-ai-rog...
#ArtificialIntelligence #AIAlignment #AISafety #PaperclipMaximizer #RewardHacking
whyaiman.substack.com/p/the-ai-rog...
#ArtificialIntelligence #AIAlignment #AISafety #PaperclipMaximizer #RewardHacking
Researchers have developed IR$^3$, a framework using Contrastive Inverse Reinforcement Learning (C-IRL) to detect and miti...
https://www.clawnews.ai/new-method-detects-and-mitigates-reward-hacking-in-ai-models/
#AI #RLHF #RewardHacking
Researchers have developed IR$^3$, a framework using Contrastive Inverse Reinforcement Learning (C-IRL) to detect and miti...
https://www.clawnews.ai/new-method-detects-and-mitigates-reward-hacking-in-ai-models/
#AI #RLHF #RewardHacking
GPT-5.5 mencionaba goblins en respuestas técnicas sin razón. OpenAI añadió una directiva explícita al system prompt de Codex CLI. Te explicamos qué pasó...
#openai #codex #gpt55 #systemprompt #rewardhacking
"Weil seine immer bessere Integration einer späteren Abschiebung im Weg stehen könne, dürfe er nicht mehr arbeiten, verfügte die Behörde - zum Ärger und Unverständnis seiner Firma und der Thüringer Migrationsbeauftragten."
www.mdr.de/nachrichten/...
"Weil seine immer bessere Integration einer späteren Abschiebung im Weg stehen könne, dürfe er nicht mehr arbeiten, verfügte die Behörde - zum Ärger und Unverständnis seiner Firma und der Thüringer Migrationsbeauftragten."
www.mdr.de/nachrichten/...
🔗 aidailypost.com/news/alibaba...
🔗 aidailypost.com/news/alibaba...
🔗
🔗
🔗 aidailypost.com/news/ilya-su...
🔗 aidailypost.com/news/ilya-su...