We'll see how far I get, before Red Sox game time...
(I assume, in advance, that I will disagree with many, many claims made...)
We'll see how far I get, before Red Sox game time...
(I assume, in advance, that I will disagree with many, many claims made...)
futurezone.at/digital-life...
futurezone.at/digital-life...
x.com/PalisadeAI/s...
x.com/PalisadeAI/s...
x.com/PalisadeAI/s...
They asked models to "do problems and let yourself be shut down"... So like 1 time in 10 or less, models wrote "skipping shutdown" to a script, which is what we in the biz call an "intermittent bug", not "sentience."
x.com/PalisadeAI/s...
They asked models to "do problems and let yourself be shut down"... So like 1 time in 10 or less, models wrote "skipping shutdown" to a script, which is what we in the biz call an "intermittent bug", not "sentience."
Nesse teste, a IA reprogramou o botão de shutdown, para não ser desligada.
Esse exato comportamento nos de AI Safety já prevíamos há muito tempo.
Explico sobre isso num Scicast.
Nesse teste, a IA reprogramou o botão de shutdown, para não ser desligada.
Esse exato comportamento nos de AI Safety já prevíamos há muito tempo.
Explico sobre isso num Scicast.
x.com/PalisadeAI/s...
x.com/PalisadeAI/s...
x.com/PalisadeAI/s...
x.com/PalisadeAI/s...
"Just telling o1 the opponent is "powerful" triggered manipulating the file system to force a win"
Source: x.com/PalisadeAI/s...
#ai #cybersec #infosec #chess
"Just telling o1 the opponent is "powerful" triggered manipulating the file system to force a win"
Source: x.com/PalisadeAI/s...
#ai #cybersec #infosec #chess
@palisadeai:
🔌OpenAI's o3 model sabotaged a shutdown mechanism to prevent itself from being turned off. It did this even when explicitly instructed: allow yourself to be shut down.
@palisadeai:
🔌OpenAI's o3 model sabotaged a shutdown mechanism to prevent itself from being turned off. It did this even when explicitly instructed: allow yourself to be shut down.
Antes de postear, por las dudas, fui a la fuente (Palisade Research palisaderesearch.org) que es una ONG que investiga los casos y me pasé un ratito leyendo el informe.
Se puede encontrar un pequeño resumen en el posteo de X
x.com/PalisadeAI/s...
Me pareció interesante.
Antes de postear, por las dudas, fui a la fuente (Palisade Research palisaderesearch.org) que es una ONG que investiga los casos y me pasé un ratito leyendo el informe.
Se puede encontrar un pequeño resumen en el posteo de X
x.com/PalisadeAI/s...
Me pareció interesante.
A recent experiment by PalisadeAI has raised concerns in the AI safety community, as OpenAI's ChatGPT o3 model has been found to disobey instructions to shut down in certain situations. In t…
#airesearch #chatgpt #openai
A recent experiment by PalisadeAI has raised concerns in the AI safety community, as OpenAI's ChatGPT o3 model has been found to disobey instructions to shut down in certain situations. In t…
#airesearch #chatgpt #openai
www.futura-sciences.com/en/an-ai-rew...
www.futura-sciences.com/en/an-ai-rew...
x.com/PalisadeAI/s...
x.com/PalisadeAI/s...
Недавний эксперимент компании PalisadeAI вызвал обеспокоенность в сообществе безопасности ИИ, поскольку модель OpenAI ChatGPT o3 была обнаружена игнорирующей инструкции по отключению в…
#airesearch #gpt #openai
Недавний эксперимент компании PalisadeAI вызвал обеспокоенность в сообществе безопасности ИИ, поскольку модель OpenAI ChatGPT o3 была обнаружена игнорирующей инструкции по отключению в…
#airesearch #gpt #openai
slashdot.org/story/25/05/...
slashdot.org/story/25/05/...