https://fazlbarez.com
I was glad to join @dwdderwetterdienst.bsky.social news to discuss unexpected AI behaviour and what it means for safety.
We still have agency over AIs and we should only build AGI if we can align them!
I was glad to join @dwdderwetterdienst.bsky.social news to discuss unexpected AI behaviour and what it means for safety.
We still have agency over AIs and we should only build AGI if we can align them!
@eccv.bsky.social !
For the best part of history we relied on understanding to arrive at answers. Yet, AI's can produce answers we cannot understand fully.
In my talk i'll discuss what understanding may look like in the age of AGI!
@eccv.bsky.social !
For the best part of history we relied on understanding to arrive at answers. Yet, AI's can produce answers we cannot understand fully.
In my talk i'll discuss what understanding may look like in the age of AGI!
The technology is exciting, but as AI systems begin acting in the physical world, it becomes even more important to test them carefully and make sure people remain in control.
The technology is exciting, but as AI systems begin acting in the physical world, it becomes even more important to test them carefully and make sure people remain in control.
Also hiring 2 RAs (interp + continual learning) at Oxford: what happens inside models that keep learning after deployment?
Come chat in Seoul, or apply 👇
tsglab.github.io/vacancies/
Please share with folks!
Also hiring 2 RAs (interp + continual learning) at Oxford: what happens inside models that keep learning after deployment?
Come chat in Seoul, or apply 👇
tsglab.github.io/vacancies/
Please share with folks!
10 papers: 7 in the main conf, 3 WS: interpretability, AI evaluation, and governance, two oral spotlights
Giving an invited talk at the EIML WS
Grateful to the students, collaborators, mentors and Claude who made it happen!
10 papers: 7 in the main conf, 3 WS: interpretability, AI evaluation, and governance, two oral spotlights
Giving an invited talk at the EIML WS
Grateful to the students, collaborators, mentors and Claude who made it happen!
Massive thanks to all my collaborators—I’ve been lucky to work with such brilliant people
#ICML2026
Massive thanks to all my collaborators—I’ve been lucky to work with such brilliant people
#ICML2026
Motion: This House Believes that AI is the Great Equalizer
Is it? Or isn't it? I'm speaking for the proposition--which might surprise those who know my work. That's rather the point!
We'll find out which way the House votes
Motion: This House Believes that AI is the Great Equalizer
Is it? Or isn't it? I'm speaking for the proposition--which might surprise those who know my work. That's rather the point!
We'll find out which way the House votes
A recent article co-authored by @aigioxfordmartin.bsky.social researcher Fazl Barez, explores safety challenges and the importance of being context-aware
🔗Read the research in full: www.science.org/doi/10.1126/...
A recent article co-authored by @aigioxfordmartin.bsky.social researcher Fazl Barez, explores safety challenges and the importance of being context-aware
🔗Read the research in full: www.science.org/doi/10.1126/...
The goal is to turns hard interpretability questions into tools for human empowerment, oversight and governance.
The goal is to turns hard interpretability questions into tools for human empowerment, oversight and governance.
@aims_oxford
!
I’m thrilled to launch a new called AI Safety & Alignment (AISAA) course on the foundations & frontier research of making advanced AI systems safe and aligned at
@UniofOxford
what to expect 👇
robots.ox.ac.uk/~fazl/aisaa/
@aims_oxford
!
I’m thrilled to launch a new called AI Safety & Alignment (AISAA) course on the foundations & frontier research of making advanced AI systems safe and aligned at
@UniofOxford
what to expect 👇
robots.ox.ac.uk/~fazl/aisaa/
🧵
My latest paper tries to solve a longstanding problem afflicting fields such as decision theory, economics, and ethics — the problem of infinities.
Let me explain a bit about what causes the problem and how my solution avoids it.
1/N
arxiv.org/abs/2509.19389
🧵
My latest paper tries to solve a longstanding problem afflicting fields such as decision theory, economics, and ethics — the problem of infinities.
Let me explain a bit about what causes the problem and how my solution avoids it.
1/N
arxiv.org/abs/2509.19389
More details (and more bragging) soon! and maybe even more news on sep 25 👀
See you all in… Mexico? San Diego? Copenhagen? Who knows! 🌍✈️
More details (and more bragging) soon! and maybe even more news on sep 25 👀
See you all in… Mexico? San Diego? Copenhagen? Who knows! 🌍✈️
arxiv.org/pdf/2509.00117
Co-authors Jared Perlo @fbarez.bsky.social Alex Robey & @floridi.bsky.social
1\4
arxiv.org/pdf/2509.00117
Co-authors Jared Perlo @fbarez.bsky.social Alex Robey & @floridi.bsky.social
1\4
Here, FUR offers a fine-grained test if LMs latently used information from CoTs for answers!
Here, FUR offers a fine-grained test if LMs latently used information from CoTs for answers!
Today’s AI doesn’t just assist decisions; it makes them. Governments use it for surveillance, prediction, and control — often with no oversight.
Technical safeguards aren’t enough on their own — but they’re essential for AI to serve society.
Today’s AI doesn’t just assist decisions; it makes them. Governments use it for surveillance, prediction, and control — often with no oversight.
Technical safeguards aren’t enough on their own — but they’re essential for AI to serve society.
We'll draw from political theory, cooperative AI, economics, mechanism design, history, and hierarchical agency.
We'll draw from political theory, cooperative AI, economics, mechanism design, history, and hierarchical agency.
Joint work with Clara Suslik, Yihuai Hong, and @fbarez.bsky.social, advised by @megamor2.bsky.social
🔗 Paper: arxiv.org/abs/2505.22586
🔗 Code: github.com/yoavgur/PISCES
Joint work with Clara Suslik, Yihuai Hong, and @fbarez.bsky.social, advised by @megamor2.bsky.social
🔗 Paper: arxiv.org/abs/2505.22586
🔗 Code: github.com/yoavgur/PISCES
White-box LLMs & model security
Safe RL & reward hacking
Interpretability & governance tools
Remote or Oxford.
Apply by 30 May 23:59 UTC. DM with questions.
White-box LLMs & model security
Safe RL & reward hacking
Interpretability & governance tools
Remote or Oxford.
Apply by 30 May 23:59 UTC. DM with questions.
We’re hiring a Postdoc in Causal Systems Modelling to:
- Build causal & white-box models that make frontier AI safer and more transparent
- Turn technical insights into safety cases, policy briefs, and governance tools
]
DM if you have any questions.
We’re hiring a Postdoc in Causal Systems Modelling to:
- Build causal & white-box models that make frontier AI safer and more transparent
- Turn technical insights into safety cases, policy briefs, and governance tools
]
DM if you have any questions.
After suffering through unhelpful reviews as an author, I want to do right by papers in my track.
After suffering through unhelpful reviews as an author, I want to do right by papers in my track.