Fazl Barez
fbarez.bsky.social
Fazl Barez
@fbarez.bsky.social
Let's build AI's we can trust!
https://fazlbarez.com
As AI gets more autonomy, we need stronger evidence that we can keep it under control.

I was glad to join @dwdderwetterdienst.bsky.social news to discuss unexpected AI behaviour and what it means for safety.

We still have agency over AIs and we should only build AGI if we can align them!
September 18, 2026 at 5:15 PM
I’m excited to give an invited talk at eXCV workshop
@eccv.bsky.social !

For the best part of history we relied on understanding to arrive at answers. Yet, AI's can produce answers we cannot understand fully.

In my talk i'll discuss what understanding may look like in the age of AGI!
September 8, 2026 at 9:25 AM
I was glad to join Sky News to discuss some of the safety questions around humanoid robots.

The technology is exciting, but as AI systems begin acting in the physical world, it becomes even more important to test them carefully and make sure people remain in control.
September 3, 2026 at 8:36 AM
I'll be at #ICML2026 next week—7 main conference papers, 2 orals + an invited talk 🇰🇷

Also hiring 2 RAs (interp + continual learning) at Oxford: what happens inside models that keep learning after deployment?
Come chat in Seoul, or apply 👇
tsglab.github.io/vacancies/

Please share with folks!
Vacancies | TSG Lab – Technical Safety & Governance Lab
tsglab.github.io
July 4, 2026 at 2:03 PM
Reposted by Fazl Barez
A glimpse into another successful Oxford Connected Life Summit, focused on what it means to live in an increasingly connected world. This year's theme was "New Intelligence, Old Questions," and featured notable speakers and organisations. Huge thanks to the student committee for making this happen!
July 2, 2026 at 1:57 PM
Heading to #ICML2026 in Seoul next week with the TSG Lab and Martian 🇰🇷

10 papers: 7 in the main conf, 3 WS: interpretability, AI evaluation, and governance, two oral spotlights

Giving an invited talk at the EIML WS

Grateful to the students, collaborators, mentors and Claude who made it happen!
June 30, 2026 at 10:02 AM
Really grateful to have 7 papers accepted at @icmlconf.bsky.social onf 2026, including 2 spotlights!

Massive thanks to all my collaborators—I’ve been lucky to work with such brilliant people

#ICML2026
June 25, 2026 at 11:37 AM
Excited to be debating at the Oxford Union this evening

Motion: This House Believes that AI is the Great Equalizer

Is it? Or isn't it? I'm speaking for the proposition--which might surprise those who know my work. That's rather the point!

We'll find out which way the House votes
June 25, 2026 at 7:57 AM
Reposted by Fazl Barez
In film, "we'll fix it in post" is what you say when something went wrong on set and you don't want to redo it. AI research has made it our entire methodology: train the model, then patch whatever comes out. Our new ICML oral argues this can't be the basis of a science of AI. 🧵
June 10, 2026 at 7:12 PM
Reposted by Fazl Barez
How can we ensure AI-powered robots remain safe when operating in the real world? 🤖

A recent article co-authored by @aigioxfordmartin.bsky.social researcher Fazl Barez, explores safety challenges and the importance of being context-aware

🔗Read the research in full: www.science.org/doi/10.1126/...
Beyond alignment: Why robotic foundation models need context-aware safety
Because AI-enabled robots can be tricked into taking unsafe actions, they require layered, context-aware safety guardrails.
www.science.org
June 9, 2026 at 2:46 PM
Incredibly excited to announce $1 Million prize pool to solve the world’s most important scientific problem in Interpretability.

The goal is to turns hard interpretability questions into tools for human empowerment, oversight and governance.
December 7, 2025 at 6:36 PM
🚨New AI Safety Course
@aims_oxford
!

I’m thrilled to launch a new called AI Safety & Alignment (AISAA) course on the foundations & frontier research of making advanced AI systems safe and aligned at
@UniofOxford

what to expect 👇
robots.ox.ac.uk/~fazl/aisaa/
October 6, 2025 at 4:40 PM
Reposted by Fazl Barez
Evaluating the Infinite
🧵
My latest paper tries to solve a longstanding problem afflicting fields such as decision theory, economics, and ethics — the problem of infinities.
Let me explain a bit about what causes the problem and how my solution avoids it.
1/N
arxiv.org/abs/2509.19389
Evaluating the Infinite
I present a novel mathematical technique for dealing with the infinities arising from divergent sums and integrals. It assigns them fine-grained infinite values from the set of hyperreal numbers in a ...
arxiv.org
September 25, 2025 at 3:28 PM
🚀 Excited to have 2 papers accepted at #NeurIP2025! 🎉 congrats to my amazing co-authors!

More details (and more bragging) soon! and maybe even more news on sep 25 👀

See you all in… Mexico? San Diego? Copenhagen? Who knows! 🌍✈️
September 19, 2025 at 9:08 AM
Reposted by Fazl Barez
🚨 NEW PAPER 🚨: Embodied AI (incl. AI-powered drones, self-driving cars and robots) is here, but policies are lagging. We analyzed the EAI risks and found significant gaps in governance

arxiv.org/pdf/2509.00117

Co-authors Jared Perlo @fbarez.bsky.social Alex Robey & @floridi.bsky.social

1\4
September 4, 2025 at 5:51 PM
Reposted by Fazl Barez
Other works have highlighted that CoTs ≠ explainability alphaxiv.org/abs/2025.02 (@fbarez.bsky.social), and that intermediate (CoT) tokens ≠ reasoning traces arxiv.org/abs/2504.09762 (@rao2z.bsky.social).

Here, FUR offers a fine-grained test if LMs latently used information from CoTs for answers!
Chain-of-Thought Is Not Explainability | alphaXiv
View 3 comments: There should be a balance of both subjective and observable methodologies. Adhering to just one is a fools errand.
alphaxiv.org
August 21, 2025 at 3:21 PM
Reposted by Fazl Barez
It is so easy to confuse chain of thought and explainability and in fact in a lot of the media it is presented as if with current LLMs we are allowed to view their actual thought processes. It is not that!
Excited to share our paper: "Chain-of-Thought Is Not Explainability"! We unpack a critical misconception in AI: models explaining their steps (CoT) aren't necessarily revealing their true reasoning. Spoiler: the transparency can be an illusion. (1/9) 🧵
July 2, 2025 at 12:41 PM
Excited to share our paper: "Chain-of-Thought Is Not Explainability"! We unpack a critical misconception in AI: models explaining their steps (CoT) aren't necessarily revealing their true reasoning. Spoiler: the transparency can be an illusion. (1/9) 🧵
July 1, 2025 at 3:41 PM
Technology = power. AI is reshaping power — fast.

Today’s AI doesn’t just assist decisions; it makes them. Governments use it for surveillance, prediction, and control — often with no oversight.

Technical safeguards aren’t enough on their own — but they’re essential for AI to serve society.
June 27, 2025 at 8:07 AM
Reposted by Fazl Barez
And Anna Yelizarov, @fbarez.bsky.social, @scasper.bsky.social, Beatrice Erkers, among others.

We'll draw from political theory, cooperative AI, economics, mechanism design, history, and hierarchical agency.
June 18, 2025 at 6:12 PM
Reposted by Fazl Barez
This is a step toward targeted, interpretable, and robust knowledge removal — at the parameter level.

Joint work with Clara Suslik, Yihuai Hong, and @fbarez.bsky.social, advised by @megamor2.bsky.social
🔗 Paper: arxiv.org/abs/2505.22586
🔗 Code: github.com/yoavgur/PISCES
May 29, 2025 at 4:22 PM
Come work with me at Oxford this summer! Paid research opportunity to:

White-box LLMs & model security
Safe RL & reward hacking
Interpretability & governance tools

Remote or Oxford.

Apply by 30 May 23:59 UTC. DM with questions.
May 20, 2025 at 5:13 PM
Come work with me at Oxford!

We’re hiring a Postdoc in Causal Systems Modelling to:

- Build causal & white-box models that make frontier AI safer and more transparent
- Turn technical insights into safety cases, policy briefs, and governance tools
]

DM if you have any questions.
May 15, 2025 at 11:12 AM
First-time Area Chair seeking advice! What helped you most when evaluating papers beyond just averaging scores?

After suffering through unhelpful reviews as an author, I want to do right by papers in my track.
April 8, 2025 at 11:59 AM