Previously undergrad @Tsinghua CS and research intern @Berkeley AI Research @Berkeley InfoSchool.
https://yuxuanli.com/
Multi-agent is everywhere today. But put frontier LLMs in a room where each holds a different piece of the puzzle, and they fail 70% of the time.
Here's why:
📄 Paper: arxiv.org/abs/2505.11556
Multi-agent is everywhere today. But put frontier LLMs in a room where each holds a different piece of the puzzle, and they fail 70% of the time.
Here's why:
📄 Paper: arxiv.org/abs/2505.11556
📄 Paper: arxiv.org/abs/2505.11556
📊 Benchmark: huggingface.co/datasets/Yux...
#ICML
📄 Paper: arxiv.org/abs/2505.11556
📊 Benchmark: huggingface.co/datasets/Yux...
#ICML
1/n
1/n
This year, we’ve seen incredible interest from researchers across HCI, NLP, CSS, and Policy. We accepted 25 outstanding papers, with 5 selected as Best Paper nominees.
This year, we’ve seen incredible interest from researchers across HCI, NLP, CSS, and Policy. We accepted 25 outstanding papers, with 5 selected as Best Paper nominees.
🚨New preprint 🚨: "How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?"
👇 THREAD 👇
[Link to paper: arxiv.org/abs/2602.184...
[1/n]
🚨New preprint 🚨: "How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?"
👇 THREAD 👇
[Link to paper: arxiv.org/abs/2602.184...
[1/n]
We spent a year working with emergency preparedness policymakers to answer a simple question: can LLM agent simulations actually help real institutions make better decisions?
The answer is yes—but perhaps not how you'd expect.
👇 THREAD 👇
[Link to paper: arxiv.org/abs/2509.218...
[1/n]
We spent a year working with emergency preparedness policymakers to answer a simple question: can LLM agent simulations actually help real institutions make better decisions?
The answer is yes—but perhaps not how you'd expect.
👇 THREAD 👇
[Link to paper: arxiv.org/abs/2509.218...
[1/n]
We spent a year working with emergency preparedness policymakers to answer a simple question: can LLM agent simulations actually help real institutions make better decisions?
The answer is yes—but perhaps not how you'd expect.
👇 THREAD 👇
[Link to paper: arxiv.org/abs/2509.218...
[1/n]
We spent a year working with emergency preparedness policymakers to answer a simple question: can LLM agent simulations actually help real institutions make better decisions?
The answer is yes—but perhaps not how you'd expect.
👇 THREAD 👇
[Link to paper: arxiv.org/abs/2509.218...
[1/n]
Due to popular demand, the submission deadline for the PoliSim workshop at CHI’26 has been extended to February 20th! 🎉
We look forward to your submissions!
👉 polisim.net
Due to popular demand, the submission deadline for the PoliSim workshop at CHI’26 has been extended to February 20th! 🎉
We look forward to your submissions!
👉 polisim.net
Developing a new AI product? How would you figure out what are the privacy risks?
Privy help non-privacy expert practitioners create high quality privacy impact assessments for early-stage AI products.
Led by @hankhplee.bsky.social
Paper: www.sauvik.me/papers/69/s...
Developing a new AI product? How would you figure out what are the privacy risks?
Privy help non-privacy expert practitioners create high quality privacy impact assessments for early-stage AI products.
Led by @hankhplee.bsky.social
Paper: www.sauvik.me/papers/69/s...
PoliSim@CHI 2026: LLM Agent Simulation for Policy
PoliSim@CHI 2026: LLM Agent Simulation for Policy
PoliSim@CHI 2026: LLM Agent Simulation for Policy
👋I am a Postdoc Fellow at @hcii.cmu.edu working at the intersection of human-AI interaction, cognitive science, responsible AI, design, and social computing. I earned my PhD from @gtresearch.bsky.social in 2024.
👋I am a Postdoc Fellow at @hcii.cmu.edu working at the intersection of human-AI interaction, cognitive science, responsible AI, design, and social computing. I earned my PhD from @gtresearch.bsky.social in 2024.
I’ll be presenting in person in Suzhou. See you there!
Smarter LLMs are more selfish.
We show reasoning-enhanced models significantly prefer greed over cooperation. The more LLMs reason, the worse they cooperate.
👇 THREAD 👇
[Link to paper: arxiv.org/abs/2502.177...
[1/n]
I’ll be presenting in person in Suzhou. See you there!
Smarter LLMs are more selfish.
We show reasoning-enhanced models significantly prefer greed over cooperation. The more LLMs reason, the worse they cooperate.
👇 THREAD 👇
[Link to paper: arxiv.org/abs/2502.177...
[1/n]
Smarter LLMs are more selfish.
We show reasoning-enhanced models significantly prefer greed over cooperation. The more LLMs reason, the worse they cooperate.
👇 THREAD 👇
[Link to paper: arxiv.org/abs/2502.177...
[1/n]
Smarter LLMs are more selfish.
We show reasoning-enhanced models significantly prefer greed over cooperation. The more LLMs reason, the worse they cooperate.
👇 THREAD 👇
[Link to paper: arxiv.org/abs/2502.177...
[1/n]
“Groups that communicate diverse information make wiser decisions.” Right? Not always—for LLMs, just like humans.
We bring this scrutiny to AI by introducing the Hidden Profile paradigm to assess how multi-agent LLMs actually reason together.
arxiv.org/abs/2505.115...
[1/n]
“Groups that communicate diverse information make wiser decisions.” Right? Not always—for LLMs, just like humans.
We bring this scrutiny to AI by introducing the Hidden Profile paradigm to assess how multi-agent LLMs actually reason together.
arxiv.org/abs/2505.115...
[1/n]
New #Facct2025 paper led by @yuxuanli1225.bsky.social
Actions Speak Louder than Words: Agent Decisions Reveal Implicit Biases in Language Models
sauvikdas.com/papers/64/se...
New #Facct2025 paper led by @yuxuanli1225.bsky.social
Actions Speak Louder than Words: Agent Decisions Reveal Implicit Biases in Language Models
sauvikdas.com/papers/64/se...
A true testament to his hard work and determination!
It's also his FIRST Ph.D. paper!
Please boost :)
New preprint: “Actions Speak Louder Than Words: Agent Decisions Reveal Implicit Biases in Language Models”
👉 arxiv.org/abs/2501.17420
[1/n]
A true testament to his hard work and determination!
www.microsoft.com/en-us/resear...
www.microsoft.com/en-us/resear...
"Modeling End-User Affective Discomfort With Mobile App Permissions"
Paper: sauvikdas.com/papers/60/se...
"Modeling End-User Affective Discomfort With Mobile App Permissions"
Paper: sauvikdas.com/papers/60/se...
On-demand RFID: Improving Privacy, Security, and User Trust in RFID Activation through Physically-Intuitive Design
👉 Paper: sauvikdas.com/papers/62/se...
#Privacy #AcademicSky #Security
On-demand RFID: Improving Privacy, Security, and User Trust in RFID Activation through Physically-Intuitive Design
👉 Paper: sauvikdas.com/papers/62/se...
#Privacy #AcademicSky #Security
arxiv.org/abs/2409.12000
As AI technologists and data workers increasingly enter the news industry, cross-functional collaboration with journalists is becoming essential. But how do these collaborations unfold in practice?
arxiv.org/abs/2409.12000
As AI technologists and data workers increasingly enter the news industry, cross-functional collaboration with journalists is becoming essential. But how do these collaborations unfold in practice?
Following a year where he was recognized with both a CHI Best Paper and a USENIX Security Distinguished Paper, it looks like I will one day be best known for being Hank's advisor :p
advait.org/files/lee_20...
Following a year where he was recognized with both a CHI Best Paper and a USENIX Security Distinguished Paper, it looks like I will one day be best known for being Hank's advisor :p
Purpose Mode: Reducing Distraction through Toggling ACDPs on Social Media Web Sites
📝 hankhplee.com/papers/purpo...
Purpose Mode: Reducing Distraction through Toggling ACDPs on Social Media Web Sites
📝 hankhplee.com/papers/purpo...