Alzbeta Manova
banner
alzbetam.bsky.social
Alzbeta Manova
@alzbetam.bsky.social
🎓 PhD student @ University of Aberdeen | 📈 Math + 🧠 Psych + 🤖 ML.
🐕 Dog parent | 🧩 Puzzle solver
🏴󠁧󠁢󠁳󠁣󠁴󠁿🇨🇿🏳️‍🌈
Presented at the BMVA Symposium today!

Quick visit to London ✈️
Great feedback and ideas💡
And great food 🥐🍓
September 30, 2026 at 8:04 PM
Reposted by Alzbeta Manova
Shout out to the #rstats {performance} package. Its check_model() function is super helpful for model evaluation.

#statistics #STEM
June 23, 2025 at 2:52 PM
Reposted by Alzbeta Manova
Petition signable by UK citizens or residents, asking "the Government to introduce a statutory duty requiring AI developers to disclose the copyrighted works used to train generative AI in enough detail..."

💙📚 🗃 #academicsky
1/

petition.parliament.uk/petitions/76...
Petition: Protect creators' rights: require AI firms to disclose training data
We ask the Government to introduce a statutory duty requiring AI developers to disclose the copyrighted works used to train generative AI in enough detail for creators to identify use of their work, e...
petition.parliament.uk
August 6, 2026 at 6:10 PM
July in PhD:

🎤 Presented at a conference
💻 Collected data for a new experiment
🤖 Evaluated multiple VLMs
🎯 Got accepted to present at a symposium
🎂 Celebrated my birthday!
📊 Analyzed new data
🗓 planned the next 3 months of my PhD

#PhDSky #PhDLife
August 5, 2026 at 9:47 AM
Reposted by Alzbeta Manova
This headline is extremely funny given what happened (OpenAI hacked HF by accident)

openai.com/index/huggin...
OpenAI and Hugging Face partner to address security incident during model evaluation
OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities and lessons for defenders.
openai.com
July 21, 2026 at 8:16 PM
Nothing beats that feeling when your script and local server finally find each other and communicate correctly ✨️

Brings me back to my undergrad days full of debugging and codebase building 💻

#PhDLife #PhDSky #AIEvaluation #AI
July 29, 2026 at 3:25 PM
Reposted by Alzbeta Manova
New preprint from the lab asking what minimal parts of whole body posture/motion allow for rapid movement recognition

arxiv.org/abs/2607.13216
Classifying daily activities needs posture, reconstructing them needs motion
Humans recognize movements effortlessly, even from noisy and complex visual input. But what information in the stimulus allows humans to rapidly classify movements? No framework has systematically com...
arxiv.org
July 23, 2026 at 11:52 PM
Reposted by Alzbeta Manova
This is quite the finding for a device thats just a plagiarism machine, or whatever we are calling it this week.

www.nature.com/articles/s41...
Large language models can predict the results of social science experiments - Nature
Large language models can be used to estimate the results of social science experiments about as accurately as a group of human forecasters—even for experiments published after the...
www.nature.com
July 9, 2026 at 1:25 AM
Day 2 of the Cyberpsychology Section Conference 🎓

Thank you to everyone who attended my talk on VLMs in action perception, loved all the feedback and discussions✨️

#cyber26 #conference #PhDSky
July 7, 2026 at 5:51 PM
Day 1 of the Cyberpsychology Section Annual Conference! 🧠🤖

Learned so much already and we are only half way through 🎓
#cyber26 #conference #PhDLife
July 6, 2026 at 11:25 PM
Reposted by Alzbeta Manova
All AI evaluations embed assumptions about what "good" behavior looks like. Our #FAccT2026 paper explores how we can center the perspectives of impacted communities - specifically, subjects of AI-generated media - in designing LLM-as-a-judge evaluation rubrics.
June 24, 2026 at 6:31 PM
Reposted by Alzbeta Manova
WorldOdysseyBench benchmarks interactive models, focusing on long-horizon stability in action following, visual quality, memory, and physics. This evaluation highlights gaps in existing models, improving the applicability of AI-generated environments. https://arxiv.org/abs/2606.31672
WorldOdysseyBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models
ArXiv link for WorldOdysseyBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models
arxiv.org
July 4, 2026 at 9:20 AM
4 years of growth, a 40-something cm chop, and lighter heads for both of us! ✂️

Officially a two-time hair donor 🙌 Highly recommend this haircut for the summer weather.

#HairDonation #DoGood
June 30, 2026 at 4:27 PM
Reposted by Alzbeta Manova
The brain continuously optimizing its predictions. Local "leaky" circuits correct errors, while hierarchies link these computations using nonlinearities shaped by prior expectations.

Predictive Coding with Bayesian Priors via Proximal Gradients
arxiv.org/abs/2606.08374
#neuroscience
Predictive Coding with Bayesian Priors via Proximal Gradients
We recast predictive coding as continuous-time proximal gradient descent applied to a regularized maximum-a-posteriori (MAP) objective. We study first a single-level problem and then a multi-level hie...
arxiv.org
June 23, 2026 at 7:00 PM
May in PhD:

✅Successful conference application
⚙️Organised school postgrad conference
📈Piloted 1 experiment
🎞️Processed 134 videos
📚Attended a teaching certificate workshop

Now off to a 3-week road trip! 🚗
#PhDSky #PhDLife
May 31, 2026 at 6:04 PM
Reposted by Alzbeta Manova
AI can't handle the truth:

"AI is still getting things wrong, more confidently than ever"

www.axios.com/2026/05/30/a...
AI can't handle the truth
Confident chatbots could encourage users to stop fact-checking.
www.axios.com
May 30, 2026 at 1:11 PM
Helped organise & presented at our annual school postgrad conference last week!
Sitting at the tech desk on stage definitely helped with the stage fright when sharing my research in action perception!

Huge thanks to organisers, attendees & presenters ✨
#Conference #Psychology #PhDLife #PhDSky
May 29, 2026 at 3:56 PM
Reposted by Alzbeta Manova
ProcBench shifts focus from final outcomes to execution-process quality in LLM coding agents, revealing defects and enhancing reliability and governability. This method uncovers flaws often missed by traditional evaluations. https://arxiv.org/abs/2605.20251
ProcBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents
ArXiv link for ProcBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents
arxiv.org
May 23, 2026 at 8:40 AM
Always love a good MLLM evaluation paper. Over half (51%) of "correct" model deductions about video subjects are ungrounded shortcuts rather than real cue processing!
#AI #MLLM
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
Multimodal Large Language Models (MLLMs) are increasingly deployed in human-facing roles where personality perception is critical, yet existing benchmarks evaluate this capability solely on numerical ...
arxiv.org
May 22, 2026 at 12:07 PM
Reposted by Alzbeta Manova
AI can reveal your location from a single photo. Vision‑Language Models can determine where any photo is taken, without any GPS data

What does that means for your privacy?
Nowhere to Hide? Privacy Risks and Policy Implications of AI Geolocation
One of the most surprising — and concerning — capabilities of the newest Artificial Intelligence (AI) systems is their ability to infer geographic location from images. Vision‑Language Models (VLMs)...
privacyinternational.org
May 16, 2026 at 5:15 PM
Reposted by Alzbeta Manova
ArXiV has a new LLM policy

(Screenshots with alt text so you don’t have to click through to the other place and see all the stupid responses)
May 14, 2026 at 9:50 PM
5 days until our school postgrad conference! 🎓

Throwback to last year when I was just a guest and didn't realize that organizing the thing is slightly more stressful than just presenting 😅

#ConferencePrep #PhDLife #PhDSky
May 13, 2026 at 9:09 AM
Reposted by Alzbeta Manova
A 1 a.m. thought for the CogSci undergrads:
Everyone talks about the "cool" theories, but no one tells you that Mathematics is what actually builds your future. Go for the math, even if it’s not required. Your future self will thank you for the analytical edge. 🧠📉 #ScienceSky#AcademicSky
May 8, 2026 at 8:02 PM
Nothing cures research procrastination like a conference acceptance notification 😅

I suddenly have all the motivation to complete my study 📊
#PhDLife #PhDSky #Cyber26
May 9, 2026 at 7:53 PM
Reposted by Alzbeta Manova
Truly excellent work from Anthropic. I think this is a promising interpretability technique, but what I really like is that they’ve made code available for open weight models.
Natural Language Autoencoders
Turning Claude's thoughts into text
www.anthropic.com
May 7, 2026 at 10:03 PM