Interested in extracting world understanding from models and more controlled generation. 🌐 https://stefan-baumann.eu/
We built a model that does exactly this. It predicts motion, not pixels -- and it's 3,000× faster than video world models.
Myriad, accepted at
@cvprconference.bsky.social
Ulrich Prestel, @stefanabaumann.bsky.social, @rmsnorm.bsky.social, Björn Ommer
tl;dr:nuisance variable, architectural unification, and autoregressive pose learning->explicit dynamic state handling
arxiv.org/abs/2605.31535
Ulrich Prestel, @stefanabaumann.bsky.social, @rmsnorm.bsky.social, Björn Ommer
tl;dr:nuisance variable, architectural unification, and autoregressive pose learning->explicit dynamic state handling
arxiv.org/abs/2605.31535
→ Same number of steps. Same compute.
But images aren’t uniform. 🤔
Some regions are easy, others are hard.
So why force the model to treat them the same? 🧵
→ Same number of steps. Same compute.
But images aren’t uniform. 🤔
Some regions are easy, others are hard.
So why force the model to treat them the same? 🧵
But motion itself is much lower-dimensional.
We introduce 64× temporally compressed motion embeddings that directly capture scene dynamics.
This enables efficient planning -> 10,000× faster than video models.
🧵👇
But motion itself is much lower-dimensional.
We introduce 64× temporally compressed motion embeddings that directly capture scene dynamics.
This enables efficient planning -> 10,000× faster than video models.
🧵👇
Check out this concurrent great work from
@stefanabaumann.bsky.social, @jannik-w.bsky.social, @tommymarto.bsky.social, Mahdi M. Kalayeh, and Björn Ommer (@compvis.bsky.social)
We built a model that does exactly this. It predicts motion, not pixels -- and it's 3,000× faster than video world models.
Myriad, accepted at
@cvprconference.bsky.social
Check out this concurrent great work from
@stefanabaumann.bsky.social, @jannik-w.bsky.social, @tommymarto.bsky.social, Mahdi M. Kalayeh, and Björn Ommer (@compvis.bsky.social)
We built a model that does exactly this. It predicts motion, not pixels -- and it's 3,000× faster than video world models.
Myriad, accepted at
@cvprconference.bsky.social
We built a model that does exactly this. It predicts motion, not pixels -- and it's 3,000× faster than video world models.
Myriad, accepted at
@cvprconference.bsky.social
Here are the main improvements we made since RoMa:
Here are the main improvements we made since RoMa:
Come say hi in Honolulu!
👋 Pingchuan, Ming, Felix, Stefan, Timy, and Björn Ommer will be attending.
Come say hi in Honolulu!
👋 Pingchuan, Ming, Felix, Stefan, Timy, and Björn Ommer will be attending.
💡 It works if we leverage a self-supervised representation!
Meet RepTok🦎: A generative model that encodes an image into a single continuous latent while keeping realism and semantics. 🧵 👇
💡 It works if we leverage a self-supervised representation!
Meet RepTok🦎: A generative model that encodes an image into a single continuous latent while keeping realism and semantics. 🧵 👇
We built the Flow Poke Transformer (FPT) to model multi-modal scene dynamics from sparse interactions.
It learns to predict the 𝘥𝘪𝘴𝘵𝘳𝘪𝘣𝘶𝘵𝘪𝘰𝘯 of motion itself 🧵👇
We built the Flow Poke Transformer (FPT) to model multi-modal scene dynamics from sparse interactions.
It learns to predict the 𝘥𝘪𝘴𝘵𝘳𝘪𝘣𝘶𝘵𝘪𝘰𝘯 of motion itself 🧵👇
In short: we emphasize how autoencoders are implemented—but not always what they represent (and some of the implications of that representation).🧵
In short: we emphasize how autoencoders are implemented—but not always what they represent (and some of the implications of that representation).🧵
It makes so much more sense intuitively to work on a sequence rather than on a token level when our rewards are on a sequence level.
It makes so much more sense intuitively to work on a sequence rather than on a token level when our rewards are on a sequence level.
Come say hi in Nashville!
👋 Johannes, Ming, Kolja, Stefan, and Björn will be attending.
Come say hi in Nashville!
👋 Johannes, Ming, Kolja, Stefan, and Björn will be attending.
📌 Poster Session 6, Sunday 4:00 to 6:00 PM, Poster #208
📌 Poster Session 6, Sunday 4:00 to 6:00 PM, Poster #208
The other two interviewees' research played a pivotal role in the rise of diffusion models, whereas I just like to yap about them 😬 this was a wonderful opportunity to do exactly that!
The other two interviewees' research played a pivotal role in the rise of diffusion models, whereas I just like to yap about them 😬 this was a wonderful opportunity to do exactly that!
"Entropic Time Schedulers for Generative Diffusion Models"
We find that the conditional entropy offers a natural data-dependent notion of time during generation
Link: arxiv.org/abs/2504.13612
"Entropic Time Schedulers for Generative Diffusion Models"
We find that the conditional entropy offers a natural data-dependent notion of time during generation
Link: arxiv.org/abs/2504.13612
sander.ai/2025/04/15/l...
sander.ai/2025/04/15/l...
Project Page: vgg-t.github.io
Code & Weights: github.com/facebookrese...
Project Page: vgg-t.github.io
Code & Weights: github.com/facebookrese...
As this will get pretty long, this will be two threads.
The first will go into the RL part, and the second on the emergence and distillation.
As this will get pretty long, this will be two threads.
The first will go into the RL part, and the second on the emergence and distillation.
🤨Interested? Check out our latest work at #AAAI25:
💻Code and 📝Paper at: github.com/CompVis/DisCLIP
🧵👇
🤨Interested? Check out our latest work at #AAAI25:
💻Code and 📝Paper at: github.com/CompVis/DisCLIP
🧵👇
Recently with NICO: Apart from a class, the images have a context, one of them is 'autumn'. There is also a pumpkin class. Surprise surprise, many autumn images contain pumpkins.
Recently with NICO: Apart from a class, the images have a context, one of them is 'autumn'. There is also a pumpkin class. Surprise surprise, many autumn images contain pumpkins.