If you want to generate this kind of images, it's actually fairly simple: train an autoregressive MLP on a cifar10 (with positional encoding).
If you want to generate this kind of images, it's actually fairly simple: train an autoregressive MLP on a cifar10 (with positional encoding).
Anyone else interested in this?
#julia #julialang
Anyone else interested in this?
#julia #julialang
In experiments across MLPs and ResNets on CIFAR10 and ViTs on ImageNet1K, we show that 𝝁P² indeed jointly transfers optimal learning rate and perturbation radius across model scales and can improve training stability and generalization.
🧵 8/10
In experiments across MLPs and ResNets on CIFAR10 and ViTs on ImageNet1K, we show that 𝝁P² indeed jointly transfers optimal learning rate and perturbation radius across model scales and can improve training stability and generalization.
🧵 8/10
🔥🚀 we are excited to keep pushing this line of work 💪
🔥🚀 we are excited to keep pushing this line of work 💪
Also CIFAR10/100 results are really terrible. I have got >90% in 2016, and yours are around 70%
Also CIFAR10/100 results are really terrible. I have got >90% in 2016, and yours are around 70%
Payman Behnam, Uday Kamal, Sanjana Vijay Ganesh et al.
Action editor: Naigang Wang
https://openreview.net/forum?id=ubrOSWyTS8
#imagenet #cifar10 #optimized
Payman Behnam, Uday Kamal, Sanjana Vijay Ganesh et al.
Action editor: Naigang Wang
https://openreview.net/forum?id=ubrOSWyTS8
#imagenet #cifar10 #optimized
🐙 github.com/dcasbol/biol...
🐙 github.com/dcasbol/biol...
Are you sure this makes any difference in the competitive setting? Seems like choosing hyper params makes more of a difference
arxiv.org/abs/2206.00364
Are you sure this makes any difference in the competitive setting? Seems like choosing hyper params makes more of a difference
arxiv.org/abs/2206.00364
🔗
🔗
Using a ViT-Tiny, we observe an average 38% improvement in linear probing performance compared to MAEs with the standard 75% masking ratio.
Using a ViT-Tiny, we observe an average 38% improvement in linear probing performance compared to MAEs with the standard 75% masking ratio.