bsky.app/profile/theo...
bsky.app/profile/theo...
by @tkrusch.bsky.social @mmbronstein.bsky.social
They propose a regularization approach for exploiting symmetries over data (penalizing variable predictions over augmented data).
arxiv.org/abs/2410.17878
by @tkrusch.bsky.social @mmbronstein.bsky.social
They propose a regularization approach for exploiting symmetries over data (penalizing variable predictions over augmented data).
arxiv.org/abs/2410.17878
A JEPA architecture that learns sparse, non-negative, informative representations through principled distributional regularization.
A JEPA architecture that learns sparse, non-negative, informative representations through principled distributional regularization.
But more importantly
Thou shalt not overestimate the
ability of regularization approaches to prevent overfitting issues
But more importantly
Thou shalt not overestimate the
ability of regularization approaches to prevent overfitting issues
Both types of structural regularization reduce degeneracy across all levels. Regularization nudges networks toward more consistent, shared solutions.
Both types of structural regularization reduce degeneracy across all levels. Regularization nudges networks toward more consistent, shared solutions.
direct.mit.edu/jocn/article...
(still uncorrected proofs, but they should post the corrected one soon--also OA is forthcoming, for now PDF at brainandexperience.org/pdf/10.1162-...)
direct.mit.edu/jocn/article...
(still uncorrected proofs, but they should post the corrected one soon--also OA is forthcoming, for now PDF at brainandexperience.org/pdf/10.1162-...)
Why do we see localized receptive fields so often, even in models without sparisity regularization?
We present a theory in the minimal setting from @aingrosso.bsky.social and @sebgoldt.bsky.social
Why do we see localized receptive fields so often, even in models without sparisity regularization?
We present a theory in the minimal setting from @aingrosso.bsky.social and @sebgoldt.bsky.social