#VisionTransformer
MR-Transformer: A Vision Transformer-based Deep Learning Model for Total Knee Replacement Prediction Using MRI https://doi.org/10.1148/ryai.240373 @cem.bsky.social @cai2r.net #MSKRad #DJD #VisionTransformer
September 6, 2025 at 7:15 PM
MR-Transformer: A Vision Transformer-based Deep Learning Model for Total Knee Replacement Prediction Using MRI https://doi.org/10.1148/ryai.240373 @cem.bsky.social @cai2r.net #MSKRad #MachineLearning #VisionTransformer
October 22, 2025 at 7:15 PM
MR-Transformer predicts knee osteoarthritis progression to joint replacement using MRI https://doi.org/10.1148/ryai.240373 @cem.bsky.social @cai2r.net #DJD #VisionTransformer #ML
September 3, 2025 at 7:15 PM
LoRA fine‑tunes a Vision Transformer on synthetic aperture sonar, letting AI spot underwater mines even among rocks. The new DINOv3 + hard‑negative mining tricks boost detection like never before. Dive into the tech! #LoRA #VisionTransformer #MineDetection

🔗 aidailypost.com/news/lora-bo...
September 21, 2026 at 1:50 PM
Towards Global Glacier Mapping with Deep Learning and Open Earth Observation Data. (arXiv:2401.15113v1 [cs.CV])
Towards Global Glacier Mapping with Deep Learning and Open Earth Observation Data
Accurate global glacier mapping is critical for understanding climate change impacts. It is challenged by glacier diversity, difficult-to-classify debris and big data processing. Here we propose Glacier-VisionTransformer-U-Net (GlaViTU), a convolutional-transformer deep learning model, and five strategies for multitemporal global-scale glacier mapping using open satellite imagery. Assessing the spatial, temporal and cross-sensor generalisation shows that our best strategy achieves intersection over union >0.85 on previously unobserved images in most cases, which drops to >0.75 for debris-rich areas such as High-Mountain Asia and increases to >0.90 for regions dominated by clean ice. Additionally, adding synthetic aperture radar data, namely, backscatter and interferometric coherence, increases the accuracy in all regions where available. The calibrated confidence for glacier extents is reported making the predictions more reliable and interpretable. We also release a benchmark dataset that covers 9% of glaciers worldwide. Our results support efforts towards automated multitemporal and global glacier mapping.
arxiv.org
January 30, 2024 at 9:07 PM
#3 GlaViTU

This paper presents a novel world-wide dataset and a novel convolutional-transformer, named Glacier-VisionTransformer-U-Net (GlaViTU), for multitemporal and global glacier mapping.

⬆️ relevant task and nice results
⬇️ weak zero-shot transferability?

www.nature.com/articles/s41...
Globally scalable glacier mapping by deep learning matches expert delineation accuracy - Nature Communications
A deep learning model using open satellite data for scalable, global glacier mapping is developed. This model matches expert-level accuracy, facilitating more reliable glacier monitoring to support cl...
www.nature.com
February 5, 2025 at 1:24 PM
Ever wondered how many black-box queries it takes to peek inside a ViT's hidden layers? Researchers cracked the CIFAR-10 vision transformer with just 8,193 queries, revealing feed-forward dynamics and Hessian curvature. Dive in! #VisionTransformer #BlackBoxQueries #CIFAR10

🔗
September 1, 2026 at 5:58 AM
画像セットからの視覚概念推論:非言語的AI学習の新境地「Show Me Examples」

画像セットから視覚概念を推論する新手法「VICIS」を発表。

#コンピュータビジョン #生成AI #概念学習 #Few-shot学習 #VisionTransformer
画像セットからの視覚概念推論:非言語的AI学習の新境地「Show Me Examples」
画像セットから視覚概念を推論する新手法「VICIS」を発表。
ai.warp-studio.com
July 17, 2026 at 10:40 PM
Matthis Dallain, Laurent Rodriguez, Laurent Udo Perrinet, Beno\^it Miramond: A saccade-inspired approach to image classification using visiontransformer attention maps https://arxiv.org/abs/2603.09613 https://arxiv.org/pdf/2603.09613 https://arxiv.org/html/2603.09613
March 11, 2026 at 6:32 AM
VPNeXt, a simplified Vision Transformer model, beats the prior state‑of‑the‑art mIoU on the VOC2012 benchmark, with its version 3 released on September 27 2025. Read more: https://getnews.me/vpnext-rethinks-dense-decoding-for-vision-transformers/ #vpnext #visiontransformer #semanticsegmentation
October 1, 2025 at 11:34 AM
PVTAdpNet combines U‑Net and a Pyramid Vision Transformer, achieving a Dice score of 0.8851 and mIoU 0.8167 for real‑time polyp segmentation on standard GPUs. https://getnews.me/pvtadpnet-boosts-real-time-polyp-segmentation-with-vision-transformers/ #polypsegmentation #visiontransformer
September 30, 2025 at 1:22 PM
Orthogonal Residual Updates add only the component orthogonal to the activation stream; a ViT‑B model gained +4.3% top‑1 accuracy on ImageNet‑1k benchmark tests. Read more: https://getnews.me/orthogonal-residual-updates-improve-stability-and-accuracy-in-deep-networks/ #neurips2025 #visiontransformer
September 27, 2025 at 1:56 AM
A new hyperspectral adapter linking vision transformers achieved segmentation accuracy on three autonomous‑driving benchmarks, training on small datasets. Read more: https://getnews.me/hyperspectral-adapter-improves-semantic-segmentation-via-vision-models/ #hyperspectral #visiontransformer
September 26, 2025 at 8:00 PM
EfficienT‑HDR cuts FLOPS by ~67% and boosts CPU inference speed over fivefold, with about 2.5× speedup on edge processors, delivering HDR quality without ghosting. Read more: https://getnews.me/efficient-hdr-lightweight-transformer-improves-edge-hdr-imaging/ #efficienthdr #visiontransformer #edgeai
September 26, 2025 at 5:13 PM
A new progressive adaptation for Swin Transformers improves accuracy on PASCAL and NYUD‑v2 while using only about 20% of the trainable parameters of a fully fine‑tuned model. https://getnews.me/parameter-efficient-multi-task-learning-reduces-model-size-by-fivefold/ #multitask #visiontransformer
September 26, 2025 at 3:10 PM
ViTP embeds a Vision Transformer in a Vision‑Language Model, was tested on 16 remote‑sensing & medical benchmarks and achieved scores. Code on GitHub. Read more: https://getnews.me/visual-instruction-pretraining-boosts-domain-specific-vision-models/ #visualinstructionpretraining #visiontransformer
September 24, 2025 at 11:48 PM
Researchers reproduced the Vision Transformer study and found Diffusion Denoised Smoothing boosts explanation‑map robustness, though it adds extra overhead. Read more: https://getnews.me/reproducing-vision-transformers-with-diffusion-denoised-smoothing/ #visiontransformer #diffusiondenoisedsmoothing
September 19, 2025 at 11:30 PM
Activation‑space tuning (NoRA) updates only 0.4% of a vision transformer’s parameters (~0.02 M) and yields +0.17% accuracy on CIFAR‑10 and +0.27% on CIFAR‑100. Read more: https://getnews.me/activation-space-tuning-improves-parameter-efficient-fine-tuning/ #activationtuning #peft #visiontransformer
September 18, 2025 at 5:23 PM
I‑Segmenter is the first fully integer‑only Vision Transformer for semantic segmentation, achieving only a 5.1 % accuracy gap to FP32 while shrinking model size up to 3.8×. Read more: https://getnews.me/i-segmenter-vision-transformer-for-efficient-segmentation/ #visiontransformer #segmentation
September 17, 2025 at 6:00 AM
A framework merging an autoencoder with a Vision Transformer raised dental age‑estimation accuracy to 0.815 for second molars and 0.543 for third molars. https://getnews.me/autoencoder-vision-transformer-boosts-dental-age-estimation-accuracy/ #dentalage #visiontransformer
September 17, 2025 at 12:54 AM