#Ego4D
Join us at 2nd EgoVis workshop @cvprconference.bsky.social #CVPR2025 in #Nashville with 28 open challenges [DL 19 May] across datasets #Ego4D #Meta-Aria #HoloAssist and EPIC-KITCHENS, w great keynotes - Abstract DL 2 May
Lead organisers: Siddhant Bansal @antoninofurnari.bsky.social &Tushar Nagarajan
January 31, 2025 at 5:03 PM
EgoVLA - they train a VLA on egocentric human data and transfer it to bimanual manipulation.
By Yang et al., UC San Diego / UIUC / MIT / Nvidia

(No Ego4D nor EPIC Kitchen in the data mix?)

arxiv.org/abs/2507.12440
July 17, 2025 at 6:29 AM
Ego4D is an incredible anthropological dataset. With over 3600 hours of videos (5 months of human activities) in first person view with written narrations from all over the world ego4d-data.org/fig1.html. (Always reminds me of this movie by #IsabelCoixet youtu.be/bYI3A6WLMBU)
Spain in a day - Tráiler
YouTube video by GRUP MEDIAPRO
youtu.be
March 25, 2025 at 2:30 AM
But how do we know if our data is any good? We asked 35k humans in Prolific about how correct and useful they found the suggested actions in reference to every video. Now we opensource the new dataset alongside the human ranks. github.com/google/parse... and all the data generation prompts.
GitHub - google/parse-ego4d: Dataset and code for generating 18,000 action suggestions via LLM querries for EGO4D dataset validated by 36,171 human annotations.
Dataset and code for generating 18,000 action suggestions via LLM querries for EGO4D dataset validated by 36,171 human annotations. - google/parse-ego4d
github.com
March 25, 2025 at 2:30 AM
ego4d-data.org

epic-kitchens.github.io/2024

ego-exo4d-data.org

Ma conosco gente che ci sta lavorando almeno alla parte di raccolta 😅
July 1, 2024 at 9:09 PM
You can check out the work next month at @iclr-conf.bsky.social Bi-Align (bialign-workshop.github.io) and FM-wild (fm-wild-community.github.io) workshops:
“PARSE-Ego4D: Personal Action Recommendation Suggestions for Egocentric Videos”

Or start playing right now :)
March 25, 2025 at 4:24 AM
For the dataset, we leverage Ego4D videos and annotations, and augment them with an automated pipeline to generate data samples for our task.
December 6, 2024 at 6:33 PM
Thanks @kid-icarus.bsky.social ! I couldn't find the Ego3D, but I did browse the Ego4D dataset and it could be useful. Nonetheless, I was surprised to see very few public HRI datasets.
February 17, 2025 at 10:37 AM
Steven Abreu, Tiffany D. Do, Karan Ahuja, Eric J. Gonzalez, Lee Payne, Daniel McDuff, Mar Gonzalez-Franco
PARSE-Ego4D: Personal Action Recommendation Suggestions for Egocentric Videos
https://arxiv.org/abs/2407.09503
July 26, 2024 at 2:32 PM
Steven Abreu, Tiffany D. Do, Karan Ahuja, Eric J. Gonzalez, Lee Payne, Daniel McDuff, Mar Gonzalez-Franco
PARSE-Ego4D: Personal Action Recommendation Suggestions for Egocentric Videos
https://arxiv.org/abs/2407.09503
July 16, 2024 at 4:02 AM
But really as it happens with all datasets, the key is on “How to make data useful?” That was our goal with PARSE-Ego4D parse-ego4d.github.io we augment the existing narrations and videos by providing 18k suggested actions using the latest of VLMs. This can be useful to train a future AI agents.
March 25, 2025 at 2:30 AM
パナソニック、動画を見せると“専門家”として答えるAI 世界最高峰コンペで第2位に
ascii.jp/elem/000/004...
>質問に応じて適切な専門家AIを動的に生成し、監督役のAIが意見をまとめて回答を選択する仕組み。これにより、人間の正解率76%に迫る71%の正解率を達成した。

なにそれすごい
パナソニック、動画を見せると“専門家”として答えるAI 世界最高峰コンペで第2位に
パナソニックグループで法人向けソリューションを扱うパナソニック コネクトは7月17日、画像認識分野で世界最高峰とされる学会CVPR2024のコンペティション「Ego4D EgoSchema Challenge」で世界第2位の評価を獲得したと発表。
ascii.jp
July 19, 2024 at 8:05 AM
Chaoyang Wang, Lexuan Xu: FROST-STA: Frozen Dense Features for the Ego4D Short-Term Object Interaction Anticipation https://arxiv.org/abs/2606.00694 https://arxiv.org/pdf/2606.00694 https://arxiv.org/html/2606.00694
June 2, 2026 at 6:41 AM
Andrea Zenotto, Simone Alberto Peirone, Francesca Pistilli, Giuseppe Averta: HiERO-StepG @ Ego4D Step Grounding Challenge: hierarchical activity understanding enables zero-shot step grounding https://arxiv.org/abs/2605.31227 https://arxiv.org/pdf/2605.31227 https://arxiv.org/html/2605.31227
June 1, 2026 at 6:41 AM
Qiaohui Chu, Haoyu Zhang, Yisen Feng, Meng Liu, Weili Guan, Dongmei Jiang, Liqiang Nie: VISTA: Technical Report for the Ego4D Short-Term Object Interaction Anticipation at EgoVis 2026 https://arxiv.org/abs/2605.20901 https://arxiv.org/pdf/2605.20901 https://arxiv.org/html/2605.20901
May 21, 2026 at 6:41 AM
Yisen Feng, Leigang Qu, Haoyu Zhang, Qiaohui Chu, Meng Liu, Xuemeng Song, Weili Guan, Liqiang Nie: OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026 https://arxiv.org/abs/2605.20818 https://arxiv.org/pdf/2605.20818 https://arxiv.org/html/2605.20818
May 21, 2026 at 6:41 AM
arXiv:2506.05782v1 Announce Type: new
Abstract: This report presents our solution to the Ego4D Natural Language Queries (NLQ) Challenge at CVPR 2025. Egocentric video captures the scene from the wearer's perspective, where gaze serves as a key [1/4 of https://arxiv.org/abs/2506.05782v1]
June 9, 2025 at 6:06 AM
Wei-Cheng Lin, Chih-Ming Lien, Chen Lo, Chia-Hung Yeh: GazeNLQ @ Ego4D Natural Language Queries Challenge 2025 https://arxiv.org/abs/2506.05782 https://arxiv.org/pdf/2506.05782 https://arxiv.org/html/2506.05782
June 9, 2025 at 6:06 AM
arXiv:2506.03710v1 Announce Type: new
Abstract: In this report, we present our champion solutions for the three egocentric video localization tracks of the Ego4D Episodic Memory Challenge at CVPR 2025. All tracks require precise localization of the [1/4 of https://arxiv.org/abs/2506.03710v1]
June 5, 2025 at 6:08 AM
Yisen Feng, Haoyu Zhang, Qiaohui Chu, Meng Liu, Weili Guan, Yaowei Wang, Liqiang Nie: OSGNet @ Ego4D Episodic Memory Challenge 2025 https://arxiv.org/abs/2506.03710 https://arxiv.org/pdf/2506.03710 https://arxiv.org/html/2506.03710
June 5, 2025 at 6:08 AM
exo-centric video datasets -- constructed globally with substantial effort -- of Exo-Ego4D to extract diverse manipulation trajectories at scale. From these extracted trajectories with the associated textual action description, we develop trajectory [3/5 of https://arxiv.org/abs/2506.03605v1]
June 5, 2025 at 6:04 AM
achieves first place in this challenge at CVPR 2025, establishing a new state-of-the-art in long-term action prediction. Our code will be released at https://github.com/CorrineQiu/Ego4D-LTA-Challenge-2025. [4/4 of https://arxiv.org/abs/2506.02550v1]
June 4, 2025 at 6:07 AM
arXiv:2506.02550v1 Announce Type: new
Abstract: In this report, we present a novel three-stage framework developed for the Ego4D Long-Term Action Anticipation (LTA) task. Inspired by recent advances in foundation models, our method consists of three [1/4 of https://arxiv.org/abs/2506.02550v1]
June 4, 2025 at 6:07 AM
Qiaohui Chu, Haoyu Zhang, Yisen Feng, Meng Liu, Weili Guan, Yaowei Wang, Liqiang Nie: Technical Report for Ego4D Long-Term Action Anticipation Challenge 2025 https://arxiv.org/abs/2506.02550 https://arxiv.org/pdf/2506.02550 https://arxiv.org/html/2506.02550
June 4, 2025 at 6:07 AM