#VirtualHome
I finished doing the outdoors ^^
#FF14 #FFXIV #FFXIVhousing #VirtualHome
February 1, 2026 at 12:52 PM
alternate universe where instead of deltarune dotcom/chapter5, it's virtualhome dotcom/myfinalhome
August 7, 2026 at 8:24 PM
SR improvement in VirtualHome and a +3.30% GCR and +2.11% SR improvement in Habitat over the best-performing baselines. [8/8 of https://arxiv.org/abs/2502.20742v1]
March 3, 2025 at 5:59 AM
benchmark covering 1,509 tasks across VirtualHome and Habitat 2.0, categorized into ultra-short, short, medium, and long tasks. Experimental results demonstrate that SPO significantly improves reasoning quality and final [6/8 of https://arxiv.org/abs/2502.20742v1]
March 3, 2025 at 5:59 AM
Jiaqi Xu, Tao Huang, Kai Zhang: Structured Self-Consistency:A Multi-Task Evaluation of LLMs on VirtualHome https://arxiv.org/abs/2602.00611 https://arxiv.org/pdf/2602.00611 https://arxiv.org/html/2602.00611
February 4, 2026 at 7:24 AM
Jiaqi Xu, Tao Huang, Kai Zhang: Structured Self-Consistency:A Multi-Task Evaluation of LLMs on VirtualHome https://arxiv.org/abs/2602.00611 https://arxiv.org/pdf/2602.00611 https://arxiv.org/html/2602.00611
February 3, 2026 at 6:29 AM
triggers the knowledge refinement process across domains. Experiments conducted on diverse embodied task benchmarks-including ALFWorld, VirtualHome, Minecraft, RLBench, and a real-world robotic scenario-demonstrate that [6/7 of https://arxiv.org/abs/2503.00870v1]
March 4, 2025 at 5:54 AM
of up to $3\times$. Finally, we expanded this pipeline by applying it to simulate complex household tasks in real-world scenarios, specifically in VirtualHome, enhancing the handling of failure cases. We release our code and [7/8 of https://arxiv.org/abs/2502.13170v1]
February 20, 2025 at 5:53 AM
some of the real use cases available in the domestic environments of VirtualHome. Additionally, since experiments with VirtualHome have shown the need to reduce the response time (which increases as the agent's decision space grows), we have proposed [5/6 of https://arxiv.org/abs/2505.02144v1]
May 6, 2025 at 5:58 AM
mid-level instructions while justifying the selection of these instructions. To validate its use in real applications we present a framework that integrates the reasoner into the VirtualHome simulator and compares its accuracy with GPT-4o, running [4/6 of https://arxiv.org/abs/2505.02144v1]
May 6, 2025 at 5:58 AM
fine-grained, intermediate reward for effective agent training. Extensive experiments on common agent benchmarks (including Webshop, ALFWorld, and VirtualHome) demonstrate that SPA consistently outperforms the state-of-the-art method in both success [6/8 of https://arxiv.org/abs/2505.20732v1]
May 28, 2025 at 6:04 AM