#offlineRL
Full trajectory data enables efficient offline RL policy evaluation under linear realizability and concentrability; tighter analysis cuts needed trajectories. Submitted 3 Oct 2025. https://getnews.me/trajectory-data-enables-efficient-offline-rl-policy-evaluation/ #offlinerl #trajectorydata
October 7, 2025 at 3:24 PM
One-Step Flow Q-Learning (OFQL) uses a single forward pass instead of multi-step denoising, halving inference time and beating Diffusion Q-Learning on the D4RL benchmark. Read more: https://getnews.me/one-step-flow-q-learning-boosts-offline-rl-performance/ #offlinerl #diffusion
October 3, 2025 at 9:23 AM
DiSA‑IQL adds a robustness term to Implicit Q‑Learning, penalizing unreliable state‑action pairs; in simulated tests it outperformed BC, CQL and IQL with higher success. Read more: https://getnews.me/disa-iql-enhances-offline-reinforcement-learning-for-soft-robot-control/ #offlinerl #softrobots
October 2, 2025 at 9:36 PM
ICQL applies linear Transformers to infer local Q‑functions, yielding up to 16.4% higher returns on kitchen tasks and 6.3%–8.6% gains on Gym and Adroit benchmarks. Read more: https://getnews.me/in-context-compositional-q-learning-boosts-offline-rl/ #offlinerl #transformers #reinforcementlearning
September 30, 2025 at 5:20 PM
Researchers combined offline reinforcement learning with GPT‑4o to speed up multi‑agent path finding, cutting training from weeks to hours and boosting success while cutting collisions. Read more: https://getnews.me/offline-rl-boosts-multi-agent-path-finding-via-gpt-4o/ #offlinerl #gpt4o #mapf
September 29, 2025 at 12:02 PM
DAWM splits generation: a diffusion model predicts future states and rewards, while an inverse dynamics model infers actions. DAWM‑augmented data boosted TD3BC and IQL on D4RL. Read more: https://getnews.me/dawm-model-boosts-offline-reinforcement-learning-diffusion-actions/ #offlinerl #diffusion
September 26, 2025 at 2:30 PM
Researchers unveiled CDQAC, an offline RL algorithm that learns job‑shop scheduling from only 10‑20 historic instances and beats the heuristics and RL baselines. (12 Sep 2025) Read more: https://getnews.me/offline-rl-improves-job-shop-scheduling-using-limited-data/ #offlineRL #jobscheduling
September 17, 2025 at 5:38 AM
HACO achieved an AUC of ~0.81 and set a risk threshold τ≈0.038 (α=0.10) to safely guide Medicaid population health decisions. Read more: https://getnews.me/hybrid-adaptive-conformal-offline-rl-improves-fair-medicaid-care/ #offlinerl #medicaid
September 16, 2025 at 11:15 PM