ICRA 2026 Highlights: Reinforcement Learning, Complex Behavior Generation, and the Future of Robotics
Яндекс
Huawei
NVIDIA
OpenDriveLab
Amazon/AWS
Waabi AI
The International Conference on Robotics and Automation (ICRA) 2026 in Vienna showcased key trends: Reinforcement Learning (RL), generation of rare edge cases, and the evolution of ML in robotics. Yandex autonomous driving team highlights best papers, workshops, and notable research on robot learning, perception, and autonomous vehicles.
At ICRA 2026 in Vienna, the main topics included Reinforcement Learning, generation of complex behavior scenarios, and the future of robotics. Best Conference Paper Award went to SymSkill for data-efficient long-horizon manipulation and OmniRetarget for humanoid whole-body loco-manipulation. In Robot Learning, a paper on view-invariant policy learning with camera conditioning (Do You Know Where Your Camera Is?) showed that VLA models degrade significantly when cameras are moved, and proposed encoding camera positions via Plücker coordinates per pixel, improving performance even without camera randomization. In Robot Perception, FindAnything presented an open-vocabulary object-centric mapping system using multiple foundation models for exploration. At workshops, World Engine from Huawei, NVIDIA Research, and OpenDriveLab enabled an autonomous car to drive 200 km in Shanghai without intervention, with code released open-source. NVIDIA announced the AlpaSim challenge with a dataset from 25 countries, 2500 cities, and 1700 hours of driving. Other notable works include Residual Off-Policy RL for finetuning behavior cloning policies (Amazon Frontier AI & Robotics), FP3 (3D Foundation Policy using Uni3D encoder), Search3D (hierarchical open-vocabulary 3D segmentation), and COMPASS (cross-embodiment mobility policy framework from NVIDIA Research). In autonomous driving, Diffusion-guided Generalizable Enhancer by Waabi AI improved 3D reconstruction near tracks, Attention BEV enhanced BEV fusion with attention blocks, TreeIRL combined MCTS with IRL for safe urban driving in Las Vegas, and a system for annotating delayed and false AEB events addressed extreme class imbalance and label noise. Another paper proposed learning to drive by imitating surrounding vehicles, using nuPlan scenes and transforming good agents' trajectories to ego perspective.
- Abbreviations
- ICRA = International Conference on Robotics and Automation — Международная конференция по робототехнике и автоматизации
- RL = Reinforcement Learning — обучение с подкреплением
- ML = Machine Learning — машинное обучение
- VLA = Vision-Language-Action — модели, объединяющие зрение, язык и действия
- 3D = Three-dimensional — трёхмерный
- 2D = Two-dimensional — двумерный
- IL = Imitation Learning — обучение имитацией
- NDS = NuScenes Detection Score — метрика для оценки детекции
- mAP = mean Average Precision — средняя средняя точность
- MCTS = Monte-Carlo Tree Search — поиск по дереву Монте-Карло
- IRL = Inverse Reinforcement Learning — обратное обучение с подкреплением
- AEB = Automated Emergency Braking — автоматическое экстренное торможение
- BEV = Bird's Eye View — вид сверху
Source: Habr — хаб ML —
original
