Perception, Evidence & Efficiency for AI in the Physical World
This is an accompanying post for my ECCV 2026 CONTEXTUS workshop talk.
- Workshop: CONTEXTUS @ ECCV 2026
- Slides: Download slides
Abstract
To enable AI assistive systems that can operate in physical, real-world environments, we need to go beyond short clips, improve transparency, and enable efficient operation over hours of continuous video. This talk presents three foundational capabilities required for embodied AI systems that work in dynamic, high-stakes settings: Predictive Motion Understanding via UniEgoMotion [1], which anticipates human action from egocentric video to enable perfectly-timed assistance; Evidence-Backed Reasoning via E-VQA [2] , which grounds AI claims in dense spatio-temporal grounding so systems show their work rather than just confidence; and Streaming Efficiency via StateKV [3], which enables real-time inference over long videos.
References
Enjoy Reading This Article?
Here are some more articles you might like to read next: