TR2026-136

Plan and Double-Check: Streaming Multimodal Q-Former for Online Robot Action Generation


    •  Hori, C., Korekata, R., Kambara, M., Masuyama, Y., Jain, S., Corcodel, R., Romeres, D., Le Roux, J., "Plan and Double-Check: Streaming Multimodal Q-Former for Online Robot Action Generation", Interspeech, September 2026.
      BibTeX TR2026-136 PDF
      • @inproceedings{Hori2026sep,
      • author = {{Hori, Chiori and Korekata, Ryosuke and Kambara, Motonari and Masuyama, Yoshiki and Jain, Siddarth and Corcodel, Radu and Romeres, Diego and Le Roux, Jonathan}},
      • title = {{Plan and Double-Check: Streaming Multimodal Q-Former for Online Robot Action Generation}},
      • booktitle = {Interspeech},
      • year = 2026,
      • month = sep,
      • url = {https://www.merl.com/publications/TR2026-136}
      • }
  • MERL Contacts:
  • Research Areas:

    Artificial Intelligence, Computer Vision, Machine Learning, Robotics, Speech & Audio

Abstract:

Integrating humanoid robots into daily life demands effective collaboration, where robots must accurately perceive human actions and context to work toward common objectives. Previous studies generate robot action sequences for human instructional videos, but most rely on offline processing with pre-segmented clips. In real-world interaction, however, robots must handle unsegmented multimodal inputs in a streaming manner. In this paper, we extend a Q-Former-based robot action generation framework to streaming processing. Our method extracts audiovisual features from short temporal chunks, incorporates leftcontext information into query embeddings, and incrementally generates action sequences and confirmation messages using a large language model. By carefully designing attention masks, the model can be trained efficiently in parallel, similar to offline methods. Experiments show that our streaming approach achieves low latency with less than 10% accuracy degradation compared to offline processing.