Challenge Description: The workshop hosts a challenge on scene-aware referential gesture generation, built on the MM-Conv dataset, which captures multimodal conversational interactions in 3D environments with synchronized speech, motion capture, and 3D scene representations. Its annotations for referential gestures provide a unique testbed for evaluating whether generated motion is temporally aligned with speech and spatially grounded in the environment.
Task: Given a spoken utterance, conversational context, and a virtual 3D scene, participants must generate co-speech referential gestures that are both communicatively appropriate and spatially consistent with the referenced objects. This task requires jointly reasoning about language semantics, scene geometry, and motion dynamics.
Evaluation: Submissions are evaluated along three independent axes. There is no single composite score; each axis is reported and ranked separately (Grounding STG (Score₁ₐ × Score₁ᵦ), Motion quality FGD, Human preference Perceptual study). STG combines temporal alignment (Score₁ₐ) and spatial grounding (Score₁ᵦ) into a single spatio-temporal grounding score. FGD measures motion quality against ground-truth statistics. The perceptual study is an organizer-run human evaluation of naturalness. The same frozen apex detector and objective evaluation pipeline are applied to every submission.
Submission: Please find detailed submission information below.
For full details, see the challenge paper below and the HuggingFace page. Access the dataset here .
Submission deadline: August 8, 2026
Notification: August 10, 2026
All deadlines are AOE.
Contact us at hsi-workshop@googlegroups.com for any question regarding the challenge.
Test set submission is now open!
The test set contains 100 anonymized examples. Generate one SMPL-X motion for every sample id, specified in the manifest in the test submission folder. Please read the README for further instructions.
Each participating team must submit a system description paper (up to 6 pages, excluding references). The paper should describe your method, training procedure, and any ablations or design choices. Accepted papers will appear in the ECCV 2026 Workshop proceedings.
Submit your motions and paper together in a single ZIP:
Submission format:
team-name_submission.zip
└── motions/
├── test_2.npz
├── test_11.npz
└── ...
└── paper.pdf
Submit here. Don't forget to validate your /motions folder before zipping with the provided validator script.
Interactive visualizer - Inspect pointing gesture samples from dataset:
Explore pointing gesture samples from the dataset. Each clip shows motion around the gesture peak frame. Full motion sequences and annotations are available on the HuggingFace dataset page.
Read the challenge paper: