ICASSP 2027 Signal Processing Grand Challenge

The First Embodied Audition and Reasoning Challenge

Embodied Speech Perception under Oracle Navigation in Multi-Party Conversations

EAR-1 asks a foundational question: if navigation is executed correctly, to what extent can movement itself improve multi-party speech perception?

Official status: EAR-1 has been accepted as an ICASSP 2027 Signal Processing Grand Challenge.

Paired
fixed and moving arrays
Real
multi-party meetings
Disjoint
rooms and speakers
2
core tasks
4
leaderboards

Recording paradigm

One conversation, two acoustic viewpoints

Paired fixed and moving microphone arrays recording the same multi-party conversation under oracle navigation.
Paired fixed and mobile recordings support controlled analysis of the acoustic value of movement for speaker diarization and speech recognition.

Scientific motivation

Is a better acoustic viewpoint itself valuable?

Fixed far-field systems are vulnerable to distance, background noise, reverberation, and overlapping speech. Embodied devices can rotate toward active speakers or approach sound sources, potentially improving direct-to-reverberant ratio and spatial separability.

Movement also introduces self-noise and rapidly varying acoustic transfer functions. EAR-1 isolates this trade-off by using a human proxy to execute a predefined oracle-navigation policy, while participating systems receive only multi-channel audio.

Navigation is controlled. Speech perception remains the challenge.

Challenge tracks

Two tasks, paired fixed and moving conditions

Each task is evaluated separately on synchronized Fixed-Array and Moving-Array recordings of the same conversations.

01

ESD-ON

Embodied Speaker Diarization with Oracle Navigation

Assign consistent speaker labels to continuous multi-channel recordings, including regions with overlapping speech. Oracle voice activity detection regions are provided without revealing active-speaker identities.

Metric DER Input Audio + oracle VAD
Fixed-Array Track Moving-Array Track
02

ESR-ON

Embodied Speech Recognition with Oracle Navigation

Generate speaker-attributed transcripts from continuous, unsegmented multi-channel recordings. Systems jointly handle speech activity detection, speaker attribution, and recognition.

Metric cpWER Input Continuous audio
Fixed-Array Track Moving-Array Track

Dataset

Real, paired, and disjoint by room and speaker

The EAR-1 corpus will provide paired multi-channel recordings of real multi-party conversations captured synchronously by fixed and moving microphone arrays. A labeled development set will support model training and system development, while an evaluation set with hidden annotations will be used for final ranking. The two sets will be disjoint in both rooms and speakers.

What participants receive

  • Synchronized fixed-array and moving-array multi-channel audio
  • Ground-truth speaker activity, speaker identities, and attributed transcripts for development
  • Oracle VAD regions for ESD-ON, including evaluation
  • Official formats, scoring scripts, baselines, and evaluation configurations

What remains hidden

  • Navigation trajectories and device poses
  • Video, inertial signals, and close-talking audio
  • Evaluation speaker labels and reference transcripts

Evaluation

Transparent metrics and separate rankings

Fixed and moving conditions use the same protocols and are ranked independently, enabling controlled within-conversation comparison.

DER

Speaker diarization

Zero forgiveness collar. Overlapping speech is scored. With oracle VAD, the metric mainly reflects speaker confusion and overlap attribution.

cpWER

Speaker-attributed recognition

Speaker-wise hypotheses are concatenated in time and scored under the minimum-error mapping between reference and hypothesized speakers.

Schedule

EAR-1 timeline

All deadlines use Anywhere on Earth unless otherwise announced.

  1. Registration opens
  2. Development set and baseline systems released
  3. Registration deadline
  4. Evaluation set released and evaluation leaderboard opens
  5. Evaluation leaderboard closes
  6. System-description reports due
  7. Invited 2-page papers due
  8. Paper acceptance notification
  9. Camera-ready 2-page papers due
  10. EAR-1 session at ICASSP 2027 in Toronto, Canada

Participation

Core rules

01

External resources

Public datasets and pretrained models are allowed when cited, accessible to other teams, and fully disclosed. New resources cannot be introduced after evaluation data release.

02

Permitted inputs

Systems may use only the officially released audio and task metadata. Trajectories, poses, auxiliary sensors, and organizer-only annotations are prohibited.

03

Final submission

Each entry includes outputs in the official format and a report describing architecture, training, external resources, and decoding configuration.

Organizing team

EAR-1 organizers

Join the challenge

Study what movement adds to speech perception.