ICASSP 2026 in Barcelona once again proved why it remains one of the leading global events for audio, speech, and signal processing. It brought together the research and the people driving the next generation of speech and voice AI. 

A consistent theme was the field is shifting from more data to better data. ICASSP keynote speaker, and Stanford Professor, Emmanuel Candès, emphasized this in his opening presentation, along with the importance of context-rich signals. This is particularly significant given that speech-related AI tasks now span thousands of nuanced dimensions (for example, research into emotional modelling references up to ~34,000 distinct human emotional states), underscoring the complexity of truly human-centric AI systems. 

 

Here are five standout sessions at ICASSP 2026:

  1. NVIDIA: ASR Guided Preference Optimisation for Low Resource TTS
    NVIDIA’s work demonstrated a meaningful shift in how text-to-speech systems can be improved in low-resource settings. By leveraging automatic speech recognition (ASR) as a feedback mechanism, models can iteratively refine outputs without relying on large volumes of paired speech-text data. This approach directly addresses one of the most persistent bottlenecks in speech AI: the scarcity of high-quality multilingual datasets. DATAmundi addresses this – human data solutions for multilingual voice and speech AI. Clients and users need this level of quality and authenticity to deliver in the real world.  
  2. Alibaba: Continuous Emotional Control in LLM-based TTS
    Alibaba’s research pushes beyond categorical emotion tagging toward continuous emotional modelling. Rather than assigning simple labels like “happy” or “sad,” their approach enables fine-grained control across a vast emotional spectrum and they referenced tens of thousands of emotional states. This has profound implications for conversational AI, where nuance, tone, and adaptability are essential. It marks a step toward emotionally intelligent systems capable of responding dynamically in customer service, gaming, and digital assistants.
  3. Meta: Directional Multi-Talker Speech Understanding
    A standout theme across ICASSP was the gap between lab-grade datasets and real-world audio environments. Meta’s work directly addresses this by enabling large language models to interpret complex, multi-speaker scenarios. This identifies not just what is being said, but who is speaking to whom. DATAmundi has been working on this with several clients and AI Labs team –> to ensure AI systems are truly reflecting and understanding the messiness of voice and audio in the real world. This introduces a new layer of contextual awareness critical for applications such as meeting assistants, smart glasses, automotive systems, and robotics. The session reinforced this broader industry shift: clean, single speaker datasets are no longer sufficient. Future systems must be trained on overlapping speech, spatial audio, and real conversational dynamics.
  4. Adobe + MIT: STEMPHONIC Multi-Stem Music Generation
    STEMPHONIC presented a highly practical advancement in generative audio: the ability to produce music as editable stems separated by instrument. This enables real world workflows for creators in media, gaming, and advertising, where control and editability are as important as generation itself. The work highlights a growing demand for structured datasets where data is not only abundant but organized in ways that support downstream manipulation and creative control.
  5. Meta Workshop Insights: Context-Rich Audio Over Volume
    Complementing the technical papers, workshop discussions emphasised a consistent message: natural voice interaction depends less on sheer data scale and more on context-rich, high-quality audio. Real conversations include pauses, interruptions, background noise, and latency. These features often excluded from traditional datasets. Capturing this complexity is essential for building systems that feel natural and responsive.

 

Implications for Human Data in Voice AI

The next generation of audio and voice AI systems will not be defined solely by larger models, but by the authenticity and representativeness of the data used to train them. 

Multilingual coverage, emotional depth, multi-speaker understanding, and structured audio formats are emerging as foundational requirements. 

For organizations building voice AI systems, this creates both a challenge and an opportunity. The challenge lies in sourcing and curating data that reflects real human communication in all its complexity. The opportunity lies in leveraging high quality, human centric datasets to unlock more robust, adaptive, and globally relevant AI systems.

You can read more about how we helped convert large volumes of complex multilingual audio into reliable ASR training data for one of the world’s largest AI platforms using AI transcription and human-in-the-loop QA.

CASE STUDY: Transforming Real World Multilingual Audio into Quality Training Data

 


 

DATAmundi’s human data solutions span multilingual and multimodal speech data, ASR and TTS datasets (including low-resource language collection), transcription and human validation, enriched metadata and human-in-the-loop QA. We support use cases from conversational and domain-specific audio to creative audio applications.

 Discover how we can help you build better speech and voice AI. Connect with us here.