The Client
One of the world’s largest AI-driven technology platforms needed multilingual audio transcription at a scale. The brief was specific: tens of thousands of short audio snippets per month, across 14 languages, transcribed and annotated to a standard that could feed directly into ASR model training.
The files were real-world audio variable quality, multiple speakers, background noise, and the annotations had to reflect that accurately. Getting it wrong would corrupt the training data. The client needed a partner who could handle the volume without cutting corners on quality, and who had the infrastructure to scale when demand increased.
The Brief
Each week, the client sent batches of MP3 files, typically 10 to 30 seconds each, covering 14 languages: German, English, Spanish, Italian, French, Portuguese, Vietnamese, Indonesian, Russian, Romanian, Mandarin, Dutch, Turkish, and Polish.
Files fell into two categories. Single-speaker and multi-speaker recordings. Some multi speaker files sometimes overlapped and often had variable acoustic conditions. Both types required transcription plus annotation, tagging events like laughter, coughing, and overlapping speech per the client’s specific guidelines.
The project started with a few hundred files per week and has since scaled to 500 transcription hours per month.
Challenges
- Multi-speaker audio was the hardest part technically. When two people talk at once, AI transcription breaks down and overlapping segments had to be handled by human annotators who could separate the voices and tag the overlap correctly.
- Language labels came without dialect information, which created misassignment risk. A Spanish file could be Castilian or Mexican, and the wrong linguist could produce errors. DATAmundi built a triage process to catch mismatches and began developing dialect-detection functionality within our proprietary data management platform, AIDA Hub.
- Volumes changed week to week. The project started with five people on one language then ramped to 40–70 active linguists across all 14. That kind of ramp requires a recruitment pipeline and workflows that move fast without compromising quality.
- Keeping annotation consistent across a globally distributed team, spanning IST, PST, and EST time zones, required more than a style guide. The annotation standard had to be enforced through structured onboarding and ongoing calibration.
DATAmundi Solution
- Proprietary AI transcription handled the first pass; human specialists recruited through DATAmundi’s DATAtalent network handled the rest. DATAmundi ran incoming files through AIDA platform, using our proprietary Audio Transcription Tool (ATT) to generate draft transcriptions. Output was sent to linguists for review and correction.
- Every batch went through a 30% quality spot check. Specialist linguists handled transcription and annotation, then DATAmundi’s internal QC team reviewed a random sample of each language batch – 24,000 files a week -> a real investment in quality and the results reflected this.
- Linguists and annotators completed a 50-file pilot before taking on full volume to catch any potential guideline misunderstandings early and kept the annotation standard consistent as the team scaled to 100–120 files per linguist per day.
- The post-processing pipeline was built from scratch: .ass to JSON conversion, timestamp validation, batch packaging. When requirements changed, the scripts were updated.
- QC ran across IST, PST, and EST, which meant the team could absorb volume spikes without pushing deadlines. Weekly delivery targets were held consistently as the project tripled in scale.

The Results
The project ran weekly delivery cycles with no deadlines missed. DATAmundi processes 20,000–25,000 files per week across 14 languages.
The client ran their own automated QA after each delivery: an LLM-based script that checked transcription accuracy against their internal benchmarks. The pass rate stayed above 90% throughout, with no outright failures recorded. The Word Error Rate (WER) held below 25%, the client’s defined pass threshold, across every batch.
The Business Impact
ASR Model Development depends on clean, well-annotated training data. A poor transcription does not just affect one file; it affects what the model learns. The client needed a partner who understood that, and who built their process around it rather than hitting volume targets at any cost.
DATAmundi delivered the volume and the quality. The client got a reliable weekly pipeline they could plan against, with enough flexibility to absorb demand changes without renegotiating scope every few weeks.
Explore DATAmundi’s ready to use speech datasets and managed services for AI data workflows to see how this model appliesto other projects.
Running a multilingual AI data project?
Talk to us about what you need. Learn more about DATAmundi’s AI Data Solutions, or get in touch with us directly.


