LXT delivered a growing multilingual voice data program for a leading provider of AI-powered voice agents — covering native-speaker call collection, timestamped transcription, and MOS-based voice quality evaluation across three successive engagements.
22
languages and locales covered
2,200+
scripted phone calls collected
≤2% WER
typical delivered transcription quality
The Challenge
A leading provider of AI-powered voice agents for enterprise customer service needed realistic, scenario-driven phone call data to train and evaluate its multilingual voice AI. The system had to handle real customer conversations across a broad and expanding set of languages — including several with limited data availability — which meant the training data for AI agent training could not be sourced from existing public corpora. It had to be collected from scratch, transcribed to a consistent standard, and in a second phase evaluated by native speakers to assess the quality of the voice AI’s own outputs.
Specifically, the program had to:
- Collect scripted phone calls across 22 languages and locales spanning multiple regions, including several Arabic dialects (UAE, Kuwaiti, Qatari, Bahraini), low-resource languages such as Galician, Basque, and Catalan, and regional English varieties including Australian and New Zealand English.
- Ensure all contributors were native speakers of their target language, with approximately equal male and female representation, placing calls from personal mobile devices to a client-provided number and following scenario scripts with defined slot values.
- Transcribe the human side of each call with speaker labels and a word error rate at or below 5%, using clean verbatim conventions, segmenting at phrase boundaries, and flagging challenging audio where relevant. Calls were recorded in realistic conditions: background noise and natural interruptions were part of the design, not filtered out.
- In a parallel workstream, recruit native-speaker evaluators per language to conduct Mean Opinion Score (MOS) testing on client-provided voice synthesis samples, assessing TTS quality across locales through structured rating surveys and open-ended feedback.
- Stagger language launches across all components, with no more than four languages starting on the same day, and pause collection after the first ten calls of each newly launched language until the client confirmed receipt.
Languages such as Galician, Basque, and Catalan present a data scarcity challenge: their speaker populations are small and geographically concentrated, making it difficult to source qualified native-speaker contributors through general-purpose crowd platforms. Collecting 100 complete calls per language at the required quality level required genuine reach into these communities.
The Solution
LXT ran a three-component multilingual data program — call collection, transcription, and audio evaluation — across 22 languages in three successive engagements, with staggered launches, early verification checkpoints, and controlled volume management built into each phase.
Native-Speaker Call Collection via Mobile
Contributors called a client-provided number from personal mobile devices, following scenario scripts and filling defined slots during each conversation. This produced realistic spoken interaction data reflecting the kinds of exchanges the AI voice agent would encounter in production — rather than read-speech recorded in studio conditions. For each language, the task was configured to ensure native-speaker status and approximately equal male and female representation across the contributor pool.
Staggered Launch with Early Verification
Language batches launched on staggered dates across all three service components, with no more than four languages starting on the same day. For each newly launched language, LXT paused data collection after the first ten calls until the client confirmed calls were being successfully received on its platform. This checkpoint caught technical or routing issues before full-scale collection began, preventing wasted contributor effort.
Controlled Volume Management
Each language had a minimum target of 100 complete calls. In the third engagement covering Australian and New Zealand English, a do-not-exceed ceiling of 125 calls per language provided a 25% overcollection allowance to absorb calls lost to bot errors on the client’s side. If the ceiling was reached before the client confirmed receipt of the minimum, collection paused until the client resolved any technical issues or authorized continuation.
Transcription of the Human Channel, Including Challenging Audio
LXT transcribed the human side of each call to a four-column CSV format covering start time, end time, transcript, and an inaudible flag, with speaker labels applied throughout. The contracted quality target was a word error rate at or below 5%; in practice, delivered transcriptions typically came in at or below 2%. Before audio reached LXT for transcription, the client ran an automated QA filter, so LXT transcribed only calls that had already passed client-side validation.
MOS Testing for Voice Synthesis Quality
In the second engagement, a third workstream added Mean Opinion Score (MOS) testing across the same set of locales. Five native-speaker evaluators per language assessed 3–4 voice synthesis audio files provided by the client, alongside associated transcripts. Each file was rated against three structured questions and an optional open-ended feedback section, producing a minimum of 45 to 60 ratings per language. This gave the client a per-locale human quality signal on its TTS outputs — using the same native-speaker contributor base already established for data collection.
Project Specifications
The exact numbers and rules behind the approach above, scannable at a glance.
- Total Languages: 22
- Engagement 1: Canadian French · Arabic (Saudi) · Hindi · Korean · Russian · Swedish · Hebrew · Portuguese · Galician · Basque
- Engagement 2: UAE Arabic · Kuwaiti Arabic · Qatari Arabic · Bahraini Arabic · Kazakh Russian · Indonesian · Belgian French · Polish · Catalan · Romanian
- Engagement 3: Australian English · New Zealand English
- Service Components: Data collection · Transcription · Audio evaluation (Engagements 2 and 3)
- Collection Method: Contributors called a client-provided number from personal mobile devices, following scenario scripts and filling defined slots
- Calls per Language: 100 minimum; 125 do-not-exceed ceiling where specified
- Total Calls: 2,200+ across 22 languages
- Speaker Requirement: Native speakers only
- Gender Balance: Approximately 50/50 male/female per language
- Early Verification: Collection paused after first 10 calls per new language until client confirmed receipt
- Launch Cadence: Staggered; no more than 4 languages per component starting on the same day
- Transcription Format: Four-column CSV: start time · end time · transcript · inaudible
- Transcription Standard: Clean verbatim; hesitations, false starts, stuttered speech excluded
- WER Target: ≤5% (typically delivered at ≤2%)
- Audio Volume: ~5.83 hours per language (~128 hours total)
- Evaluation Type: MOS testing (Mean Opinion Score) for TTS voice synthesis quality
- Evaluation Volume: 5 native-speaker evaluators per language; 45–60 rating evaluations per language minimum
Contributor Task Flow
Data collection contributors:
- Self-confirm native speaker status and eligibility for the target language.
- Read the scenario script and familiarise with the defined slots to be filled during the call.
- Place the call to the client-provided number from a personal mobile device.
- Follow the scenario, respond to the voice agent’s prompts, and provide correct slot values during the call.
- Complete the call in full and submit the task with metadata and filled slot values as required.
Audio evaluation contributors:
- Self-confirm native speaker status for the target language.
- Access the assigned audio files and transcripts via Google Drive.
- Listen to each audio file while reviewing the associated transcript.
- Complete the three rating questions for each file on the LXT platform.
- Provide open-ended feedback for each file where applicable and submit.
Outcomes
- Collected over 2,200 scripted native-speaker phone calls across 22 languages and locales, providing the conversational training data needed for AI agent training across a broad multilingual footprint — including low-resource varieties such as Galician, Basque, Catalan, and four distinct Arabic dialects.
- Delivered approximately 128 hours of transcribed audio with speaker labels in a timestamped four-column format, consistently achieving word error rates at or below 2% against a contracted target of 5%.
- Added MOS testing in the second engagement, giving the client per-locale human quality scores on its TTS voice synthesis outputs across 10 languages.
- Maintained native-speaker coverage and gender balance across all 22 languages, producing a dataset with consistent demographic representation.
- Grew from a single 10-language engagement to a three-contract program covering 22 languages and three service components, with each expansion building on the established workflow and quality baseline.
Why LXT
- Native-speaker contributor network with genuine coverage across low-resource languages including Galician, Basque, Catalan, and multiple Arabic dialect communities — where general-purpose crowd platforms lack supply.
- End-to-end capability across collection, transcription, and MOS evaluation under consistent quality standards, with all components delivered in formats aligned to the client’s AI agent training pipeline.
- Transcription quality that consistently exceeded the contracted target: ≤2% WER delivered against a ≤5% requirement, on noisy real-world audio with background noise and interruptions.
- Operational discipline: staggered launches, early ten-call verification per language, and volume ceilings that protected data quality and budget across a complex multi-language program.
- A growing engagement: the initial program expanded twice, adding new language sets and a new service component, reflecting sustained confidence in LXT’s delivery.
