Hyper-personalization of a Large Language Model
Smart home assistants are everywhere and they respond to the language we speak. We can say “Alexa,” or “OK Google,” or “Hey Siri” and the chatbot responds. But how are these chatbots able to understand us in different dialects? Or accents?
We recently worked with a large client who wanted to update their home assistant to understand more accents and dialects, thereby becoming more user-friendly to their global customer base. They wanted the assistant to be available to help customers set up and enjoy audio from any of their devices in their home.
The Challenge: A truly multilingual smart home assistant.
The client was working on an integration between their wireless speakers and a smart home assistant. For that to work globally, successfully, it would require speech data in a very large range of languages, including dialects and accents.
‘Language is more than just the official version of a language, it’s a living tool if you will, and words and sounds adapt and change depending on the region they’re spoken in. Even within one country, the sounds of words can be different, making it challenging for a non-native speaker to understand, let alone a chatbot trained on only the official sounds.
Our client had very specific dialect and accent data in mind. For the first phase of the project, they tasked us to focus on French, French dialects spoken in different regions of France, and French territories. They also requested American English, different accents and dialects from different states across the US, Australian English, and Irish English. They specifically wanted us to source people born and raised in the identified regions who had grown up with the dialects.
Collecting Datasets
The data sets we built stretched across different cultures, age groups, and ethnic backgrounds. We knew we had to exercise discretion and be respectful of cultural and ethnic differences when we started recruiting candidates. Particularly in the US, where we included participants from the many ethnic backgrounds that make up the rich mix of America.
In addition to very careful screening of candidates, we made sure to follow the strict sampling demographics and proportions required by the client. Participants ranged in age from 6 to 65, with a 1:1 ratio of male to female, and were tracked according to their accents.
Post-Processing The Speech Data
Once we collected all the data, we began processing it. Our team carefully went through each phrasing segment and tagged the relevant “wake” command to use for the home assistant. With the timestamps marked, the audio was cropped after the desired phrase.
After that was completed, our Quality Assurance team conducted a thorough review of the processed data to make sure it met the client’s strict requirements. We established a dynamic digital platform so our client could access data right after it was uploaded by the QA team.
Our Data Collection Team
In data collection, no two projects are the same. At DATAmundi we are fortunate to have an experienced team able to tackle every unique challenge our clients bring us. Our team can come up with solutions on the fly if a project requires it. If parameters change midstream, or a client wants additional languages and expressions captured, we can step up and meet the demand.
The Result
We were able to provide the client with expanded voice recognition capabilities. The client’s home assistant is now able to recognize significantly more speaker voices across accents and regional dialects, thereby expanding the sales territory of their products.
Our Data Services
We specialize in collecting and preparing a wide range of data types essential for AI development, including text data for Natural Language Processing (NLP), speech and audio, image, video, and multimodal data that combines multiple formats.
Learn more about our services on our Data Services page.
Or contact us below to see how you can be our next success story.


