Behind the Screens: Fighting Hate with Data
Data annotation to help AI detect toxic in-game speech
One of our most challenging projects to date was one for a large, global gaming company. One of their most popular games is played in real-time by gamers all over the world, and they are able to interact with each other in the game; however, the rules of civility were often not observed.
It got so bad, and there were so many complaints that the company decided to build a bot that could detect and block bad language and – sometimes truly heinous – threats to other players. To achieve this, they contracted with DATAmundi to annotate and tag the speech they wanted to block from players. We were also asked to use the tagged data to create a script for their AI bot to learn from. We were asked to tag both the audio and the transcribed text.
The company also planned to contact those players who broke the rules with the intent to speak to them, limit their access to the game for a set period of time, or ban them from ever playing that game again.
That’s a lot of swear words!
To create a clean set of audio and text data, we first created scripts from over 100,000 pieces of chat data and audio recordings of people swearing, spewing hate speech, and threatening other players. Then we recruited people who agreed, in writing, to record themselves saying these swear words, hate speech, and threats.
We recruited people across a broad demographic range and focused on people who spoke English, Chinese, Brazilian Portuguese, or Korean. Including within those languages, the different accents you might find there.
We found the hardest part of this project was hearing and reading the absolutely horrible things some people said. That meant building in some time to just gather and talk through what we were exposed to.
Creating the Dataset
Once we had the scripts and the team assembled, we started recording. We checked in with our people regularly because a lot of the material was more than just swearing—it was truly harmful hate speech. Even then, it proved to be a slow process simply because of the amount of text in the scripts. There was such a variety of swear words and hate speech it took time to get good, clean recordings of all of it.
To meet our client’s request we knew we knew we had to get it right. Their reputation depended on being able to somehow mitigate the abuse that went on within their most popular game.
Post-Processing
To tag and annotate the data we’d recorded meant we had to listen to it all again. But if we did a good job, it would mean far fewer people would be exposed to this level of online abuse. Knowing that made the difficult task more bearable. In the end we were able to provide the gaming company with a very large dataset to use to train their AI to recognize and stop online abuse in the game.
The Result
The client took the dataset we provided and fed it into their LLM. The goal was to address the many complaints they had received from players, while not taking away from the overall online experience of their players.
The client wanted to maintain their reputation and comply with hate speech laws in the countries they operated in. They are now able to send alerts to offenders found by the system, and ban those who continue abusing other players.
Our Data Services
We specialize in collecting and preparing a wide range of data types essential for AI development, including text data for Natural Language Processing (NLP), speech and audio, image, video, and multimodal data that combines multiple formats.
Learn more about our services on our Data Services page.
Or contact us below to see how you can be our next success story.


