In this first episode of the DATAmundi Podcast, host Louise Law is joined by CEO Veronique Ozkaya to explore the various dynamics in the global data for AI data market. As compute power and AI investment surge, the conversation focuses on a key success factor: high quality human data. 

Veronique highlights the key shifts shaping the industry, from growing demand for data sourcing, augmentation, and evaluation to the rapid rise of multimodal and audio data. The conversation explores why linguistically accurate, culturally aware AI is becoming essential, how AI adoption is maturing, and why demand for expert human judgment in specialist AI models continues to increase. 

LISTEN TO EPISODE ONE NOW

— Transcript

Louise Law: Hi everyone! Welcome to the very first episode of the DATAmundi podcast. I’m Louise Law, your host, and I’m delighted to have you all with us today as we kick off this series dedicated to exploring the world of human data for AI. In each episode, we’re gonna dive into the trends and real-world challenges shaping the data space, and we’ll feature conversations with leaders and experts across the industry.

Louise Law: Today, I’m joined by our CEO, Veronique Ozkaya, to get her take on the state of the AI data market.

Louise Law: What trends have caused us by surprise this year, and what are some of the opportunities ahead. Veronique, welcome.

Veronique Ozkaya: Thank you, Louise. Pleasure to be here. Thank you for having me. Happy to kick this off.

Louise Law: Brilliant. So, for the sake of our listeners that may not have, heard of you or heard of DATAmundi, do you think you could tell us a little bit about yourself and a little bit about DATAmundi?

Veronique Ozkaya: Sure. So, I’m the CEO at DATAmundi, and before that, I had a career in the language services, also known as the localization industry, and worked for a number of organizations, very much focused on my leadership as a growth leader.

Veronique Ozkaya: So I’ve been known for taking companies from small, relatively small size to doubling or tripling their sizes, and for working across a variety of sectors. So, I have worked for American organizations, very much focused on high-tech and life science. I’ve worked for European companies that had more of a public sector, e-commerce, so different sectors but similar types of services and at different levels of scales.

Veronique Ozkaya: And about a year and a half ago, I was recruited by a company called Sumalinguae that also was in that content and localization industry, but had a division that was dedicated to data services. And that was the reason I joined.

Veronique Ozkaya: And very quickly, I wanted to ensure that the company focused very much on that part of the business, because I saw, really, the enormous potential.

Veronique Ozkaya: It’s a segment, it’s a sector that’s growing very quickly. Why? Because as the compute is growing in AI.

Veronique Ozkaya: The fuel for all of these AI models and systems is data. So, we literally took that service and decided to build our strategy around it and to grow it significantly, and that also meant changing our identity.

Veronique Ozkaya: And rebranding to DATAmundi. DATAmundi was actually the name of a company that we acquired in 2021, so the name didn’t come, you know, out of the sky, but rather it was focusing on what we had, acquired and were going to develop. So that’s the background.

Louise Law: Sounds good. Well, I think there’s a lovely alignment there. Obviously, you’ve got a great experience as a growth leader, and you’re leading a company in a very, very high growth area, so I think there’s a really nice fit there, isn’t there?

Veronique Ozkaya: Absolutely, and if you look at the massive investment that’s made in AI, just to give you a data point, over 52% of VC investment to AI in 2025. It’s a sector, data specifically, that’s growing about 30% year-on-year. It’s estimated to reach over $100 billion by…

Veronique Ozkaya: four years’ time. And again, because that compute power, you know, all these data centers being, built, all of the hyperscalers investing in infrastructure, you know what? This infrastructure, it’s built to run on data.

Louise Law: So is that how you’d summarize the landscape at the moment, is it?

Veronique Ozkaya: Yes, and what’s shifting now is, like, so…

Veronique Ozkaya: Data is a critical factor for AI success, right?

Veronique Ozkaya: Also, the second thing to think about is not any type of data. It’s all about high quality, accessible, ethical data sets, and they’re essential for these AI applications. So we kind of see three things that are driving our services.

Veronique Ozkaya: Are driving, in general, the human data companies and providers. You’re seeing that AI adoption is maturing, right? That organizations are implementing AI, and they realize that they need the human quality data. If you’re going to set up an internal bot in your organization.

Veronique Ozkaya: to help, knowledge management. Instead of having to go on… I am, and ask a colleague for something, if you could ask a bot and say, how do I, in marketing, get access to this collateral? Then this internal system could give you that information without you having to ask a couple of people.

Veronique Ozkaya: For this to work, the input into the system, into the bot, needs to be clean and correct. So that’s what we’re seeing, and more and more companies are adopting AI, but if they don’t have the right sector-specific, topic-specific, curated data, it’s very complicated. And that’s why you see a rise in

Veronique Ozkaya: the demand for experts that are going to train, train AI. That’s one. The second thing, and that’s our specialty, that’s our differentiator, is the language gap.

Veronique Ozkaya: So, AI is in English. 90% of the models are in English. And, of course, now, the AI foundation models and builders, they’re looking for multilingual data.

Veronique Ozkaya: So they can spread that, you know, to… across different countries. And then the third, and not least, is ethics.

Veronique Ozkaya: So, governance, and what we’re seeing as a trend, and as a request or demand, which is very justified, is that clients are asking to, you know, where is the data coming from? Consent, all of these aspects around regulations are definitely growing and important. I think it’s a good thing. It’s the guardrails, right, in.

Louise Law: Yes, absolutely. I mean, I think I’ve read, you know, in a couple of areas, that data is the new oil. There’s all this huge investment in AI systems, implementing AI… without data, it’s never gonna work, really, is it? Something that I think is that you mentioned high quality, the importance of experts, because we don’t have good data without humans. We still need that human judgment and that human guidance to be able to make sure that the data is accurate and clean to train the AI models and, the value of language and culture, which is obviously where DATAmundi has its heritage and has its roots, so there’s, again, there’s something that… there’s a really nice alignment there of all the different factors coming together..

Veronique Ozkaya: I was going to give an example of data, because that question has come to me from a number of people saying, data, what is data? And data, of course, it’s text. It’s audio. It’s video. A very, very simple example that, you know, is not related to LLMs.

Veronique Ozkaya: Take autonomous driving. You’re in a car in, in your self-driven car, right? That car has been trained. Right? The system has been trained to recognize things, so… and take, unusual events. There’s fog, and somebody is crossing the street. This is, like, something that doesn’t happen a lot, right? But you still need to capture that.

Veronique Ozkaya: And there’s a danger, it needs to stop, or if it’s snowing, and I know in your part of the UK, it snows, right, in the winter a bit.

Veronique Ozkaya: signs, the road signs, if… right? If the road signs are covered in snow, how will the car understand this is happening? You need to capture this. So this is data, and this is fed into the system for the car to actually recognize what’s happening. So I just want to kind of explain that data is not just text. It’s also multilingual data sets, it’s also, you know, multimodal.

Louise Law: Yeah, absolutely, absolutely. So, was there anything that surprised you this year?

Louise Law: You know, it’s an area that’s constantly changing, constantly evolving. We’re hearing there’s new research coming out all the time. What surprised you about this year, Veronique, in terms of the industry, and also maybe some of the challenges that we’re seeing our clients facing?

Veronique Ozkaya: The main thing that surprised me was the significant increase in requests for data sourcing. So we offer three sets of services, or three buckets of services. You’ve got the sourcing, which is data collection, essentially, and also off-the-shelf data. Then, we have the services that are really related around data augmentation. We annotate, we label, we enhance these, we make the answers in the data sets better. I’ll give you an example of that. Let’s say you have a…

Veronique Ozkaya: an empty output, and all the professions, you know, can have a… have a gender in some languages. So, a nurse is… [nurse in French], If your MPT only says, like, the nurses are women, you have a bias. So this, you know, what we do is we fix that, right? And then you’ve got the evaluation. The evaluation is about trust and safety and benchmarking. So these are the three buckets.

Veronique Ozkaya: Bucket number one, the data sourcing, I was quite surprised to see a real resurgence. I was thinking, what’s happening? And essentially, we see three things. Number one, you know, the models initially were built by using publicly available data, quality data, but scraping the web and using all of this, you know, big volumes of data. So we’ve gone from volume to quality, very much. And the sourcing is that there’s a big push for a human-generated data set.

Veronique Ozkaya: And there’s scarcity there, so that’s tough, because you’ve got very, very specific domain data that is, that companies are looking for, and everybody’s looking for that. So if you’re going to build a model about drones, for example, and you need experts in drones in Slovakia.

Veronique Ozkaya: There is not tons of these people who can actually contribute to improving the world.

Louise Law: I was just gonna say, you know, you talk about, sort of, domain-specific activities, like any area, like law, medical, healthcare, you know, some of the STEM areas.

Louise Law: There’s no wiggle room for error, really, is there? So, that’s the importance of when you’re building and collecting that data, or sourcing that data, as we’re talking about at the moment. It’s very important to be able to have that trust and reliability in the data.

Veronique Ozkaya: And some of it is really hard to get by. So the second.

Veronique Ozkaya: The reason why we’re seeing so much data sourcing is because the AI adoption is growing. So you’ve got sectors even that would have been quite conservative, that have a mandate to adopt AI.

Veronique Ozkaya: Let’s take an example, let’s say take manufacturing. If you need to capture how a factory floor is functioning, a lot of the information is in PDFs, it’s in people’s mind. You know, how do you capture that? And of course, confidential data. So that’s also a challenge, you know, the scarcity is related to domain, but also the type of data that you’re trying to get. And then the third thing, of course.

Veronique Ozkaya: regulation, right? We’re back to that. Like, I mentioned the confidential aspect here, but also, where does the data come from? Did the person consent to give it? So, if the data is not clean, that’s it. You cannot defend your model. I think that’s a very important, so compliance, it’s like a risk, but it’s also an opportunity, because it’s also this kind of stamp of… Ethics, or ethical aspect that you put on the model.

Louise Law: Absolutely, absolutely. And I think that the three different areas you’ve mentioned before, I think that’s where it’s quite important for both the, you know, for us as a human data provider to know what the use case is for the actual AI system that the data’s being used, because if we know that, we know we can apply the right procedure, the right solution, the right technology to make sure it makes, you know, makes the marks. Would you like to talk a little bit about some of the use cases that you’ve seen this year, and how that’s affected, how we go about developing the actual data solution.

Veronique Ozkaya: That’s a super good point, Louise, and I really see that the use cases are getting more… varied, more complex.

Veronique Ozkaya: And they require more ingenuity in how to either source the data, or work on the data, and also being able to scale, right? Because you need to look at how you can make some tasks, repetitive tasks, more automated, so you can scale, and also, as we work across languages, that’s quite important. And…

Veronique Ozkaya: Sometimes it’s not very obvious, right, when you get a request from a client, and you work with the data scientist of, what is this going to be used for? So, very key to understand what is the goal with the data, what are the KPIs? So, if a company says, we need to improve the quality of a data set by 0.2%

Veronique Ozkaya: Why is that? What are you trying to achieve? So we’re always trying to figure that out, and to give you, example of, of use case. So we, we, we tend to specialize in LLMs, but… and also all the aspects around, trust and safety, so that’s the whole, the whole, is the model grounded, working on data sets, golden answers, all of these things, you know, is this compliant, etc. But let’s give an example of, of, LLM.

Veronique Ozkaya: Let’s say that you are, working on a legal LM for a big law firm. Well, the firm is going to use us to review and validate the data sets before it’s deployed. And that means that we need to identify the right profile of lawyers that are going to be reviewing these data sets before they use, and… because you cannot have mistakes there, and a generic annotator will not be able to do that. Let me give you an example of something that… that… that failed, last year when I was in an AI conference in New York.

Veronique Ozkaya: A big consulting company explained how they deployed an internal assistant. A text, you know, assistant internally.

Veronique Ozkaya: That was supposed to help all of the, their staff, and mostly it’s advisors and consultants and so on and so forth, getting answers and being able to, pick into the company global database of answers to standard questions.

Veronique Ozkaya: The usage, the adoption, was 95% on the first day of launch, so everybody was already eager to use this assistant.

Veronique Ozkaya: Within a week, it dropped to 5%. And the lady from the consulting company that was explaining what happened, she said, you know what? We had a fantastic technology and we really didn’t consider cleaning the data much better.

Veronique Ozkaya: You lose the trust. It’s exactly that. Another example of use case is, let’s say, you know, conversational agents.

Veronique Ozkaya: Right? Well, they need human-generated dialogues.

Veronique Ozkaya: That are going to reflect a certain expertise. So, if you are going to use an agent with your insurance company. You want to make sure that what the agent is telling you is based on cctual, real, human-generated.

Veronique Ozkaya: conversations that are happening. So, sourcing that, not, not, not an easy case. Or, another maybe example related that we did, we did a couple, a few months ago was working on a, on a voice assistant in Arabic, but it had to be Gulf Arabic.

Veronique Ozkaya: So what you need to do is, and you need to get that human-generated speech, but with the actual variant of Gulf Arabic, so that the system will understand. Exactly, accent and dialect, otherwise the AI will not understand what is being said.

Louise Law: That’s some really interesting use cases. I’m sure there’s a lot more to come as well, isn’t there? Absolutely. Do you think… I’ll throw a quick, you know, before we start wrapping up, do you think we’ll ever reach a ceiling where the… all the AI models are absolutely trained, and they’re brilliant, and there’ll be no… there’ll be no need to… to put more data in the tank, so to speak, so it will ever hit a sealing.

Veronique Ozkaya: I think it’s not going to stop for a long, long, long time.

Veronique Ozkaya: simply AI adoption. Right now, of course, the big consumers of data are in high tech and, you know.

Veronique Ozkaya: technology companies, etc. But if you think at AI adoption, we’re only at the very beginning, and I’m sure, you see, right, in the media, a lot of, data points, excuse the pun, that explain that a lot of companies are only at the beginning of the journey.

Veronique Ozkaya: Right? If you also see the amount of investment that is being put into computing power, it just shows you that this is really something that is completely changing how we work, how we will work, so I believe that data needs, I just, are going to not only increase, but they’re going to for sure evolve.

Veronique Ozkaya: I don’t think the end is in sight. You’re going to get to points where maybe for a specific kind of use case in a domain, etc, the system will self-learn, and it’ll get to a certain level of autonomy, if I can say, and then you’re going to have another 10 things that are going to open up. I mean, right now, there are… robots.

Veronique Ozkaya: Robotics is actually a big, a big field right now. Robots are being trained to do all sorts of tasks, including folding napkins, like a butler. And what does that mean, that you have professional butlers who are showing how a napkin should be folded. That’s something that I never thought.

Louise Law: Sort of letting it from humans.

Veronique Ozkaya: Yes, I think so. Also, it is also an exaggeration to some of the ways it’s going, but it also shows what the technology is capable of doing, so it’s fascinating to see what needs to get in, in order to have good output.

Louise Law: So, final question, Veronique, as we’re kind of heading into a new year, we’re looking at the road ahead. What are you feeling really passionate about? What’s really, really getting you energized for the road ahead?

Veronique Ozkaya: What’s really exciting is that, while continuing to provide the same services that we’ve been, you know, supplying in terms of data sourcing, augmentation, evaluation.

Veronique Ozkaya: that… The use cases are all going to be different and  I think that’s very exciting, because you see, really, the effect of what you’re producing. You also see that we’re contributing to making AI truly global.

Veronique Ozkaya: Not just in English, and giving access to these amazing tools and technology to billions of people around the world, so very exciting. And you really work with some of the best companies in the world on the most advanced.

Louise Law: We’re privileged to be right at the coalface of where all the innovation is happening.

Veronique Ozkaya: And coming with our background, coming from language services. I see this also as an opportunity.

Veronique Ozkaya: For what we call the contributors, so these are our… Partners in the supply chain to reinvent themselves and to leverage all of this knowledge that they have about contexts and cultures and how things are done in order to, to train the models. So that’s also very exciting to see that, there is a bright future. It’s a matter of changing a little bit the… the way you look at things and embracing, embracing the technology, and participating.

Louise Law: Of course, that sounds great. Veronique, thank you so much for joining me today and sharing your insights, not just on where the AI data landscape is, but where it’s going, what you feel very passionate about. I really appreciate your time.

Louise Law: And, you know, we’re excited to continue to explore different topics in the world of human data with other episodes, and I’m sure we’ll have you back on the podcast, you know, in the next couple of months as well. So, everybody, stay tuned, thanks for listening, and we’ll see you next time.

Veronique Ozkaya: Thank you, Louise.

Louise Law: Thanks, Veronique.