You know you want the best quality data for your LLMs, and to get that, you already know you need to benchmark your data from time to time. There are benchmarking tools, such as fixed test sets and metric-only scoring, but those run the risk of introducing data contamination, ecological invalidity, and even ‘gaming’ the metric. Though there are fewer error-prone options available, too, such as multimodal evaluation, which provides better coverage.
Compare and contrast
We looked at some other companies in the data services and solutions sphere and were able to favorably compare our offerings against theirs. Overall, you’ll find we prize human review quite highly in benchmarking, primarily because humans are able to contextualize data, especially multilingual data, better than an AI tool can.
Human data raters are still considered the gold standard across the industry when it comes to catching tone, hallucinations, and bias. Humans are often also better at following instructions because they can understand and see how what they’re reading fits in with what comes before and after. Having human evaluators offers greater trust and safety in the data you’re using for LLMs.
The difficulty is that human review is slower and costs more, when what you want is to quickly get ready for the next iteration of your system.
Hybrid options
Some companies are recommending a hybrid solution with automation and human judgment at scale, and while that is certainly an option, at DATAmundi we employ a different method. You could still call it hybrid, but we emphasize human judgment in benchmarking, putting it first, especially when dealing with multilingual, domain-specific datasets.
Where others look at a hybrid model with a strong focus on automation, adding in human judgment at scale (potentially limited scale), DATAmundi, offers structured benchmarking with expert review and preference rankings.
The Hub
Most data solutions companies offer some kind of proprietary Hub to track or compare data. We offer SME/linguistic scoring, side-by-side model comparisons, all tracked and accessible in our AIDA Hub. Our AIDA Hub allows us to offer speed without sacrificing accuracy. We’ve automated repetitive steps, like task assignment and final reviews, while maintaining high standards that include human review and expert oversight.
When it comes to metrics, we focus on concrete deliverable metrics: accuracy, precision, NDCG (normalized discounted cumulative gain), and linguistic quality. It’s not uncommon in the industry for benchmarking and metrics to come from a BLUE/ROUGE scoring to evaluate translation and summarization by comparing reference text. Or to use perplexity to measure how well a model predicts the next text token; however, the higher the perplexity is, the more uncertain the data is.
This is why we choose human-in-the-loop every time, especially for evaluating multilingual content. We bring a strong multilingual and linguistic QA emphasis, along with a focus on ethical human AI, bias mitigation, and compliance with GDPR and other regulations.
Future forward
Looking to the future, there are developments on the horizon we’re keeping an eye on. Depending on how they develop, we may implement one or two of them for select projects. Developments we’re watching are:
- Humanlikeness Evaluation, which includes the Human Likeness Benchmark (HLB) test to see how closely LLM outputs match human writing and reasoning.
- Programmatic Task Benchmarks. These are benchmarks that organize human feedback, preferences, corrections, and explanations to automatically evaluate AI models.
- Human-AI Collaboration Benchmark Refinement is when humans use LLM-generated candidate judgments to iteratively refine evaluation criteria. An example of evolving human judgment at scale, while keeping a human-in-the-loop for final decisions.
Until that time, we agree that for your most important data projects and LLM benchmarking, particularly in critical industries such as healthcare, banking, or science, having human review as part of the process remains essential.
That’s an area where we excel. To learn more and to discuss your LLM benchmarking needs, contact us.
