The Client
A major technology company approached DATAmundi with a critical task: review the AI-generated responses behind their enterprise search platform. These responses covered highly specialized domains including finance, healthcare, and technology, so accuracy mattered. The client needed to confirm that outputs were grounded in real source material and did not introduce unsupported information that could damage user trust. With such high-quality expectations, they were looking for a partner able to deliver expert data annotation on a global scale.
The Brief
Execute a comprehensive evaluation framework to assess AI-generated summaries against source passages and human-generated responses across diverse specialized domains. Each evaluation needed to verify grounding accuracy, identify hallucinations, and meet the client’s exacting quality standards.
The project also required measurable reviewer consistency using inter annotator agreement (IAA) metrics, including Percent Agreement (how often reviewers selected the same label) and Fleiss’ Kappa (agreement beyond random chance). The client defined a target Fleiss’ Kappa score, a benchmark used to confirm that expert reviewers reached consistent conclusions rather than agreeing by coincidence.
The Challenge
The difficulty was not defining the evaluation criteria but applying them reliably at scale.
The material covered multiple highly technical subject areas, which meant judgement calls could not rely on generic annotation guidelines alone. Evaluators needed enough subject familiarity to recognize when an answer was technically correct, partially supported, or subtly misleading. Small differences in interpretation between reviewers could quickly introduce inconsistency into the dataset.
Consistency, therefore, became a central risk. The project required tight calibration across reviewers, clear decision frameworks, and continuous quality monitoring to ensure that identical cases received identical judgements regardless of who reviewed them.
Early measurement confirmed this challenge: when reviewing the initial Fleiss’ Kappa scores, inter-rater reliability analysis showed low statistical agreement, despite reviewers appearing to agree in roughly 80–90% of cases. This discrepancy suggested that alignment was often superficial, with evaluation criteria being interpreted differently, particularly when determining whether an AI response introduced unsupported information
At the same time, the client’s internal review process was extremely thorough. Any inconsistencies or weak justification in the annotations would be flagged during expert validation. The delivery approach, therefore, had to prioritize reviewer alignment, documentation clarity, and defensible evaluation decisions from the outset.

Our Approach
- Structured expert review model: We deployed 12 domain specialists across three subject areas, finance, healthcare, and technology, organized into review groups of three annotators supported by one arbitrator per domain. Independent annotations were completed first, followed by arbitration to resolve disagreements and ensure consistent final decisions.
- Deploying domain experts: DATAmundi assigned evaluators with proven subject matter expertise in each required field. Using AIDA Hub, DATAmundi’s proprietary data management platform, candidates were matched to tasks based on their technical background, with additional external specialists brought in where needed to ensure full coverage.
- Configuring a project-specific evaluation setup: DATAmundi collaborated with the client to update an existing evaluation tool previously developed for their environment. The two-week refinement process involved iterative back-and-forth validation with the client to ensure reviewers could efficiently compare AI outputs, source passages, and human responses side by side.
- Maintaining strict quality oversight: Quality checks were built into the process from the start. Annotator performance was monitored continuously, inconsistencies were flagged for review, and calibration steps ensured standards stayed consistent across all domains.
- Improving reviewer alignment through calibration: Outlier analysis (identifying reviewers whose decisions consistently differed from others), guideline refinement, arbitration reviews, and structured feedback cycles were introduced to address disagreement patterns and standardize interpretation across evaluators.
The Results
- Strong validation: The client’s technical lead, known for maintaining very high engineering and quality standards, confirmed that the delivery met their expectations. The project included multiple structured rework and calibration stages, which progressively improved reviewer alignment and strengthened annotation consistency across domains.
- Expert-led, technically sound evaluations: By aligning domain specialists to each subject area, the project ensured that every assessment reflected genuine subject matter understanding. This allowed evaluators to identify subtle grounding issues and unsupported claims that non-specialists could easily overlook.
- Establishing reliable evaluation consistency: Following iterative refinement, all domains achieved the client’s target Fleiss’ Kappa score, demonstrating strong and reliable agreement between independent expert reviewers.
The final dataset was accepted after collaborative quality improvement cycles, confirming that the evaluation framework could meet strict internal review requirements while handling complex and subjective AI assessment tasks.
The Business Impact
The completed evaluation gave the client a reliable way to assess their AI system’s performance across specialized domains, strengthening confidence in the accuracy of their enterprise search platform and the quality of the responses it delivers to users.
By establishing measurable and repeatable evaluation standards, the client gained confidence that improvements observed in model performance reflected genuine progress rather than reviewer variability.
The resulting dataset also gave their technical team clear visibility into where outputs were well supported and where further refinement could improve performance.
Beyond the immediate delivery, the project established a repeatable evaluation approach that the client can use as their AI systems evolve. With access to experienced domain evaluators, a configurable review environment, and established quality controls, they now have a trusted setup in place for future validation work across additional subject areas.
Need support with AI data evaluation?
Get in touch to discuss your next AI project with one of our experts.


