By Erik Vogt

ACL 2026 made Explainability of NLP Models its special theme. Although I was expecting to learn about models, new architectures, and expanding capabilities, it was clear that explainability was higher on the ACL agenda. How we understand AI systems and decide whether their outputs deserve our trust.  

The central question is no longer simply whether language models can perform a task. Increasingly, researchers are asking: 

  • How do we know when they are right? 
  • How do we identify consequential failures? 
  • What does a useful explanation look like? 
  • How can evaluation improve not only the model, but also the data and systems around it? 

The “Success Catastrophe” of Modern AI

Philip Resnik opened the conference with a keynote on the changing relationship between computational linguistics and AI. 

He described the field as confronting a kind of “success catastrophe.” Large language models (LLMs) have brought enormous attention, investment, and practical capability to computational linguistics. But that success also risks narrowing the field. 

A concern was that the industry can become dominated by benchmark improvement, model access, and industry priorities. A few points on an edit distance score may receive more attention than deeper questions about language, meaning, representation, and human communication. As I have asked often before, what exactly is a +5 edit distance score worth?  

Resnik’s argument was not that the field should reject LLMs of course. It was that computational linguistics should remember what makes it distinct and that LLMs are not the ONLY tool that matters.  

Language is not merely another input and output format for AI. It is a complex human system shaped by culture, context, intention, relationships, and history. Every computational approach contains assumptions about language, even when those assumptions are hidden inside a model. 

What Makes an Explanation Useful? 

Tania Lombrozo’s ACL keynote approached explainability from the standpoint of a cognitive scientist. Her question was not how machines should explain themselves, it was more fundamental: what do explanations do for humans? What happens when we explain, and why do we do it?  

Her argument was function-first. We should evaluate an explanation according to what it is intended to achieve. Explanations can support learning, discovery, generalization, prediction, or decision-making. A satisfying explanation is not necessarily a faithful one, and a technically accurate explanation may still be useless to the person receiving it. 

She also highlighted a difficult fact: humans are poor judges of our own understanding. We often believe we understand something until we are asked to explain it. It made me think about one of Richard Feynman’s many thoughtful quotes: “If you want to master something, teach it.”  Very appropriate for an academic conference.  

But that complicates AI explainability. We are asking models to produce explanations for users who may not reliably distinguish between genuine understanding and a plausible story. 

Why Evaluation is Becoming Part of the Architecture of AI

These topics appeared throughout the ACL conference in work on uncertainty, automated judging, data selection, annotation, red teaming, benchmark validity, and human-AI collaboration. The signal seems distinct: evaluation can no longer be treated as a final test performed after an AI system has been built. It must become part of the system itself. 

AI tools can help evaluate and improve training data by identifying likely annotation errors, duplicated examples, inconsistent labels, gaps in coverage, and samples on which human and machine judgments diverge. 

They can also help determine which data is useful. More data is not always better data. A carefully selected set of examples may improve a model more than a much larger set of repetitive or poorly targeted material. Or particularly bad data could do harm.  

The right question is not whether AI should replace human evaluators. It is how to divide the work intelligently. Machines are good at scale, comparison, repetition, and pattern detection. Humans are needed for context, consequence, ambiguity, and accountability. 

This is also why I don’t think human-only red teaming could ever be sufficient. Human experts remain essential, but they cannot explore the volume of scenarios, languages, personas, and interaction histories required to test modern AI systems adequately. Automated tools can broaden the search for failures, while humans determine which failures matter and what should change.

About the Author Erik Vogt, Vice President of Solutions at DATAmundi.

Erik Vogt is VP of Solutions at DATAmundi, where he helps frontier labs and high tech organizations design and deliver multilingual AI data, model evaluation and enterprise AI solutions. With more than 25 years of experience spanning AI, localization, language technology and business transformation, Erik has held leadership roles across the global language and AI industry.  

Connect with the DATAmundi Team