Written by Gert Van Assche

Last week, a series of seemingly unrelated events pointed to one big question: Can large language models (LLMs) actually think? Some experts say yes, others emphatically no — and the lines between hype, capability, and risk are becoming harder to draw.

 

Google, Nvidia and Mistral push forward

Nvidia’s Jensen Huang described how inference-time reasoning (like Chain-of-Thought prompting) lets LLMs solve multi-step problems dynamically — a key shift from training-based knowledge to real-time logic.

Mistral upped the ante with 2 new models: Magistral Small (open source) and Magistral Medium (enterprise-grade). Both emphasize a “reasoning-first” architecture, combining logical step-by-step problem solving with tool integration and database access.

Google’s latest preview of Gemini 2.5 Pro boasts major improvements in coding, reasoning, STEM, and multimodal tasks — along with user-requested fixes.
The competition for the smartest model is clearly heating up.

Wikipedia, Apple and Epoch AI  push back

Wikipedia recently suspended its pilot of AI-generated article summaries after an intense backlash from its editor community. The summaries, powered by Cohere’s Aya model and marked with a “yellow unverified” label, were meant to improve accessibility. But editors warned they were misleading, lacked transparency, and risked damaging Wikipedia’s hard-earned trust.

In its “research” piece, “The Illusion of Thinking”, Apple reminded the community that LLMs don’t genuinely reason — a view many researchers published about in the past year. LLMs may sound smart, but don’t confuse fluency with understanding.

In the fascinating Epoch AI study (https://epochai.substack.com/p/beyond-benchmark-scores-analyzing)  14 mathematicians analyzed how OpenAI’s o3-mini-high model handles complex problems. The results were weirdly human — not always in a good way. The model often attempted multiple solution strategies, even coding mini-experiments. But 75% of its reasoning paths included hallucinations, and it struggled to produce rigorous proofs.
Impressive? Yes. Reliable? Not yet.

And Users Are Already Trusting AI With Big Decisions

Sam Altman says how people use ChatGPT reflects their age – and college students are relying on it to make ‘life decisions’. Whether or not the models are ready for that kind of trust is up for debate — but it’s clear that the expectation is already here.

Our Take at Datamundi

If you’ve worked with LLMs seriously, you know this: the output can be dazzling — but to assess its truth, you need patience, subject matter knowledge, and the ability to dissect both the logic and the “facts.”

At Datamundi, we support AI system builders with the infrastructure, experts, and tools needed to test before you trust. Pre-deployment validation is no longer a nice-to-have — it’s essential.

A Final Thought

Tools like RealAvatar let you speak with avatars of real experts — like Andrew Ng. But to paraphrase B.F. Skinner, the sociologist who invented the famous operant conditioning chamber, the Skinner box: The real question isn’t whether avatars think. It’s whether humans do.

The challenge isn’t AI. The challenge is us.  We’ll need to be smarter, sharper, and better — because the future won’t wait.
PS — Andrew Ng’s avatar loves this post. Do you?