GPT-5 Can't Count the D's in DEEPSEEK

Hey

Hope you had an excellent weekend wherever you may be. Mine was equal parts chilled, hanging out with visiting family, and battling my impostor syndrome as I prepare to moderate Monday’s upcoming Nobel Prize dialogue on AI’s role in the Future of Learning and Teaching. Done all the prep I can and decided it’ll all be fine in the end but wish me luck!

Anyway, a few notes from the world of AI to kickstart your week:

Welcome to the Trench of Despair: Why AI in Education Feels Like Institutional Chaos | Academic Reality Check 🏫

UQ's Jason M. Lodge just got back from UNESCO and EARLI conferences (here and here, respectively) with a bit of a reality check: we're having completely the wrong conversations about AI in education. Lodge identifies what he calls the "trench of despair" - that moment when AI hype collides with institutional breakdown. Policymakers build strategies on debunked concepts like "AI natives" and mythical AI detectors, while learning scientists studying actual impact operate in isolation from classroom realities where 90% of students already use these tools. The result is institutional paralysis disguised as progress.

Digital Learning WeekDigital Learning WeekFrom 8 to 11 September 2026, UNESCO Headquarters, Paris, France The influence of artificial intelligence (AI) in education continues to escalate. AI tutors are reaching learners who have never had access to personalized support, ‘agentic’ AI systems...unesco.org

Lodge warns we're facing a timeline crisis education can't afford. Unlike previous technology cycles that took decades to mature, AI's rapid deployment means systems risk being shaped entirely by corporate interests rather than pedagogical evidence. This timeline compression - where AI capabilities arrived years ahead of expert predictions - creates the institutional paralysis where policymakers can't keep pace with technological reality. The choice isn't between resistance and adoption - it's between evidence-based integration and letting Silicon Valley decide what learning looks like while democratic input becomes irrelevant. #nothanks

Research, policy, and pedagogy: The three conversations on AI in educationResearch, policy, and pedagogy: The three conversations on AI in educationI have spent the last couple of weeks at the European Association for Research on Learning and Instruction (EARLI) conference in Graz, Austria and UNElinkedin.com

The AI Empire Has No Clothes: How Silicon Valley Convinced Us Their Dominance Was Inevitable | Corporate Power Exposed 👑

Karen Hao's new book "Empire of AI" exposes how AI companies function like historical empires: extracting resources they don't own, exploiting cheap labor, and justifying it as a civilising mission. Kenyan workers like Mophat Okinyi suffered psychological devastation moderating sexual content for ChatGPT - losing his family after months of reading the internet's worst material for $2/hour. These "sacrifice zones" externalise human costs while companies promise educational transformation, spending billions to convince us their dominance is inevitable using the same rhetoric empires always use.

‘It’s destroyed me completely’: Kenyan moderators decry toll of training of AI models‘It’s destroyed me completely’: Kenyan moderators decry toll of training of AI modelsEmployees describe the psychological trauma of reading and viewing graphic content, low pay and abrupt dismissalsthe Guardian

Nothing about this trajectory is inevitable - it's manufactured. These companies need our data, energy, and labor to function, which means we have leverage. Whether it's artists glazing work to break AI training or communities blocking exploitative data centres, resistance is happening. As Hao emphasises throughout her analysis, these aren't natural laws but design choices - we could build task-specific systems that admit limitations rather than everything machines optimised for the appearance of omniscience. For the full analysis, Hao's book is essential reading, but Hasan Minhaj's recent interview with her offers an excellent preview of how we got trapped in systems that promise everything while delivering shareholder value.

Welcome to the Age of Wizards: Why AI Stopped Being Your Co-Worker and Became Your Oracle | User Experience Revolution 🧙

AI oracle Ethan Mollick has just identified something crucial arising with the newest round of models: we're no longer working with AI systems, we're summoning them. In his latest post, Mollick describes the shift from "co-intelligence" to the "wizard model" - where you prompt, wait, and hope the magic works. When he fed GPT-5 Pro his academic paper for analysis, nine minutes later he got sophisticated critique that found an error no human reviewer had caught. The problem? He has no idea how the AI did it or whether its other claims are accurate. We've moved from collaboration to being an audience for technological performance art.

On Working with WizardsOn Working with WizardsVerifying magic on the jagged frontieroneusefulthing.org

This shift happened faster than anyone predicted. Those 2022 forecasts saying AI wouldn't achieve gold-medal math performance until 2030? Both OpenAI and Google hit that benchmark this year. That timeline compression creates Mollick's dilemma: how do you train someone to verify work in fields they haven't mastered when the AI itself prevents developing that mastery? Every time we hand complex work to a wizard, we lose the chance to build judgment needed to evaluate its output. The fairy tale lesson applies: the better the magic, the deeper the mystery.

Assessing Near-Term Accuracy in the Existential Risk Persuasion Tournament — Forecasting Research InstituteIn June–October 2022, we convened 169 people to participate in the “Existential Risk Persuasion Tournament” (XPT). The XPT participants included both superforecasters with proven forecasting track records and domain experts with subject-matter...Forecasting Research Institute

OpenAI's Confession: Why Language Models Can't Help But Make Things Up | Technical Breakdown 🧠

OpenAI just published research that basically admits what critics have been saying all along: hallucinations aren't a bug, they're a mathematical inevitability. Their new paper "Why Language Models Hallucinate" reveals that these systems are essentially trained to guess confidently rather than admit uncertainty - and the problem starts at the statistical foundation, not just in deployment. When they tested state-of-the-art models with simple questions like "How many D's are in DEEPSEEK?", the models returned answers ranging from "2" to "7" when the correct answer is one. Funny red-face moment when OpenAI's own GPT-5 Pro found an error in lead researcher Adam Kalai's dissertation that no human reviewer had caught in over two decades - but he has no idea how it did it or whether to trust the analysis... 🤔

why-language-models-hallucinate.pdfWhy Language Models Hallucinate Adam Tauman Kalai∗ OpenAI Ofir Nachum OpenAI Santosh S. Vempala† Georgia Tech Edwin Zhang OpenAI Septembe...cdn.openai.com

The research connects hallucinations to basic classification problems: if you can't reliably distinguish truth from fiction in training data, you'll generate fiction during output. But here's the socio-technical twist that cuts deeper than the math - current evaluation systems actually reward hallucination. Most AI benchmarks use binary scoring that penalises "I don't know" responses while rewarding confident guesses, creating what the researchers call an "epidemic of penalising uncertainty." This same binary thinking explains why institutions build strategies on debunked concepts like AI detectors - systems designed to reward confidence over accuracy. The solution isn't better detection tools - it's changing how we score AI systems to stop rewarding overconfident bluffing over honest uncertainty.

How to make a 100% accurate AI detector – or, why we need specificity in discussions of AI detector accuracyHow to make a 100% accurate AI detector – or, why we need specificity in discussions of AI detector accuracyOpenAI reportedly has a ChatGPT detector that’s reportedly 99.9% accurate. Well I can do you one better. I can write you an artificial intelligence delinkedin.com

When AI Makes Movies: Runway's Video Revolution Hits Professional Territory | Creative Disruption 🎬

Runway's Aleph model, released in July, crossed into professional utility with this week's fifth Gen:48 competition winners. Over 1,600 submissions showcased Hollywood-grade video editing through text prompts: changing camera angles mid-scene, adding objects, transforming lighting, even aging actors. The results produced unexpected magical realism - kids with Xbox controllers alongside glowing cats in 90s lounges, men with earth-heads talking to counsellors with shifting glass faces. The films don't replicate traditional cinematography but create something distinctly AI-native, where the impossible feels normal.

Runway Research | Introducing Runway AlephRunway Research | Introducing Runway AlephRunway Aleph is a state-of-the-art in-context video model, setting a new frontier for multi-task visual generation, with the ability to perform a wide range of edits on an input video such as adding, removing, and transforming objects, generating any...runwayml.com

This represents graduation from clip generation to comprehensive post-production workflows that traditionally required expensive VFX teams. The $50,000 prizes and Lionsgate partnerships signal industry recognition these aren't experiments anymore. We're watching real-time democratisation of professional filmmaking tools, where studio-quality effects become conversational prompts. But this cuts both ways - more creators can access sophisticated tools, yet it also floods the market with AI-enhanced content. The question is whether this elevates amateur work or devalues professional expertise when anyone can generate reverse shots, seamless continuations, and complex style transfers through simple text commands.


These aren't theoretical concerns - they're documented realities reshaping education right now. When AI capabilities arrive five years ahead of expert predictions, when evaluation systems mathematically reward confident fabrication over uncertainty, and when creative tools democratise overnight whilst Kenyan workers subsidise Silicon Valley promises, institutions can't keep debating citation policies. Lodge's timeline crisis, Hao's empire analysis, and Mollick's wizard problem converge: the choice between evidence-based integration and corporate-shaped learning has already been made by default. Universities will either acknowledge the mathematical inevitability of hallucination and systematic extraction driving these systems, or discover these realities through policy failures and competitive displacement.

Want more analysis? Check out the latest edition of Adjunct Intelligence where Dale Leszczynski and I dig into the idea of an AI Winter - what does it mean, is one coming. Speaking of AI video, he went hard on this one so well worth a look! 👀🍿