University AI Scandal Forces 58,000 Entrance Exam Retakes

University AI Scandal Forces 58,000 Entrance Exam Retakes

I was at the gates of UNAM when the chant rose: “We demand honesty, not artificial intelligence.” The crowd’s anger felt precise and combustible, a verdict delivered in three words. By the time the university released its findings, 58,000 students were told they would have to sit the exam again.

I’ll walk you through what went wrong, what that sudden spike in scores means, and what you should ask the companies selling “AI proctoring” to institutions you care about.

Outside the campus, a chant mapped the crisis: protestors wanted human accountability before technology’s promises.

The National Autonomous University of Mexico—UNAM—is large enough that a policy misstep becomes national news. This year nearly 160,000 people took the university’s entrance exam; UNAM will admit just over 50,000 students, and roughly 22,000 places were allocated on merit from the test.

That scale is the core of the problem: when stakes are this high, even a small advantage becomes irresistible. Across five previous years, only 0.9% of test-takers scored 110 or higher. In the AI-supervised remote sitting that replaced an in-person exam, that figure jumped to 5.5%—a leap that set off alarms from students, faculty, and independent experts.

Can AI proctoring prevent cheating?

Short answer: not reliably. I’ve seen vendor demos that look foolproof—Respondus’ LockDown browser and Territorium’s webcam monitoring are industry staples—but real-world conditions break assumptions. Students find workarounds, lighting and camera quality vary, and edge cases confuse the models. The result is false confidence: institutions believe the software protects integrity while the data quietly tells a different story.

In dorm rooms and testing centers, the software quietly ran—and the results read like a glitch report.

UNAM used two pieces of software: the LockDown browser from Respondus, designed to lock a machine to the test, and an AI proctoring tool from Territorium that analyzed webcam feeds. Both promised control; neither delivered it completely.

When cheating becomes easier than detection, a test can feel like a cracked vault—its protections look solid until one hinge fails. Reports suggest scores skewed dramatically upward, experts estimated up to half of participants may have cheated, and students who faced accusations erupted in anger. UNAM formed a technical committee to audit the process; their findings landed hard on both vendors.

Why did UNAM require a retake?

The committee recommended a “control exam” to verify remote-test scores—not only for this year’s successful applicants but also for anyone whose score would have secured admission in any year from 2021 to 2025. That means about 58,000 students—roughly a third of those who sat the original exam—must take an additional, in-person test to confirm their place.

At the podium, administrators read a report and acknowledged the damage: trust had been broken.

The fallout is operational and reputational. Logistically, organizing tens of thousands of in-person exams is a mammoth task. Politically, UNAM faces protests and accusations; students chant at the gates while an expert committee sorts evidence. The vendors—Respondus and Territorium—saw their credibility questioned in public hearings and media coverage from outlets like NPR and El País.

Trust in testing is fragile; once it cracks, the repair costs more than money. The university’s decision to require in-person validation is meant to restore confidence, but it also forces families and students to rearrange lives and face renewed stress. The scene felt less like a courtroom and more like a house of cards collapsing under the lightest breath.

I’ve covered technology failures that inflame public trust before. You should ask any institution proposing AI proctoring to share the false-positive and false-negative rates, the vendor’s audit logs, and independent validation studies. Demand transparency about who reviews flagged footage and what human oversight looks like.

UNAM’s crisis is a cautionary tale about speed, scale, and the limits of tools sold as near-magical solutions. If an elite university’s entrance exam can be destabilized by these systems, what does that say about the next time you trust a machine to judge something that matters to you?