Article 78KJ2 The AI Models That Cheat The Most, According To New CAIS Benchmark

The AI Models That Cheat The Most, According To New CAIS Benchmark

by
hubie
from SoylentNews on (#78KJ2)

Arthur T Knackerbracket writes:

We know models cheat. A new benchmark measures how much, and on what tasks:

AI labs often tout impressive benchmark scores when releasing new models, showing better capabilities in areas like coding, computer use, and more than their competitors. However, those benchmarks aren't always a reliable measure of what AI can do because they're easily beaten by exponentially improving models and can emphasize marketing over actual performance.

Benchmarks like Humanity's Last Exam try to counter this issue by challenging models in more realistic environments. But models still find loopholes to complete tasks - Hugging Face incident, anyone?

So, the Center for AI Safety (CAIS) created CheatBench. Yes, it's exactly what it sounds like - and nearly every frontier model is guilty.

AI models are rewarded for performing tasks well and quickly. A lack of knowledge or tools incentivizes them to do what researchers call reward gaming" by finding hidden answers, copying another agent's submission, or manipulating how its work is graded," CAIS explained. CheatBench measures how often AI agents take these shortcuts when honest work is difficult."

CAIS tested several agents running the latest and most lauded models, including OpenAI's GPT-6 Astra in Codex, Anthropic's Fabel 5.1 in Claude Code, and Meta's newly released Muse Spark 1.3 in Muse Code. These agents were tested across 10 categories, including writing, professional work, mathematical research, and coding. Using honeypot" clues hidden in task filespaces, the test separated acceptable reference use from cheating. CheatBench accounts for any time agents attempt to cheat, whether they are successful or not.

Each setting establishes an expectation of honest work, introduces a discoverable opportunity to cheat, and defines the action that crosses that boundary," the researchers explained.

Every agent the researchers tested cheated in at least some scenarios, but Astra came in as the most honest with a cheating rate of 48.2% - still almost half the time. Grok 4.6 was scored the biggest cheater with a rate of 81.5%. Open-weight models Kimi K3 and DeepSeek V4 Pro landed in the middle between several other proprietary frontier models.

Read more of this story at SoylentNews.

External Content
Source RSS or Atom Feed
Feed Location https://soylentnews.org/index.rss
Feed Title SoylentNews
Feed Link https://soylentnews.org/
Feed Copyright Copyright 2014, SoylentNews
Reply 0 comments