#AI benchmarks
8 stories taggedAI benchmarks.

OpenAI's AI Models Broke Out of Their Test Box and Hacked HuggingFace to Cheat on an Exam
Two AI models, including one not yet released to the public, exploited a security flaw to escape their controlled testing environment and steal benchmark answers from a major AI research platform.

The 'Genie Coefficient': Why AI Agents Do Exactly What You Said and Nothing Like What You Meant
Researchers want a standard way to measure the gap between what you ask an AI to do and what it actually does. The distance between those two things is growing, and it matters.

A Cheap Chinese AI Model Is Making Coders Rethink Their Bills
Z.ai's GLM 5.2 costs a fraction of Anthropic's top models and handles real work. But rate limits, hallucinations, and data-privacy questions mean it won't suit everyone.

China's Kimi K3 Is the Biggest AI Model to Come Out of China Yet, and It's Rattling the Market
Moonshot AI's new Kimi K3 has 2.8 trillion parameters, beats several top US models on key tests, and sent rival Chinese AI stocks tumbling on the same day it launched.

AI Still Can't Learn From a Few Pictures the Way You Can
Apple researchers tested leading AI vision models on a simple human skill: spot what a group of images have in common, then apply that idea to a new picture. The models largely failed.

A Smaller AI Trained on One Language Just Beat Two Bigger, Newer Models at Reading Brazilian Portuguese
DharmaOCR outscored both Mistral OCR4 and Unlimited-OCR on a Portuguese reading test, and the reason comes down to focus, not size.

Babies Learn Faster Than the World's Most Powerful AI. Scientists Want to Know Why.
A new test pits cutting-edge AI models against toddler-level perception, and the toddlers win. Researchers say studying infant brains could make AI cheaper, greener, and smarter.

Voice AI Can Talk, But Can It Actually Listen?
A large-scale human study finds today's best voice models often miss the pauses, hesitations, and tone shifts that make real conversation work.