#benchmarking
Every summary, chronological. Filter by category, tag, or source from the rail.
Fixing Computer Use Benchmarks: Beyond Replay Exploits
Current computer use benchmarks are often gamed by 'replay agents' that blindly repeat successful trajectories. Robust evaluation requires stochastic, verified environments and honest statistical uncertainty to avoid costly deployment errors.
AI EngineerThe Benchmaxxing Plague: Why AI Benchmarks Fail Reality
Benchmarks are increasingly gamed by labs to inflate performance scores, leading to a disconnect between leaderboard rankings and real-world utility. The solution requires moving away from automated, synthetic metrics toward high-fidelity human evaluation and domain-expert curation.
AI EngineerDeepSWE: A Contamination-Resistant Coding Benchmark
DeepSWE is a long-horizon coding benchmark using 113 original, human-authored tasks to prevent model contamination and reward hacking, providing a more accurate assessment of frontier model capabilities.
AI EngineerTraining AI Models to Out-Think Hackers via Logic Benchmarks
Current AI models struggle with the 'logic leaps' required for cyber defense. Arithmetic and Hugging Face are building a new benchmark using human-discovered zero-days in blackbox environments to train models that can reason through complex, multi-step exploitation chains.
AI EngineerWorldLines: Benchmarking Long-Horizon Stateful Embodied Agents
WorldLines introduces a new benchmark and modeling framework designed to evaluate how embodied AI agents maintain state and execute complex, long-horizon tasks over extended periods.
SentinelBench: Evaluating Long-Running AI Monitoring Agents
SentinelBench provides a standardized framework for evaluating AI agents tasked with continuous, long-running monitoring, addressing the critical gap in testing agent reliability over extended time horizons.
Why Micro-Benchmarks Often Fail to Predict Production Performance
Benchmarks often report false improvements because they measure performance under ideal conditions—like warm caches—that rarely exist in real-world production environments.
Showing 7 of 7