GPT-5.5 Scores 10.6% on ActiveVision, Humans Hit 96.1% [R]
The interesting finding from a new [arXiv paper]( https://arxiv.org/abs/2607.16165 ) isn't that a frontier vision model failed a new benchmark, that happens weekly, but the specific shape of the failure and the fact that the models cannot patch it by writing their own code. The benchmark, called ActiveVision, contains 17 tasks across 3 categories designed, in the authors' words, to "force repeated visual perception rather than a single static description." GPT-5.5 at the highest exposed reasoning-effort tier solves 10.6% of items and scores zero on 11 of the 17 tasks. Claude Fable 5, which the authors note tops most reasoning and coding leaderboards, manages 3.5%. Three human participants averaged 96.1%. submitted by /u/Justgototheeffinmoon [link] [留言]