BTC
ETH
HTX
SOL
BNB
View Market
简中
繁中
English
日本語
한국어
ภาษาไทย
Tiếng Việt

Kimi open-sources PerceptionBench visual perception benchmark, GPT-5.6-Sol leads in accuracy but no model surpasses 60%

2026-08-02 15:17

Odaily News: Kimi announced the open-sourcing of PerceptionBench, a multimodal large model visual perception evaluation benchmark, designed to break down visual perception capabilities into 10 atomic-level abilities for independent assessment, covering dimensions such as visual relationships, counting, attributes, depth and 3D, localization, comparison, fine-grained recognition, context integration, OCR, and hallucination detection. The benchmark is built on model failure cases from 42 existing evaluation sets, containing a total of 3,000 manually verified questions, each testing only a single visual capability without requiring reasoning or external knowledge.

Evaluation results show that among 16 cutting-edge multimodal large models, none achieved an overall accuracy exceeding 60%. GPT-5.6-Sol ranked first with 59.7% accuracy, followed by Kimi K3 (58.5%), Claude-Fable-5 (57.2%), Gemini-3.1-Pro (56.2%), and GPT-5.5 (55.8%) in the top five. The report notes that visual hallucination remains the weakest capability across all models, indicating significant room for improvement in overall perception abilities.