Kimi open-sourced the PerceptionBench visual perception benchmark; GPT-5.6-Sol has the highest accuracy, but no model has surpassed 60%

robot
Abstract generation in progress

PANews July 29 reported that Kimi announced the open-sourcing of its multimodal model visual perception evaluation benchmark, PerceptionBench, aiming to decompose visual perception capabilities into 10 atomic abilities for independent assessment. It covers dimensions including visual relationships, counting, attributes, depth and 3D, localization, comparison, fine-grained recognition, contextual integration, OCR, and hallucination detection. The benchmark is built from failure cases of models in 42 existing evaluation sets. It includes 3,000 manually verified questions; each question tests only a single visual ability and requires no reasoning or external knowledge.

The evaluation results show that among 16 leading multimodal models, no model’s overall accuracy exceeds 60%. GPT-5.6-Sol ranked first with 59.7% accuracy, followed by Kimi K3 (58.5%), Claude-Fable-5 (57.2%), Gemini-3.1-Pro (56.2%), and GPT-5.5 (55.8%) in the top five. The report notes that visual hallucination remains the weakest capability across all models, and there is still significant room for improvement in overall perception performance.

View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned