Skip to main content
Back to timeline
arXivSource publication:

MBZUAI team evaluates GPT-6 Astra and five other general systems across 34 capabilities and 55 benchmarks, finding semantic understanding near reference levels while precise geometry and faithful reconstruction remain gaps

Synopsis

Researchers at MBZUAI and Apertix placed GPT-6 Astra alongside Fable 5, Kimi K3, Gemini 3.1 Pro, Qwen 3.8-Max, and Muse Spark 1.3 under identical instructions, inputs, and scoring across 34 capabilities, 55 benchmarks, and nine areas of computer vision, comparing them with dedicated models and humans, and found that semantic interpretation, reasoning, and object-centric prediction approach or reach reference levels while metric geometric accuracy, faithful reconstruction, temporally consistent dense prediction, and specialized fine-grained visual knowledge retain larger gaps.

AI-generated editorial illustration: Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision

Interpretation

The study produces a capability map spanning nine areas: recognition, perception and visual reading; visual reasoning; spatial reasoning; 2D grounding, detection, and segmentation; 3D perception and geometric prediction; video understanding and segmentation; image generation, editing, and restoration; robotics; and expert-domain vision, totaling 34 capabilities and 55 benchmarks. Prior evidence was spread across tasks, domains, and model comparisons, making it hard to see what individual advances collectively mean for computer vision; this work places six frontier general-purpose systems on a common evaluation basis alongside specialist and human references. All six systems receive the same task instructions, visual inputs, and evaluation samples, with outputs scored by the corresponding benchmark metric and the strongest available reasoning setting used where supported; reference scores come from dedicated specialists and human performance per capability.

Semantic capabilities are becoming shared strengths of frontier systems: OCR reaches 93.7–98.1 against human 98.0, all six models surpass reported human performance on document, chart, and infographic understanding by +0.2 to +3.5 points, and all six reach or exceed human 78.7 on mathematical reasoning with scores of 78.8–92.5. These results move the question from whether general models can read images to which capabilities have become a shared baseline across the frontier, and indicate that visual reading and document understanding have passed the human reference. The conclusion rests on convergent performance across multiple systems on the same benchmarks, with score ranges for all six models and human comparison values reported for both OCR and document understanding.

GPT-6 Astra opens margins in structured prediction and spatial reasoning: 96.0 versus human 95.8 in 2D spatial reasoning and 89.6 versus 94.1 in 3D spatial and multiview reasoning; it exceeds the specialist in object detection and segmentation by +10.7 and +4.3 points, leads the next-best frontier model by +26.6 points in video segmentation, +22.7 in pose estimation, +10.4 in image segmentation, and +10.7 in 3D visual grounding. These gains show the advantage extending from image-level structured prediction to dense prediction across video frames, while 2D localization shifts from dedicated models toward a general-purpose interface. Scores are reported on the corresponding benchmark metrics and compared directly with specialist references and the next-best frontier model; Qwen 3.8-Max also surpasses the specialist detector by +8.6 points, indicating the trend is not limited to a single model.

Remaining gaps concentrate where precise geometry, temporal consistency, faithful reconstruction, and fine-grained domain knowledge are required: image restoration clusters all generalists at 17.2–17.7 dB PSNR against the specialist's 30.7, microscopy and pathology scores 14.1–23.8 against specialist 57.1 and human 82.0, and the best frontier model in video segmentation remains 7.9 points below the dedicated model. This separates shared strengths from shared limitations: the uniformly low restoration and pathology scores across models indicate a common frontier rather than an individual system's weakness. Both restoration and pathology report narrow cross-model score ranges alongside specialist and human reference values; qualitative observations show models can localize relevant structures yet struggle to distinguish their fine-grained categories.

Perspective

The work speaks to the computer vision research community and to engineering teams building vision systems, in settings where one must judge whether a visual capability can be handled through a general-purpose interface or needs a specialist tool. It reports capability-maturity tiers on the evaluated benchmarks rather than implying the underlying tasks are solved; references may come from a dedicated specialist, a prior generalist state of the art, or human performance, and the choice of reference affects interpretation, as when generalists surpass the evaluated specialist in 3D visual grounding yet remain substantially below human performance. The study also suggests a division of labor: semantic interpretation, language-conditioned reasoning, and task planning and coordination are natural capabilities to strengthen within the generalist, while metric geometry, faithful reconstruction, dense correspondence, high-fidelity restoration, and fine-grained discrimination in rare expert domains may remain better provided through specialist tools.

A careful reader will still watch how added reasoning and tools close selected gaps but with benefits that vary across capabilities: tool use costs 1.04–1.20× while higher reasoning can reach 13.22× the baseline, sometimes for only small improvements, and segmentation gains plateau at higher effort. When the generalist should act directly, when it should call a specialist, and how it should verify the resulting output therefore remain open questions. The maturity tiers also depend on the chosen reference, and substituting a reference could change how a capability's tier is read. This evidence bundle is full text, but the main results table and some figures are not expanded in the text, so per-capability numbers and the complete tier detail of Figure 9 still require the original.

Sources