415.tech
AI & tech, from the frontlines of Silicon Valley
Roboflow benchmarks GPT-5.6 Sol at 46.2 mAP, triple GPT-5.5's detection score

Roboflow benchmarks GPT-5.6 Sol at 46.2 mAP, triple GPT-5.5's detection score

Roboflow's VLM benchmark scores OpenAI's GPT-5.6 Sol at 46.2 mAP@50 on object detection against GPT-5.5's 13.8, and 73.0% on counting against 64.9%. The gain is uneven — OCR and text extraction slipped to 90.7% and 82.5% — but detection moves from an OpenAI weak spot to a capability document workflows and screen-understanding agents can actually use. Gemini 3.5 Flash still leads both tasks at 0.8 cents per image against Sol's 2.5, so the economics of high-volume vision work have not shifted.

Source: blog.roboflow.com

Post on XEmail

Detection moved from a weak point to a usable capability, and counting improved across the full model family.

Roboflow

Why this matters

  • → Object detection jumped 3.3x, unblocking document workflows and UI agents that depend on visual understanding.
  • → Sol remains 3x more expensive than Gemini 3.5 Flash, limiting adoption for high-volume vision work.
  • → OpenAI finally closes a long-standing weakness in computer vision, reshaping VLM capability rankings.
OpenAI's vision pivot