Pelicans on bicycles get better — but even in 2026 top labs still fall off the bike

Asked:

“Show me a timeline of models being evaluated by their ability of drawing a pelican riding a bike”

25 dated model evaluations, November 2024 to July 2026, mostly from Simon Willison’s published pelican-benchmark posts. The recurring canonical task is “Generate an SVG of a pelican riding a bicycle”; variants cover video, image/vector generation, agentic iteration, a POV-Ray edition and one controlled study. Tiers (Excellent, Strong, Mixed, Poor, Not rated) code the evaluator’s own wording — they are not an official numeric benchmark.

ExcellentStrongMixedPoorNot rated — tier colour codes the evaluator’s wording; flagged entries are milestones
🚲 Gemini 3 Deep Think was called “the best one I’ve seen so far” on 12 Feb 2026 — the same day GPT-5.3-Codex-Spark drew an orange duck merged with a bicycle. simonwillison.net
🐦 Open and local models caught up: GLM-5.1 became the favourite open-weights result on 7 Apr 2026, and Qwen3.6-27B was judged outstanding for a 16.8 GB local model two weeks later. simonwillison.net
📉 Progress is not monotonic: Poor results still appeared late — Qwen3-30B-A3B in Jul 2025, GPT-5.1 in Nov 2025, Codex-Spark and Opus in Feb 2026. simonwillison.net
🔬 A controlled study on 22 Jul 2026 found no evidence that labs optimize specifically for the pelican-bicycle prompt: pelicans aren’t drawn better than other animals, nor bicycles than other vehicles. simonwillison.net

All 25 evaluations

DateModelTierModeAssessmentSource

Data: 25 dated model evaluations of the “pelican riding a bicycle” task, 2024-11-15 to 2026-07-22, one row per model/date, drawn mainly from Simon Willison’s weblog — a deliberately unscientific benchmark by its author’s own description. Ability tiers are a visualization-friendly coding of the cited evaluator wording, not an official score. Later reporting notes perceived general capability correlates with better drawings, but the 2026 “pelicanmaxxing” controlled study found the combination was not unusually optimized versus other animals and vehicles. Assessments in the table are lightly truncated for space.

This report was generated automatically by Keenable SELECT at a user's request, from publicly available web sources linked herein. Keenable does not review, verify, or endorse its contents and makes no representation as to accuracy, completeness, or timeliness; AI-based extraction may contain errors. Nothing in this report is investment, legal, financial, or other professional advice. All trademarks and referenced content remain the property of their respective owners; no affiliation or endorsement is implied. To report an error, rights concern, or request removal: legal@keenable.ai.

Made withKeenable SELECT · 2,742 pages in 4m 57s · Ask your own questionShare:XLinkedInReddit
Made with Keenable SELECT