“Show me a timeline of models being evaluated by their ability of drawing a pelican riding a bike”
25 dated model evaluations, November 2024 to July 2026, mostly from Simon Willison’s published pelican-benchmark posts. The recurring canonical task is “Generate an SVG of a pelican riding a bicycle”; variants cover video, image/vector generation, agentic iteration, a POV-Ray edition and one controlled study. Tiers (Excellent, Strong, Mixed, Poor, Not rated) code the evaluator’s own wording — they are not an official numeric benchmark.
| Date | Model | Tier | Mode | Assessment | Source |
|---|
Data: 25 dated model evaluations of the “pelican riding a bicycle” task, 2024-11-15 to 2026-07-22, one row per model/date, drawn mainly from Simon Willison’s weblog — a deliberately unscientific benchmark by its author’s own description. Ability tiers are a visualization-friendly coding of the cited evaluator wording, not an official score. Later reporting notes perceived general capability correlates with better drawings, but the 2026 “pelicanmaxxing” controlled study found the combination was not unusually optimized versus other animals and vehicles. Assessments in the table are lightly truncated for space.