Find and rank the 50 most valuable use cases where an AI agent needs visual evidence — a chart, table, map, diagram, or screenshot — to complete a task reliably. For each, report the request, the agent's job, the required visual, why text alone is insufficient, the corpus, observed demand, repeat frequency, error consequence, and likely customer. Rank by observed demand 30%, visual necessity 30%, business impact 20%, repeatability 10%, corpus feasibility 10%. Exclude decorative visuals and text-solvable cases; separate observed evidence from inference; cite everything; finish with the top 10 the Evidence API should support first.
A broad scan — 64 search queries in four batches over 12,321 deduplicated pages — surfaced 220 candidate tasks; a strict screen kept 36 concrete recurring use cases with visual-necessity of at least 5/10. This is a ranked 36-of-requested-50 result: 14 further candidates could not be responsibly promoted to force the total to 50. Each score is exactly 10 × (0.30 demand + 0.30 visual necessity + 0.20 impact + 0.10 repeatability + 0.10 corpus feasibility), components 0–10; sources span 13 independent hosts and no single URL backs more than half the rows.
Common primitives implied by these ten rows (labeled inference): evidence capture and fetch (screenshots, page images, frames); immutable timestamps and versioning so agents never act on stale screens; crop and region references for grounding a claim to pixels; OCR plus spatial and layout grounding; chart, table, and diagram reasoning; screenshot and UI-state verification of outcomes; provenance citations back to the source visual; freshness checks; calibrated confidence; and human-review escalation when confidence is low. Each primitive traces to failure modes named in the rows — stale-screen false success, hallucinated fields, repeat-action loops — not to invented requirements.
| # | Use case | Score | Dem | Vis | Imp | Rep | Cor | Details | Source |
|---|
Screen. From 220 screened candidates, 36 were included; 184 were excluded for benchmark-only demand without a recurring customer job, visual necessity below 5/10, duplication of a stronger normalized use case, or decorative rather than load-bearing visuals. Excluded rows were never promoted to reach the requested 50.
Pixel-essential versus API-solvable. The strongest cases are pixel-essential: unlabeled charts, canvas-drawn widgets, video frames, satellite imagery, and GUI state exist only as pixels. Weaker cases — some form and invoice extraction — could partly be solved by structured APIs or text extraction where clean source data exists; their visual-necessity scores of 7–8 rather than 9–10 reflect that.
Corpus acquisition. Feasibility ranges from 9 (public chart benchmarks, Vega-Lite galleries) down to 5–6 where the corpus is customer-owned production imagery, medical scans, or licensed broadcast footage. Blank fields in the table mean the source did not resolve the field, not that the answer is "none".
Limitations. Scores blend observed evidence with explicitly labeled inference for frequency, customer, and consequence; inference is kept under its "Inference:" label throughout. Evidence spans benchmarks, papers, GitHub and public workflow reports, documentation, and professional use cases, and skews toward what is published — production demand without public traces is undercounted.
Method: 36 definitive use cases ranked from a 220-candidate screen of 12,321 deduplicated pages found via 64 search queries; scores are 0–100 weighted composites of five 0–10 components (demand 30%, visual necessity 30%, impact 20%, repeatability 10%, feasibility 10%). Long evidence text is truncated in expandable rows for space; blank fields are unresolved. Compiled for Keenable SELECT.