Which datasets on the City of Toronto's open data portal have documented data-quality or normalization problems — inconsistent schemas across files, free-text fields that should be coded categories, PDF-only publication, or taxonomies and code lists that changed over time and broke time-series analysis? For each case: dataset name, the specific problem, and the source URL documenting it.
A documented-case scan of web-indexed City portal posts, City documentation, GitHub projects, and civic-tech writeups (as of October 9, 2026) surfaced 26 problem instances across roughly 20 datasets: 13 taxonomy or classification breaks, 7 free-text normalization cases, 4 schema-consistency cases, and 2 documented PDF/non-structured publication cases at the time described. This is a scan of documented cases, not a statistical audit of the whole catalogue, and some sources describe earlier dataset versions. The same dataset can appear more than once — each row is a distinct documented problem type.
Click a bar to filter the evidence table to that problem type; click again to clear. Hover a bar for the datasets behind it.
| Dataset | Documented problem | Source |
|---|
Documented-case scan of Toronto open data quality and normalization problems: 26 rows, one documented problem instance each, drawn from City portal posts, City documentation, GitHub issues and projects, and civic-tech writeups indexed as of 2026-10-09. Counts measure documented instances, not a statistical audit of every catalogue package; some sources describe historical publication states or earlier dataset versions. Long issue descriptions are shown in full; sources span multiple independent hosts, with no single URL supporting more than half the rows.