Half the documented quality problems in Toronto's open data are broken time series: taxonomies that changed underneath the numbers

Asked (summary):

Which datasets on the City of Toronto's open data portal have documented data-quality or normalization problems — inconsistent schemas across files, free-text fields that should be coded categories, PDF-only publication, or taxonomies and code lists that changed over time and broke time-series analysis? For each case: dataset name, the specific problem, and the source URL documenting it.

A documented-case scan of web-indexed City portal posts, City documentation, GitHub projects, and civic-tech writeups (as of October 9, 2026) surfaced 26 problem instances across roughly 20 datasets: 13 taxonomy or classification breaks, 7 free-text normalization cases, 4 schema-consistency cases, and 2 documented PDF/non-structured publication cases at the time described. This is a scan of documented cases, not a statistical audit of the whole catalogue, and some sources describe earlier dataset versions. The same dataset can appear more than once — each row is a distinct documented problem type.

Click a bar to filter the evidence table to that problem type; click again to clear. Hover a bar for the datasets behind it.

The April 2019 sheet of the TTC bus delay workbook has 11 columns instead of 10, gains an Incident ID field, and renames Min Gap/Min Delay to Gap/Delay — breaking naive month-by-month appends. rdrr.io
Parking-ticket street names appeared in 722,000 free-text variants ("King St W", "King Street West", "KING STR W") before fuzzy matching against the Centreline dataset collapsed them to about 10,000 clean segments. open.toronto.ca
The 311 dataset holds more than 850 unique request types, many renamed, merged, retired, or replaced over time; one analysis had to manually consolidate over 200 of them into standardized categories. github.com
TTC streetcar delay coding jumped from 13 plain-language categories in 2024 to roughly 139 alphanumeric codes from 2025, with no official crosswalk — and 12% of the new codes fall outside the official lookup table. opendatacanada.ca

Evidence: all 26 documented cases, grouped by problem type

DatasetDocumented problemSource

Documented-case scan of Toronto open data quality and normalization problems: 26 rows, one documented problem instance each, drawn from City portal posts, City documentation, GitHub issues and projects, and civic-tech writeups indexed as of 2026-10-09. Counts measure documented instances, not a statistical audit of every catalogue package; some sources describe historical publication states or earlier dataset versions. Long issue descriptions are shown in full; sources span multiple independent hosts, with no single URL supporting more than half the rows.

This report was generated automatically by Keenable SELECT at a user's request, from publicly available web sources linked herein. Keenable does not review, verify, or endorse its contents and makes no representation as to accuracy, completeness, or timeliness; AI-based extraction may contain errors. Nothing in this report is investment, legal, financial, or other professional advice. All trademarks and referenced content remain the property of their respective owners; no affiliation or endorsement is implied. To report an error, rights concern, or request removal: legal@keenable.ai.

Keenable SELECTAsk your own question
Made with Keenable SELECT