Undercount.

Municipal 311 deduplication. Converts how many people complained into how many problems exist. Toronto 311 open data, 2025.

Cities send repair crews where people complain, not where things are broken.

Complaint volume measures civic engagement. Deduplication converts how many people complained into how many problems exist, which is the only fair basis for sending a crew. This is that system, built on Toronto's published 311 data for 2025, and an honest account of what that data can and cannot show.

Three numbers

84.32%

of requests are located only to a postal district of thousands of addresses. The other 15.68% name an actual street corner.

500,269 requests · query location_precision_partition

26.02%

of the requests deduplication can examine sit in a cluster of two or more: the same problem, reported again.

77,833 examinable requests · query dupe_rate_intersection

34.93%

of located pothole reports are duplicates. 5,998 reports describe 4,754 potholes.

Road Pothole / Road Damage · query top_types_by_dupe_rate

What the data would not support

The obvious claim, that deduplication reorders which ward gets served first, is not true in this dataset, and the ward page argues against it rather than for it. Deduplication can only reach the 15.68% of requests that carry a corner, so it removes 2.68% of the file, and the gaps between adjacent wards are larger than that.

Chasing the largest duplicate cluster in the city also turned up a data artifact rather than a street problem. One intersection on Bathurst St carries 1,948 vehicle-noise reports, 14.4% of all redundancy in Toronto, arriving at 43 to 46 second intervals, a signature that appears at exactly two places in 500,269 rows. It is one automated submitter, not 1,948 residents, and excluding it makes every headline number in this project smaller.

Where to look

How to read anything here

Every figure carries its unit and the query that produced it. The complaint text is generated and marked synthetic wherever it appears, because the published dataset contains none. Household income is joined once, after every count is final, and is never an input to the deduplication or to any allocation.