How we test this data
We strive to make the data as accurate as possible. But accuracy is a moving target and the accuracy of data improves with use. Here we explain the process we have used to reach a publishable threshold of accuracy. As we improve the data and add new methods to detect errors they will be documented here.
This dataset has 46,148,034 rows and connects observations across 56 years. That is 56 sets of potential changes to collection procedure, measurements, and variable alignment. Our team of two at Civilytics cannot hand-verify a dataset of this extent. Instead, we have built a release threshold — a set of checks that must pass before the data are released for use. Those checks are described here so you can judge whether it is strict enough for what you are about to use it for.
Verification at this scale is an economics problem, not a checklist. Each additional round of checking costs more time, energy, and tokens and finds fewer clear errors and more and more judgment calls. Somewhere on that curve you either publish the data or you never do. Our philosophy is to do so transparently and put trust in the hands of the users to surface issues and report back so that together, the data continues to improve with use over time. Together we are building a more accurate fiscal history of states, counties, and cities on the foundation provided by the Census Bureau. To do that, you need to know what has been tested in the data and how.
Below are the checks we used.
The rules are built into the publication process
The data is created via a reproducible data pipeline. Publication only occurs if every blocking check passes. Any failed check stops the build before anything is written.
When a check does fail, the rule is that you fix the data or you catalogue the reason with a citation.
The four types of checks
1. Reconciliation against the Census Bureau’s own totals
We add our figures up and compare them to what the Census Bureau published for the same year and state. If our total for a state and government type differs from theirs beyond a set tolerance, the build stops. This is the strongest check.
There is a related check that reconciles each modern year’s assembled data against the raw source file it came from — row count, dollar total, and count of missing values — which catches the class of bug where a join silently duplicates or drops rows.
What it caught: the state-code defect (E-002/E-003 in the register). The Census Bureau used its own alphabetical state codes before FY2017 and standard FIPS codes afterward, and we read both as FIPS. For five years each state’s published figures were compared against a different state’s figures in the dataset.
What it cannot catch: anything wrong in both places at once. If a government misreported to the Census Bureau, we reconcile perfectly against their number and are perfectly wrong. Reconciliation proves we transcribed and aggregated faithfully. It cannot prove the underlying report was true.
2. Cross-checking independent paths to the same number
Many figures can be reached more than one way — a Census aggregate should equal the sum of the detail lines beneath it; a total should equal direct spending plus intergovernmental payments; a year’s data assembled in wide form should sum to the same thing as the same year in long form. When two routes to one number disagree beyond tolerance, the build stops.
What it caught: errors introduced when we combine the Census Bureau’s account codes into consistent categories across years. In one case, regrouping moved dollars between two categories that should never exchange them — a highway dollar cannot become a police dollar. In another, reshaping a year’s data from one table layout to another silently duplicated rows. Neither shows up when you look at one view of the data. Both are obvious the moment two views disagree.
What it cannot catch: an error that moves both paths together. If an account code is misclassified at the source — spending recorded as health rather than waste treatment — the aggregate and its parts still agree. They are just both filed under the wrong heading. This class is also where our worst known coverage gap is; see below.
3. Coverage accounting
Checks that the right things are present, rather than that the present things add up: does each year contain roughly the expected number of governments of each type, does every geographic code exist and match the identifier it is attached to, does a category that existed last year still exist this year.
What it caught: seven missing years of hospital spending (E-001) surfaced here — roughly $70–90 billion a year, more than half the Health & Hospitals category in each of those years, absent because a group of codes stopped being carried.
What it cannot catch: a government that never reported at all. Coverage accounting compares what arrived against what we expected to arrive, and a government missing from the source in the first place is invisible to it. Four of every five years is a sample, and the sample varies: of Wisconsin’s 608 cities, 597 reported in FY2012 and 112 in FY2019. This check will not flag it — it is a property of the data you have to reason about yourself.
4. Series-break detection
The hardest part of a series that runs from 1967 to 2024 is that the definitions move. Categories split, codes change meaning, whole families of codes stop being collected. These breaks must be documented so users have a choice in how they handle them. There is no single right answer to resolve moving categories or definitions, it is an analytical choice, so we try to give users the information they need to make the right choice.
We catalogue every known boundary between Census category definitions — currently 326 catalogued breaks — and check for uncatalogued ones by looking for year-over-year jumps too large to be real. A jump that is a genuine event gets catalogued with its citation. A jump that is not gets fixed.
What it caught: the FY2022 disappearance of an entire code family (E-006). This was a real change; the Census Bureau stopped collecting this code family.
What it cannot catch: a definition change too small to produce a visible jump. A category that quietly absorbs a slightly different set of activities, with no step in the dollars, passes everything we have. This is the failure mode we are least able to detect and the one that impacts trend analysis at the local level the most.
Threshold checks we are still working on
We calibrate thresholds for two additional types of error: a single government’s spending jumping implausibly year over year and state-level swings that cancel out in national totals. The variance in these is calibrated against the FY1967–FY2023 series.
One check is still report-only: a comparison against the state-by-type totals the Census Bureau publishes directly. It is the cheapest strong check available to us, precisely because the Bureau publishes the answer, and it is not yet blocking.
Work in progress
We are still working to improve our checks and this page will be updated to keep track of that progress. Right now, our check to verify against Census aggregates covers only 42 of 107 groups of Census codes. For the other 65 it runs and finds nothing, because there is nothing there for it to compare.
Each of those 65 now carries a written reason, and the build fails if a group is neither checked nor accounted for: 36 are deferred with a reason on record, 12 have no underlying detail to compare against, one would be circular, and 16 have no decision recorded yet. The coverage figure is written out on every build.
Closing out these 65 requires additional research rather than configuration and is on the roadmap for future work.
What none of this catches, and where you come in
Every check above tests internal consistency and faithfulness to the Census source. None of them can tell you whether what was reported to the Census was right.
If a county’s audited financial statement says one thing and this dataset says another, no check we run will ever surface it. Only somebody who has both documents can — and that person is almost never us.
That is the whole argument for the issue register, and it is why the commons stage matters. It is not a community feature bolted onto a dataset. At this scale it is the only verification strategy that scales at all.
If a figure here disagrees with your government’s audited financials, tell us — email data@civilytics.com, with the government, the year, and a link to the document. We’ll investigate, log the report in the project data pipeline register, and we’ll credit your contribution.