Changelog — US Census of Governments Finance API
Notable changes to the public API surface at
https://cog-api.civilytics.org, to the data it serves, and
to the documents published at https://www.civilytics.org/cog/. Newest first. Entries
marked BREAKING change numbers or status codes a
working integration already depends on.
2026-10-05 — The Southern guide and the API specification, republished
No API change. Both documents are rebuilt against
the current corpus (pipeline b58f729).
- The Southern guide reflects the reissued Census files below. Most of its figures moved by less than half a percent. Per person in FY2024, the two states with the lowest general revenue are now Florida and New Hampshire, where they were Florida and Georgia. Its four opening tables now name the corpus commit they were built from.
- The API
specification had last been built against an August corpus and
listed 18 endpoints. It now lists all 19, including
GET /api/v1/examples, and listscomplete=andformat=among each route’s parameters. - This changelog is now published at https://www.civilytics.org/cog/changelog/, and its three newest entries appear on the landing page.
2026-09-30 — The Census Bureau’s reissued files, and fund balances back to FY1967
BREAKING (numbers). Values for FY2017–FY2022 changed, FY2023 and FY2024 lost three items, and fund balances now reach back to FY1967. No route, parameter or status code changed.
The corpus (pipeline b58f729) now takes every year from
the newest file the Census Bureau posts. The Bureau reissued
FY2017–FY2021 in October 2024 and FY2022 in August 2026, and re-posted
FY2023–FY2024 on 4 August 2026. This API had been serving the files
those replaced.
- FY2017–FY2022. National totals moved by at most 0.5% in any one year. Individual governments moved far more: Chicago’s FY2022 property tax went from $1.57 billion to $3.35 billion.
- FY2023 and FY2024. Every figure the old and new files share is identical. The re-post drops fund balances, total salaries and wages, and sports-betting tax, so those two years no longer carry them.
- Fund balances (sinking, bond and other funds) now
run from FY1967 to FY2021, where they started at FY2012, so
/governments/{govid}/balancesreturns years it did not before. The Bureau’s files carry no fund balances after FY2021. - The full accounts are errata E-012 (the reissue, with more examples and what to do if you downloaded FY2023 or FY2024 fund balances before this release), E-013 (fund balances) and E-014 (sports-betting tax).
provenance.manifest.pipeline_commiton any data response names the corpus behind it.
2026-08-13
— complete=, and a numeric NA that was
serializing as the string "NA"
Added —
complete=true recovers wide-era census zeros
(cog-api#22)
On /spending, /revenue,
/governments/{govid}/spending and
/governments/{govid}/revenue. Default false;
nothing changes unless you ask for it.
By default a category a government did not report in a year simply
has no row. With complete=true the requested (category ×
year) grid is filled from the corpus code_set, and every
row carries value_source:
value_source |
meaning | amt_nominal |
|---|---|---|
reported |
the corpus carries this cell | as published |
census_zero |
dense-source year (≤ FY2011), absent — Census published $0 | 0 |
not_reported |
sparse-source year (≥ FY2012), absent — unknown | null |
A not_reported amount is null and
must not be read as zero. Filling a modern absence with
0 turns “we don’t know” into “they spent nothing” — a
different and much stronger claim. Only the pre-FY2012 era supports the
zero reading, because the source was dense there.
This matters more over HTTP than in R: an R user who gets an empty
result can inspect cog_explain(); a BI consumer gets an
empty cell and reads it as zero.
meta.totalcounts the filled rows, andmeta.completeechoes the parameter so a response states which contract its count is under.provenance.completionreportsapplied,rows_filledandabsence_means— the last keyed by year, because one request can span the FY2011/FY2012 seam.- Refused with a 400, never silently ignored,
alongside
recipeorexpenditure_concept=total(neither has a grid to fill), on/profile(it stacks spending and revenue into one table, so onevalue_sourcecolumn cannot describe both grids), and as an unknown parameter on/balances,/rollupsand/peer-comparison.
Fixed
— a numeric NA serialized as the string "NA"
on every endpoint
Pre-existing and live, not introduced by the above.
jsonlite::toJSON() renders a numeric NA as
"NA" — a JSON string — and the serializers set
null="null" but not na="null". Character and
logical NAs already came out as null, which is
why only the numeric case was wrong and why it went unnoticed. Measured
on ordinary requests with no new parameter set:
| endpoint | leaked |
|---|---|
/governments/{govid}/peer-comparison |
"target_rank":"NA" ×22, "amt_nominal":"NA"
×6 |
/governments/{govid}/spending, /spending,
/rollups, /balances |
"base_year":"NA" |
The harm is quiet and worst for R and BI consumers: one
"NA" string turns an entire amount column
character, so arithmetic silently stops working rather than
failing loudly. It also made complete=true’s central
contract impossible to honour — a string is not a null.
Those fields now serialize as null. If you were
string-matching "NA" to detect a missing value, match
null instead; if you were parsing these columns as text to
work around the coercion, you can stop.
2026-08-13
— categories_missing: a category that stops early is now
named (cog-api#56)
Additive. No existing field changes; three new keys
appear under provenance.scope on the per-government money
routes.
provenance.scope.govids_missing already answered “which
of my units contributed nothing?”. There was no
equivalent for categories, so a government whose
intergovernmental aid stops two years before its property tax returned
the surviving rows and said nothing at all.
That is quietly wrong rather than merely incomplete: a per-category “% change across the window” then measures a different window for each category. Spokane City, FY2020–FY2023 — property tax, sales tax and current charges run to FY2023; IG Federal and IG State stop at FY2021; IG Local at FY2017.
Added,
on /governments/{govid}/spending, /revenue,
/balances and /profile
| field | meaning |
|---|---|
years_present |
the years this result actually returned rows for |
categories_found |
every category that appeared at least once |
categories_missing |
[{category, years}] — present in some of
years_present, absent in others |
categories_missing is an empty list rather than an
absent key when nothing is missing, so it can be read unconditionally.
years is always a JSON array, even for a one-year gap.
The
reference set is years_present, not the years
requested
If a government reported nothing at all in a requested year, that is
a fact about the year, not about any category — diffing against
requested years would mark every category missing in it and bury the
real signal. A category absent throughout never reaches
categories_found, so this cannot flag a category a
government legitimately does not run.
The fleet routes deliberately do not carry it
“IG State is missing in 2022” has no single truth value across a
response spanning 20,000 governments, and the field shape carries no
govid to disambiguate. Those routes also push
limit/offset into SQL, so the handler holds
one page rather than the full result. Emitting nothing is truthful;
emitting an empty list there would read as “no gaps”.
Also corrected
data-dictionary.md described govids_missing
as a peer-comparison-only field. It is not, and never
was — every route taking an explicit govid set carries it, including the
flat /spending and /revenue with a
govid= list, where it is the field that separates a stale
id from a government that genuinely reported nothing.
2026-08-13
— A govid that cannot be an id is now a 400
(cog-api#55)
BREAKING. govid= was checked only for
non-emptiness, so any value — including a full URL — passed through,
matched nothing, and returned status: "success" with
data: []:
GET /api/v1/revenue?govid=https://example.com/foo&years=2022
-> {"status":"success","data":[],"error":null,...}
This API promises that empty-plus-success means “the government
reported nothing” — a real and common condition in this corpus.
That promise only holds while govid is well-formed. Until
now a typo and a genuine reporting gap arrived in the identical shape,
which for a finance API is the worst available confusion: the caller
writes up “no reported revenue” when they actually mis-pasted an id.
A canonical_govid is digits only, exactly 12 of
them. Anything else — a URL, a name, a truncated or FIPS-style
id — is now rejected with a 400 naming the expected format and pointing
at /api/v1/governments?q= to resolve one.
The two failure modes stay distinct
| Input | Before | After |
|---|---|---|
Cannot be an id (https://…, not-an-id,
482032010) |
200, data: [] |
400 naming the format |
Well-formed but not in the corpus (999999999999), flat
routes |
200 + govids_missing |
unchanged |
Well-formed but not in the corpus,
/governments/{govid}/… |
400 Unknown govid |
unchanged |
A flat govid= list of ten ids must not fail because one
has gone stale, so only input that could never be an id is refused. That
distinction is the point of the change, not a side effect of it.
The accepted width is read from the corpus, not hardcoded
The issue hedged on whether 12-digit-numeric holds for every vintage
(FY2012–FY2016 GOVS ids are 14 characters). The accepted widths are
derived from the mounted registry at boot rather than pinned to 12, so
the check cannot lock out a vintage the corpus actually carries, and
cannot go stale if the corpus gains another width later. Against the
deployed corpus this resolves to {12}.
2026-08-13 — The wrong-flow category gate now covers every money route (cog-api#41)
BREAKING. The 2026-08-05 entry below gave
/rollups a 400 for a category whose
category_type it cannot serve. It left the flat and
per-government money routes alone, so the same class of input got a 400
on one endpoint and a silent empty 200 on another, with
nothing in either response to predict which. That asymmetry was
documented honestly at the time and tracked as a follow-up; this closes
it.
BREAKING
— a wrong-flow category is now a 400 on
/spending, /revenue and
/balances
Previously 200 with data: []:
GET /api/v1/spending?state=WI&years=2012&category=Property%20Tax
GET /api/v1/spending?years=2022&category=Fund%20Balances
GET /api/v1/revenue?years=2022&category=Police
GET /api/v1/governments/{govid}/spending?category=Property%20Tax
GET /api/v1/governments/{govid}/revenue?category=Police
GET /api/v1/governments/{govid}/balances?category=Police
A revenue category is a real /categories value, so it
resolved fine and then matched nothing — and
/spending?state=WI&category=Property%20Tax reads as “no
Wisconsin city reported property tax in 2012.” False, and not obviously
absurd enough for a caller to catch.
Remedy: send the category to the route that serves
its category_type. The error names it, and names the
route:
category 'Property Tax' is a revenue category; /api/v1/spending serves expenditure categories only. Use /api/v1/revenue for a revenue slice, or see /api/v1/categories for each category's category_type.
NOT affected
category=All Categories— the reserved pseudo-category is exempt and keeps working on every money route, both flows.cog_categories()emits it once per flow, so itscategory_typesays nothing about which route can serve it. (cog-api#54 was filed believing this was already broken; it never was — that report’s reproduction used a govid that is not in the corpus.)/profile— deliberately ungated. It blends spending and revenue into one table (see itsflowcolumn), so both flows are legitimate there.- An omitted
category— unchanged on every route. /rollups— same behaviour as before, now via the shared helper rather than its own inline copy.
2026-08-13 — The documented subtype vocabulary was missing six real values (cog-api#44)
Docs only. No behaviour changes; the API always accepted these.
/api/v1/data-dictionary listed spending subtypes as
operations, capital,
intergovernmental and revenue as own_source,
federal, state, local_aid.
Measured against the corpus, six more exist and all six work as
subtype= filters:
| flow | undocumented |
|---|---|
| spending | assistance, interest,
insurance_benefits |
| revenue | utility, liquor_store,
insurance_trust |
A list that omits real values asserts completeness by omission, so a
client reading this document as the authority — which is its purpose —
concluded those rows could not exist and then met them in ordinary
responses with no documented meaning. assistance is
the worst case: it appears under the default expenditure
concept, so it shows up in ordinary /spending
output with no parameter set. It is public-assistance payments to
individuals (the Census J objects), carried mainly by
states and counties.
The enumeration now says which concept reveals each value
A bare list was not enough, and its absence was its own trap: filter
subtype=interest under the default concept, get zero rows,
conclude the government carries no debt service.
| appears under | spending subtypes |
|---|---|
primary (default), direct,
total |
operations, capital,
assistance |
direct, total only |
interest, insurance_benefits |
total only |
intergovernmental |
| appears under | revenue subtypes |
|---|---|
general (default), total |
own_source, federal, state,
local_aid |
revenue_concept=total only |
utility, liquor_store,
insurance_trust |
Guarded against recurrence
The new tests derive the expected vocabulary from
.known_subtypes() — the same corpus-backed source
resolve_subtype() validates against — rather than restating
it. The docs can no longer drift from the API without the suite going
red. Re-listing the values by hand would only have moved the stale copy
into another file.
2026-08-13
— GET /api/v1/examples: worked URLs, one per endpoint
(cog-api#57)
Additive. A new read-only route; nothing existing changes.
Agent fetch tools are commonly permission-gated to URLs that have appeared verbatim in the conversation — a documented pattern is not enough, the literal string has to exist somewhere first. So an agent handed a reference table of route shapes still cannot move, and asks the user to paste URLs one at a time. That is what actually happened in the report behind this issue.
GET /api/v1/examples
{"status":"success","data":[
{"endpoint":"/api/v1/revenue",
"description":"Flat revenue slice for an explicit govid list",
"path":"/api/v1/revenue?govid=532063205599&years=2024",
"example_url":"https://cog-api.civilytics.org/api/v1/revenue?govid=532063205599&years=2024"},
...]}
One request bootstraps the whole API. example_url is
absolute — the point of the route — and path is the same
thing root-relative, matching the convention GET /api/v1/
uses for its own documentation links (cog-api#13). Both are emitted so a
caller never has to join strings and risk the
/api/v1/openapi.json 404 that convention exists to
prevent.
The examples cannot rot
govid and every year are read from the mounted
corpus at request time, not hardcoded. A corpus republish that
moves the year window updates the examples with it, instead of turning
them into a list of 400s. The absolute URL is built from the requesting
host, so it is correct behind the production TLS proxy
(X-Forwarded-Proto is honoured) and equally correct
anywhere else.
Two lists, one source of truth
llms.txt’s CANONICAL EXAMPLES block carries
the same list as prose. The issue asked for one source of truth “rather
than maintaining two lists that will drift”; the test suite asserts the
set of endpoints covered by each is identical in both
directions, so drift either way is a red build. The suite also
fetches every advertised URL against a booted server and
requires 200, so a route rename cannot leave a
dead example behind — the previous block was verified by hand, which
does not scale.
2026-08-09 —
format=csv: up to 50,000 rows in one request
Additive. format defaults to
json, and the JSON body is byte-identical with or without
format=json.
/spending and /revenue accept
format=csv, which returns the rows themselves, header
first, and raises the row cap from 1,000 to 50,000 in one request. The
counts the JSON envelope carries move to response headers:
X-Total-Count, X-Limit and
X-Page. Provenance is not included; use JSON if you need
it.
GET /api/v1/spending?state=WI&type=city&years=2022&format=csv&limit=50000
Every California city’s FY2022 spending (6,857 rows) went from seven requests and 2.00 s to one request and 0.32 s.
- The raised cap applies only when
subtype=,recipe=andcomplete=true(from 2026-08-13) are all unset. With any of them set, an over-caplimitis a 400, never a silent clamp. - Accepted on the flat money routes only.
Access-Control-Expose-Headerslists theX-*headers, so a browser client can read them.- For the whole corpus, read the public parquet mirror rather than paging this API; the data dictionary has runnable snippets.
2026-08-05
— Two status-code breaks: over-cap limit and
/rollups category rejection
BREAKING. Both changes reject a request that
previously answered 200 in silence — no error, no response
flag, and (until now) no changelog entry, even though this file’s own
bar (above) says a status-code change on a working integration is
exactly what belongs here.
BREAKING
— limit above 1000 is now a 400, not a silent clamp
(cog-api#39)
Any request whose limit= exceeded the cap (1000) used to
succeed anyway, silently clamped to 1000 rows under
status: "success" — a limit=5000 caller got
1000 rows back with nothing in the response indicating 4000 were
dropped. It is now rejected outright with a 400 naming
the cap.
Remedy: keep limit at or below 1000 and
walk further pages with page=. meta.total
always reports the full unpaginated count, so the page count is
ceil(meta.total / meta.limit).
BREAKING
— /rollups?category=<a revenue or balance category>
is now a 400, not an empty 200 (cog-api#35)
/rollups wraps cog_geographic_rollup(),
which is expenditure-only. A revenue or balance category
(e.g. Property Tax, Fund Balances) is a real,
valid /categories value, so it used to resolve fine and
then silently match zero rows — a status: "success"
response with data: [] that reads exactly like “this layer
reported nothing,” not “you asked for a category this route can’t
serve.” It is now rejected with a 400 naming the
category’s actual category_type.
Remedy: use /api/v1/revenue (with
state=/type=) for a revenue slice, or
/api/v1/governments/{govid}/balances for holdings. See
/api/v1/categories for each category’s
category_type.
2026-07-31 — Census concept vocabulary for expenditure and revenue
Redeploys the API onto uscogdata’s classification model
(uscogdata#11,
uscogdata#12).
Closes cog-api#10 and cog-api#28.
Read this section if you consume /revenue or
pass expenditure_concept. Two figures change.
Neither raises an error, and neither is flagged in the response body —
the disclosure is here and in /api/v1/data-dictionary.
BREAKING — default revenue drops for governments that run utilities
revenue_concept is now a parameter, defaulting to
general — strict Census General Revenue.
Utility, liquor store and insurance trust revenue have left the
default. Previously the API returned all of them.
| government type | utility + liquor store share of the previous default |
|---|---|
| City | 15.9% |
| Township | 4.8% |
| County | 1.7% |
| State | 1.2% |
Madison, WI FY2020 — a typical city:
| request | figure |
|---|---|
| before this release (default) | $616,551,000 |
after this release (default,
general) |
$564,322,000 |
?revenue_concept=total (after) |
$616,551,000 |
To keep your previous numbers, add
revenue_concept=total. That reproduces the old
default exactly — it is the same set of subtypes, now named.
Why the default moved rather than staying put: the API is a thin
layer over the uscogdata R package, and the two surfaces
are meant to return the same figure for the same question.
general is the reader’s default and Census’s standard
headline concept for cross-government comparison.
BREAKING
— expenditure_concept=direct returns a larger figure
The expenditure vocabulary is now three nested concepts:
| concept | includes | excludes |
|---|---|---|
primary (new default) |
Running services: salaries, supplies, construction, equipment | Interest on long-term debt; payments to other governments |
direct |
primary + interest on long-term debt
(I89) — Census’s published Direct
Expenditure |
Payments to other governments |
total |
direct + payments to other governments
(Direct + M + L) |
— |
If you did not pass expenditure_concept, your
spending figures do not change. The new primary
default returns exactly what the old direct default
returned.
If you passed expenditure_concept=direct
explicitly, your figure grows by that government’s interest on
long-term debt — same parameter value, larger number. Madison, WI
FY2020:
| request | before | after |
|---|---|---|
| default | $623,347,000 | $623,347,000 (unchanged) |
?expenditure_concept=direct |
$623,347,000 | $651,051,000 |
The difference, $27,704,000 (4.4%), is item code I89,
general interest on debt. To keep your previous numbers, use
expenditure_concept=primary (or drop the
parameter).
This resolves the defect in cog-api#10: the API’s
direct was ~7% below Census’s own published “Direct
Expenditure” for the same government and year, under a parameter of that
name, with nothing on the surface disclosing it. direct now
means what Census means by it.
Added
revenue_concept=general|totalon/api/v1/revenue,/api/v1/governments/{govid}/revenue, and the revenue rows of/api/v1/governments/{govid}/profile. Reported back inmeta.revenue_concept.expenditure_concept=primary, the new default, alongsidedirectandtotal./api/v1/data-dictionarygains “Whichrevenue_conceptdo I want?”, the three-concept expenditure table, and worked dollar examples for both.
Changed
- Revenue
totalcarries no route restriction, unlike expendituretotal(which remains available only on/api/v1/governments/{govid}/spending). Expendituretotaladds money handed to other governments, which those governments report again as their own spending, so summing it across governments double-counts. Revenuetotaladds the government’s own utility, liquor store and insurance trust receipts, which nobody else reports — so it is safe to aggregate. expenditure_concepton a/revenueroute still returns 400, but the message now points atrevenue_conceptinstead of claiming revenue has no concept distinction — which stopped being true withuscogdata#12.revenue_concepton a spending route returns 400 rather than being silently ignored (the mirror of the rule above)./rollups,/peer-comparisonand/profilenow forwardexpenditure_conceptto the underlying verb instead of validating and discarding it. Previouslydirectwas the verb’s own default, so discarding it changed nothing; nowdirectandprimarydiffer, so those endpoints would have returnedprimaryfigures under an explicitdirectrequest.
Known series seam
The employee-retirement codes inside insurance trust revenue
(X01/X02/X05/X08)
stop after FY2016, when those systems moved out of the
annual finance collection. A long revenue_concept=total
series therefore steps down at FY2016/FY2017 for
collection-scope reasons, not because revenue fell. The
general default is unaffected — insurance trust is excluded
on both sides of the seam.