The Dataset That Stopped Updating: A Field Method for Auditing the Open Data Portals Cities Quietly Abandoned

When a dataset stops updating, the first instinct is to treat it as a technical failure. A cron job died. A server filled up. Someone left the agency and nobody took over the ETL script. All of that may be true. But for anyone working in procurement accountability or democratic infrastructure oversight, the more useful reading is different: a stopped dataset is a governance artifact. It tells you something about ownership, funding, and the terms under which a public digital system was acquired and is being maintained. The question is not “why is this broken?” but “who was supposed to keep it running, and what did the contract or policy actually require?”

This is a field method for answering that question. It is written for the practitioner who has to produce a finding that survives contact with a city attorney, a council staffer, or a vendor’s account manager. It assumes you have no access to internal systems, no subpoena power, and no interest in overclaiming what a stale timestamp proves.

Why a stopped dataset is a governance signal

The federal open data landscape offers a useful contrast. Data.gov is described as the federal government’s open data site, with the stated aim of making government more open and accountable. The OPEN Government Data Act made Data.gov a statutory requirement rather than a policy, requiring federal agencies to publish information online as open data in standardized, machine-readable formats with metadata included in the catalog. The same law directs GSA to work with OMB and the Office of Government Information Services to establish an online repository of tools, best practices, and schema standards for open data practices across the federal government.

That is a statutory floor. Municipal open data portals generally do not have one. Data.gov notes that numerous states, cities, and counties have launched open data sites and that it includes non-federal data in its catalog through collaboration. But collaboration is not the same as obligation. A city portal may rest on an executive order, a mayoral initiative, a vendor contract, or nothing more durable than a grant-funded pilot. When the dataset stops updating, you are often watching the difference between a policy commitment and a legal requirement become visible.

The Sunlight Foundation’s Open Cities project described its work as making municipal governments more transparent, accountable, and participatory through assistance with community-centered open data and open contracting. Its Web Integrity Project monitored changes to government websites, describing the purpose as revealing shifts in public information and access to web resources as well as changes in stated policies and priorities. That framing is the right one for this work. A portal that stops updating has changed its stated policies and priorities, whether or not anyone announced it.

Baseline the catalog’s claims about itself

The audit unit is not the data values. It is the catalog’s own claims about the data. Before you look at a single record, capture what the portal says about itself: the stated update frequency, the named publishing agency, the license, the last-modified field, the refresh cadence in the metadata. These claims are the object of study. If a portal says “updated monthly” and the most recent record is fourteen months old, the discrepancy between stated and actual cadence is the finding. You can document that without any access to internal systems, and it is much harder to dismiss than a general complaint about stale data.

This is also where you learn what the portal is willing to be held to. A catalog that publishes a specific cadence is making a commitment. A catalog that publishes no cadence at all is making a different kind of statement. Both are findings.

A repeatable staleness protocol

A single observation cannot distinguish a slow publisher from an abandoned one. You need a fixed observation window and at least two timestamped snapshots. The protocol is deliberately simple because it has to be reproducible by someone who is not a data engineer and defensible to someone who is.

Step one: capture the catalog metadata. For each dataset in scope, record the dataset title, the named publisher, the stated update frequency, the last-modified field as displayed, the license, and the URL. Take a screenshot or save the HTML. Note the timestamp of capture in UTC.

Step two: capture a content fingerprint. For tabular data, record the row count and, where the platform exposes it, a checksum or hash of the file. For geospatial data, record the feature count and the extent. For any dataset, record the most recent date value in the date column, if one exists. This is your baseline.

Step three: wait a fixed interval. The interval should be long enough that a genuinely monthly publisher would have published at least once, and short enough that you can complete the audit. Four to six weeks is a reasonable default for datasets claiming monthly or quarterly updates. For datasets claiming annual updates, the interval needs to be longer, or the finding needs to be framed as a missed annual cycle rather than a broken pipeline.

Step four: re-capture. Repeat steps one and two. Compare. Record count unchanged, checksum unchanged, last-modified field unchanged, and most recent date value unchanged across the interval is the signal. Any one of those changing is evidence the pipeline is alive, even if slowly.

Step five: record the direction of travel. Two snapshots establish whether the dataset is moving or not. Three snapshots establish a trend. For a public memo, two is usually enough to state the observed discrepancy. For a pattern finding across multiple datasets, three is better.

Separating three failure modes

Collapsing everything into “the portal is dead” produces a finding nobody can act on. The three failure modes route to different remedies and different responsible parties.

Broken pipeline. The publisher is still there, the dataset is still listed, but the automated update has stopped. This is fixable by the current operator. The remedy is a maintenance request, and the oversight question is why the maintenance obligation was not being monitored.

Publisher departure. The named publishing agency has been reorganized, defunded, or has stopped publishing. The dataset may still be listed, but nobody owns it. This needs reassignment of ownership. The oversight question is who is accountable for the dataset now, and whether the portal’s governance documents say.

Quiet deprecation. The dataset has been removed from the catalog, or the portal itself has been taken down, without a public decision record. This is the most serious because it removes the public’s ability to audit. The remedy is a public decision record. The oversight question is what process governed the removal and whether it was documented.

These are not mutually exclusive. A publisher departure often precedes quiet deprecation. But reporting them separately forces the responsible party to be named, and naming the responsible party is what turns an observation into an oversight action.

Reading the metadata layer

Metadata absence is a first-class finding, not a technical footnote. When the last-updated field is missing, auto-generated, or contradicts the record contents, the portal has removed the public’s ability to audit it. That is a distinct harm from the data being stale. A stale dataset with honest metadata tells you the pipeline is broken. A stale dataset with missing or misleading metadata tells you the portal is not designed to be audited at all.

Watch for three specific patterns. First, a last-updated field that reflects the catalog’s own refresh rather than the dataset’s content, so it always shows a recent date even when the underlying records are years old. Second, a stated update frequency that is contradicted by the record contents, such as “daily” on a dataset whose most recent record is from the previous fiscal year. Third, a publisher field that names a department that no longer exists or a vendor that no longer holds the contract. Each of these is documentable from the public interface and each is a finding in its own right.

Writing it up for oversight

The write-up should be tiered by audience. A single memo that tries to serve everyone serves no one.

The public memo states the observed discrepancy and the method. It names the dataset, the stated cadence, the observed cadence, the observation window, and the capture timestamps. It does not speculate about causes. It does not name individuals. It states what was observed and how, and it invites correction. This is the document that can be published without legal review in most jurisdictions, because it is a factual observation of a public interface.

The staff-level note names the responsible office and the specific contract, policy, or statute that governs the update obligation. This is where you connect the observed discrepancy to the instrument that was supposed to prevent it. If the portal rests on an executive order, cite the order. If it rests on a vendor contract, cite the maintenance clause, or note its absence. If it rests on a grant, cite the grant terms. This note is for the council staffer or the CIO’s office, and it should be specific enough that they can act on it without further research.

The procurement note flags missing maintenance terms for the next award. If the current contract has no data refresh obligation, no uptime requirement for the portal, and no metadata standard, that is the finding. The remedy is language in the next solicitation. The Sunlight Foundation’s Open Data Policy Hub offers a step-by-step guide to creating an open data policy, covering setting objectives, drafting language, collecting feedback, implementing the policy, and promoting the policy. That guide is a reasonable starting point for drafting the language, though it is a policy guide rather than a contract template.

What the method does not prove

State the limits once, plainly, in the method section rather than repeating them as caveats throughout. A stale dataset does not establish intent. It does not measure service quality. It does not by itself prove a legal violation. It does not tell you whether the underlying program is functioning or failing. It tells you that a public commitment to publish is not being met, and it gives you a documented, reproducible basis for asking why.

That is enough. The value of the method is not that it proves misconduct. It is that it converts a vague sense that the portal is neglected into a specific, timestamped, reproducible finding that a decision-maker has to respond to. The response may be a fix, an explanation, or a decision to deprecate the dataset publicly. All three are better than silence.

What to ask in your next meeting

Bring one dataset. Bring the stated cadence, the observed cadence, and the two capture timestamps. Ask the responsible office a single question: what is the maintenance obligation for this dataset, and who is accountable for meeting it? If the answer is that no one knows, you have found the governance gap. If the answer is that the obligation exists but is not being monitored, you have found the oversight gap. Either way, you have moved the conversation from “the portal is broken” to “the commitment is not being kept,” which is a question a decision-maker can actually answer.

FAQ

How many snapshots do I need before I can publish a finding? Two snapshots taken a fixed interval apart are enough to state an observed discrepancy. Three snapshots establish a trend and are better for a pattern finding across multiple datasets. One snapshot is not enough to distinguish a slow publisher from an abandoned one.

What if the portal does not publish a stated update frequency? That absence is itself a finding. Record it. A portal that does not state a cadence has not made a commitment the public can hold it to, which is a different governance problem from a missed cadence, but it is still a problem worth documenting.

Does the OPEN Government Data Act apply to my city’s portal? The retrieved source describes federal agency obligations under the Act. It does not establish that the Act’s requirements apply to municipal portals. Check your city’s own policy, executive order, or contract for the governing obligation.

What if the dataset is stale but the portal says it is current? That contradiction is the finding. Capture both the stated status and the record contents with timestamps. The discrepancy between the portal’s claim about itself and the observable state of the data is documentable from the public interface and is harder to dismiss than a general complaint about staleness.

Can I use this method on a state or federal portal? The protocol is platform-agnostic. The governance analysis will differ because the governing instruments differ. Federal portals have a statutory floor. State and municipal portals may rest on executive orders, contracts, or grant terms. The method of capturing the catalog’s claims about itself and comparing them to observed behavior works the same way.