Why So Much Civic Tech Never Moves the Needle—and What to Do About It

Why So Much Civic Tech Never Moves the Needle—and What to Do About It

Most civic tech projects start with a kind of moral urgency that’s hard to argue against. A team spots a broken process—voter registration that keeps eligible people off the rolls, public services buried in paper, court proceedings no one can track—and builds a digital fix meant to shrink the gap between what the state promises and what it delivers. The launch feels like a win. A press release goes out, a funder gets thanked, a handful of early adopters say enthusiastic things on social media. Then, without much noise, the project drifts into a strange limbo. Usage stays flat. The people it was built for—often folks navigating bureaucratic systems under real stress—never pick it up in large numbers. The underlying mess remains.

I’ve watched this pattern repeat across continents and contexts. The problem isn’t that civic tech teams lack talent or commitment. The problem is that the field hasn’t built a rigorous, shared framework for figuring out whether a project actually works. We mistake activity for impact. We count clicks instead of changed outcomes. And because we don’t systematically measure what matters, we keep building things that feel good but do little.

People collaborating around a table with laptops and notebooks in a community workspace
A team discusses civic tech metrics—but do they measure what counts?

Why the standard metrics mislead us

Most civic tech evaluations lean on a handful of easy-to-grab numbers: website visits, app downloads, form completions, survey responses. Those metrics aren’t useless, but they’re dangerously incomplete. They tell you whether someone found your tool, not whether the tool changed their life. They measure attention, not justice.

Take a project that builds an online portal for reporting broken streetlights. The dashboard might show five hundred reports filed in the first month. A funder sees that and feels reassured. But the metric that actually matters is whether the streetlights get fixed faster than they did before the portal existed—and whether the neighborhoods that historically wait longest for repairs see any improvement. If the portal just funnels complaints into the same unresponsive maintenance queue, the project hasn’t worked, no matter how many reports flow through it.

Another trap is the satisfaction survey. A civic tech team drops a short questionnaire after each interaction. Users might rate the experience highly because the interface is clean and the language is friendly. But a pleasant interaction with a digital tool isn’t the same as a resolved dispute with a landlord, a corrected error on a benefits application, or a successfully expunged criminal record. Satisfaction can mask stagnation.

A person staring at a laptop screen displaying charts and data
Dashboards look compelling, but do they reflect real-world change?

Building an evaluation framework that starts with people

If we want to know whether a civic tech project works, we have to anchor our evaluation in the experience of the people the project is meant to serve. That sounds obvious, but it demands a discipline most teams find uncomfortable. It means defining success together with users, not in a conference room. It means tracking outcomes that may take months or years to materialize. And it means accepting that a tool might succeed on one measure and fail on another—and being honest about both.

Define success backward from the problem

Start not with the technology but with the harm you’re trying to reduce. Is it that eligible voters are being purged from the rolls without notice? That tenants facing eviction can’t find legal help in time? That parents eligible for nutrition assistance are turned away because of paperwork errors? Frame the success of the project in terms of that harm: fewer wrongful purges, more evictions prevented, more eligible families receiving benefits. Only then do you ask what the tool must do to contribute to that outcome.

This step forces a painful clarity. It often reveals that a digital tool alone can’t produce the change you want. You might discover the real bottleneck is a policy requiring wet-ink signatures, or a court clerk who refuses to accept electronically filed documents, or a state agency that simply ignores incoming data. That knowledge is valuable. It redirects energy toward advocacy, litigation, or organizing—the non-technical work that civic tech too often sidesteps.

Map the chain of causation

Once you’ve defined the outcome, work backward through the steps that must occur for the tool to contribute to it. If the project is a chatbot that helps people understand their eligibility for food assistance, the chain might look like this: a person hears about the chatbot, trusts it enough to use it, receives accurate and locally relevant information, understands the information, feels confident enough to apply, completes the application correctly, submits it to the correct office, and eventually receives benefits. A break at any link in that chain means the project fails. Evaluation means measuring each link, not just the first one.

A diverse group of people sitting in a circle discussing documents in a bright room
Good evaluation begins with listening to the people most affected.

Choose indicators that reflect power, not just activity

Many civic tech teams default to indicators that are easy to collect automatically: page views, session duration, click paths. Those indicators tell you something about user behavior, but they don’t tell you whether the tool shifted power. A more honest set of indicators might include:

  • The percentage of users who report that the tool helped them take an action they couldn’t have taken otherwise.
  • The reduction in time between a problem arising and its resolution, measured for different demographic groups.
  • The rate at which users successfully navigate a process without needing to escalate to a human intermediary.
  • The degree to which the tool reduces disparities in outcomes across race, income, geography, or language.

Collecting these indicators requires qualitative methods—interviews, case tracking, follow-up surveys weeks or months after the interaction—and a willingness to invest in research that doesn’t produce tidy real-time dashboards. It also requires building relationships with community organizations that can provide ground-truth data when official datasets are incomplete or misleading.

The hard work of attribution

Even when outcomes improve, a civic tech team must ask whether the project caused the improvement or merely coincided with it. This is the attribution problem, and it’s especially thorny in civic contexts where multiple interventions often occur simultaneously. A drop in evictions in a particular city might result from a new legal aid chatbot, or from a change in court scheduling, or from a moratorium enacted by the mayor, or from a broader economic trend. Disentangling these factors requires careful design.

Randomized controlled trials are rarely feasible in civic tech. The sample sizes are often too small, the ethical barriers too high, and the timelines too short. But other methods can strengthen attribution claims. A difference-in-differences analysis can compare outcomes in communities that received the intervention with outcomes in similar communities that didn’t. A regression discontinuity design can exploit eligibility cutoffs. Process tracing can examine whether the hypothesized causal chain actually operated in specific cases. None of these methods is perfect, but together they can build a plausible case that a project contributed to change—or reveal that it did not.

Designing for evaluability from day one

Most civic tech projects treat evaluation as an afterthought. A team builds for a year, launches, and then scrambles to figure out what data they should have been collecting. This is a mistake. Evaluability must be designed into the architecture of the project from the beginning. That means:

  • Building consent flows that allow users to opt into follow-up contact for research purposes, with clear explanations in plain language.
  • Logging not just that an action occurred, but the context—what the user was trying to accomplish, what preceded the interaction, and what followed.
  • Structuring data collection so that outcomes can be disaggregated by race, income, language, and geography, even when those variables aren’t the focus of the project.
  • Partnering with independent evaluators who can design the measurement strategy and publish results without interference from the project team or funders.

These practices require resources and humility. They also require governance structures that protect user privacy and prevent the misuse of sensitive data. Civic tech projects collect information about people at vulnerable moments. That information must be held to the highest standards of security and ethical use, and users must have genuine control over how their data is handled.

What failure looks like—and why we need to see it

A field that can’t see its failures can’t learn. Yet civic tech is littered with projects that quietly disappear, their domains lapsing, their GitHub repositories archived without postmortems. Funders rarely demand honest accounting, because they too are invested in the narrative of innovation. The result is a collective amnesia that dooms new projects to repeat old mistakes.

An honest evaluation framework must make space for failure that is documented, analyzed, and shared. A project that fails to achieve its intended outcome isn’t a scandal; it’s a source of knowledge. We need case studies that explain why a well-designed tool didn’t change outcomes: because the implementing agency resisted it, because users didn’t trust it, because the legal framework blocked it, because the problem was misdiagnosed. Those case studies are as valuable as success stories—maybe more so.

Building a culture of honest evaluation

Shifting the norms of a field takes more than a set of methods. It takes a culture that rewards honesty over hype. Funders must stop asking for “scalable solutions” and start asking for evidence of impact at any scale. Conference organizers must give as much time to postmortems as to demos. Universities must train civic technologists in evaluation design, not just in Python and design thinking.

I’ve seen what becomes possible when a team commits to this kind of rigor. They move more slowly at first, because they’re doing the unglamorous work of defining outcomes, building measurement into their systems, and listening to people who are skeptical of technology. But when they do launch, they know what they’re looking for. They can spot early signs of impact—or early signs of failure—and adjust accordingly. They can tell funders not just how many people used their tool, but how many people’s lives changed. And they can walk away from a project that isn’t working without feeling that they’ve wasted their time, because they have generated knowledge that will outlast the code.

FAQ

What’s the most common mistake in civic tech evaluation?

The most common mistake is measuring outputs instead of outcomes. Teams track how many people visited a website or downloaded a form, but fail to measure whether those actions led to a tangible improvement in people’s lives—such as receiving a benefit, resolving a dispute, or correcting an official record. Outputs are easy to count; outcomes require deeper, often qualitative, research.

How can small teams with limited budgets evaluate impact?

Small teams can begin by embedding evaluation into their existing workflows. Partner with a university researcher who needs a project for a methods course. Build short, focused follow-up surveys that users can opt into. Conduct a handful of in-depth interviews with users several months after they interact with the tool. Even a modest evaluation design, if it asks the right questions and is honest about its limits, can produce useful insights that more elaborate studies miss.

What should a funder look for in a civic tech proposal regarding evaluation?

A funder should look for a clear theory of change that connects the tool to a specific, measurable reduction in harm. The proposal should name the outcome, describe how it will be measured, and acknowledge the assumptions that could break the causal chain. It should budget for independent evaluation and specify who will own the data and the findings. Proposals that promise only website traffic or user satisfaction scores should be treated with skepticism.