Measuring What Matters: A Framework for Evaluating Civic Tech Projects

Not long ago, I sat in a community meeting room with a group of election officials who had just spent eighteen months building a tool to help voters find polling places. The interface was clean, the maps loaded quickly, and the team had clearly poured themselves into the work. When I asked how they knew it was successful, one of them smiled and said, “We got a lot of thank-you emails.”

I appreciated the optimism. But thank-you emails are not evaluation. They are not a theory of change. They are not evidence that a civic technology project has made a public process more equitable, more accessible, or more trustworthy. If we want civic tech to be more than a feel-good exercise—if we want it to shift power, reduce harm, and serve people who have been systematically excluded—we need to evaluate it with rigor and honesty.

This is difficult work. Government technology projects often operate under political pressure, short timelines, and constrained budgets. Evaluation can feel like a luxury, or worse, a threat. But refusing to evaluate is itself a choice, and it is one that usually protects the status quo. In this article, I want to walk through a practical framework for evaluating whether a civic tech project actually works. I will draw on my experience inside and outside government, and I will be frank about the uncomfortable tensions that arise when we ask hard questions about impact.

People sitting around a table discussing documents and laptops in a collaborative setting

Start with the Theory of Change

Before you can evaluate a civic tech project, you need to know what it is supposed to do. This sounds obvious, but I have reviewed dozens of project plans that skip straight to features and outputs. A theory of change forces you to articulate the causal logic: if we build X, then Y will happen, which will lead to Z. It exposes assumptions and makes them testable.

A strong theory of change for a civic tech project answers three questions:

  • Who is the intended beneficiary? Be specific. “The public” is not a beneficiary. Low-income renters facing eviction in a specific county are beneficiaries. Elderly voters who do not use smartphones are beneficiaries. If you cannot name the group, you cannot measure whether you reached them.
  • What behavior or condition needs to change? Civic tech is not about delivering information for its own sake. It is about enabling someone to do something they could not do before, or reducing a harm they were experiencing. Maybe a person needs to complete a form, understand a right, or trust a process enough to participate.
  • What is the mechanism? How will the technology cause the change? Is it by reducing a cognitive burden? By routing a request to the correct office? By surfacing data that was previously hidden? The mechanism is the bridge between the tool and the outcome, and it must be plausible.

I once worked with a team that built a dashboard to track city council votes. Their theory of change was that if residents could see how their representatives voted, they would become more engaged advocates. That mechanism—visibility leading to advocacy—depends on many intermediate steps: residents must find the dashboard, understand the votes, care about the issues, and have the capacity to act. Without testing those steps, the team could not know whether a lack of advocacy meant the tool was failing or the theory was wrong.

Distinguish Outputs from Outcomes

One of the most common mistakes in civic tech evaluation is conflating outputs with outcomes. Outputs are things you produce: a website launched, an app downloaded, a dataset published. Outcomes are changes in the world: a higher percentage of eligible people registered to vote, a reduction in wrongful evictions, an increase in public trust. Outputs are easy to count; outcomes are hard.

I have seen project reports that boast about “10,000 unique visitors” to a benefits eligibility tool. That is an output. The outcome question is: did those visitors successfully enroll in benefits they were entitled to? Did they do so faster or with less distress than before? If you cannot answer that, you do not know whether the tool worked.

To measure outcomes, you need to define them early and build data collection into the project from the start. This often requires collaboration with the government agencies that hold administrative data—a fraught but necessary relationship. It may also require qualitative methods: interviews, observations, and user testing that probe beyond surface-level satisfaction.

Close-up of hands writing on a notepad next to a laptop and coffee cup

Ask Who Is Left Out

Evaluation must be explicitly concerned with equity. A civic tech project can look successful in aggregate while worsening disparities. If a new online portal makes it easier for college-educated, English-speaking homeowners to contest property assessments, but does nothing for renters or people with limited digital literacy, the project may have increased inequality even as aggregate metrics improved.

I recommend conducting a disaggregated analysis as a standard practice. Break down your outcome metrics by race, income, language, age, disability status, and geography. If the data does not allow that, the project has a data equity problem that needs to be addressed before meaningful evaluation is possible. In some cases, you may need to partner with community organizations to reach populations who are invisible in administrative records.

It is also worth asking whether the project’s theory of change assumes a level playing field. A tool that helps people “find their elected representative” presumes that contacting that representative is a viable path to redress. For many communities, that presumption does not hold. Evaluation should test whether the promised mechanism actually functions for the people who need it most.

Watch for Unintended Consequences

Technology interventions in public systems rarely have only the effects we intend. An online reporting tool for potholes might lead a city to redirect resources to neighborhoods where residents have smartphones and time to file reports, neglecting areas with deeper infrastructure needs. A public-facing dashboard of police complaints could expose complainants to retaliation if anonymity is not carefully protected.

Evaluating unintended consequences requires humility and a willingness to listen. It means setting up feedback channels that are accessible to people who do not come to public meetings. It means reading between the lines of user support tickets. It means explicitly asking: who might be harmed by this project, and how would we know?

In one project I advised, a team built a tool to simplify the process of applying for a criminal record expungement. The tool was genuinely easier to use than the paper form. But during evaluation, we discovered that some users were confused about whether they qualified, and a small number submitted applications that triggered review processes that actually delayed their cases. The tool had reduced one friction while introducing another. We only caught it because the evaluation plan included interviews with legal aid attorneys who saw the downstream effects.

Measure What Users Actually Experience

Civic tech teams often measure usability through standard metrics like task completion rate and time on task. Those matter, but they are not enough. People’s experience of a government service is shaped by their history with the state, their level of trust, and their emotional state. A person who has been denied benefits before may approach a new tool with suspicion, no matter how clean the interface is.

I have found it useful to adapt evaluation methods from public health and social work. Techniques like the “most significant change” approach, where participants describe in their own words what difference a service made in their lives, can surface dimensions that surveys miss. Observational research in settings where people actually use the tool—a library computer lab, a community center, a phone with a cracked screen—reveals context that a usability lab never will.

It is also important to measure trust. In the public sector, a transaction is never just a transaction. If a voter lookup tool works perfectly but leaves the user feeling surveilled or confused about data privacy, the project has failed on a dimension that matters for democratic legitimacy. Trust can be measured through validated survey instruments, but it must be measured longitudinally; a single snapshot is not enough to understand whether a project is building or eroding public confidence.

Diverse group of people collaborating over a shared laptop in a bright workspace

Build Evaluation into the Project Lifecycle

The most common failure mode I see is treating evaluation as an afterthought. A team spends months building, launches, and then someone says, “We should probably check if this worked.” By that point, the budget for evaluation is gone, baseline data was never collected, and the team is already onto the next thing.

Evaluation should begin at the project conception stage. That means writing a theory of change, identifying key outcome metrics, and planning data collection before a single line of code is written. It means budgeting for evaluation staff or external researchers. It means building instrumentation into the product itself—not just analytics tracking, but consent-based mechanisms for following up with users to understand outcomes.

I advocate for a phased approach: formative evaluation during development to catch problems early, summative evaluation after launch to measure impact, and ongoing monitoring to detect drift. Each phase has different methods and different questions. Formative evaluation might rely on prototypes and qualitative testing; summative evaluation might require quasi-experimental designs if randomization is not possible. The point is that evaluation is not a one-time event but a continuous practice.

Be Honest About Failure

Civic tech projects operate in complex systems where failure is common. A tool can fail because the theory of change was wrong, because the implementation was flawed, because the political context shifted, or because the problem was never technological in the first place. The field will not advance if we only talk about successes.

Publishing honest evaluations—including negative results—is a public good. It helps other teams avoid repeating mistakes. It builds credibility with the communities we claim to serve. And it models a kind of institutional humility that is rare in government technology. If your project did not achieve its intended outcomes, say so, and explain what you learned. That is not a failure of evaluation; it is evidence that evaluation is working.

I have written before about the importance of learning from projects that fall short. The temptation to spin or bury disappointing results is strong, especially when funders or political sponsors are watching. But the civic tech field is littered with graveyards of tools that were launched with fanfare and quietly abandoned. We owe each other the truth about what happened and why.

Practical Steps for Getting Started

If you are part of a civic tech team and you want to take evaluation seriously, here is where I would begin:

  1. Write a one-page theory of change. Do it with your team, and do it before you start building. Make sure it names beneficiaries, the behavior change, and the mechanism. If you cannot fit it on one page, you have not clarified your thinking enough.
  2. Identify two to three outcome metrics that you can track over time. These should be outcomes, not outputs. If you do not have access to the data you need, start building relationships with the people who do.
  3. Plan for a disaggregated analysis. Decide which demographic dimensions matter for equity in your context, and figure out how you will collect that data ethically and legally.
  4. Set aside at least 10% of your project budget for evaluation. If you cannot afford 10%, you are under-resourced, and you should be transparent about that with stakeholders.
  5. Share your findings publicly, even if they are uncomfortable. Write a report, give a talk, publish a case study. Make your learning part of the public record.

These steps are simple in concept and difficult in practice. They require discipline, resources, and a willingness to confront uncomfortable truths. But without them, we are just guessing. And in a field that affects people’s access to housing, healthcare, justice, and democratic participation, guessing is not good enough.

Frequently Asked Questions

What if our project is too small to justify a formal evaluation?

No project is too small to benefit from clear thinking about whether it works. For a very small project, you might not be able to do a quasi-experimental impact evaluation, but you can still write a theory of change, track a few outcome indicators, and conduct interviews with a handful of users. The key is to do something systematic rather than relying on anecdote. Even a lightweight evaluation can surface assumptions that need testing and prevent you from scaling a broken idea.

How do we evaluate a project that is meant to improve trust or legitimacy, not a transactional outcome?

Trust and legitimacy are real outcomes, and they can be measured, but it requires careful instrument design. Validated survey scales exist for measuring trust in institutions, perceived fairness of a process, and related constructs. The challenge is usually sampling: you need to reach a representative group of people, not just the ones who voluntarily provide feedback. In some cases, you can embed short trust-related questions into the service itself, with appropriate consent. Longitudinal measurement is especially important for trust, because it can shift slowly and in response to many factors outside your project.

What should we do if our evaluation shows the project is not working?

First, resist the urge to hide the results. Share them with your team, your stakeholders, and ideally the public. Then, investigate why. Is the theory of change flawed? Was the implementation poor? Did external conditions change? The answer will determine your next move. You might need to redesign the tool, pivot to a different approach, or—in some cases—shut the project down. Shutting down a project that does not work is a responsible outcome, not a failure. It frees up resources for things that might actually make a difference. Document what you learned and treat it as a contribution to the field.

How can we evaluate a project when the outcomes take years to materialize?

Long-term outcomes are the hardest to measure, but you can identify intermediate outcomes that are on the causal pathway and can be measured sooner. For example, if your project aims to increase civic participation among young people over a decade, you might measure shorter-term changes in political knowledge, sense of efficacy, or intention to vote. These intermediate outcomes are not a substitute for the ultimate outcome, but they can give you early signals about whether the theory of change is holding up. Pair them with a plan for a follow-up study when the long-term data becomes available.

Evaluating civic tech is not about checking a box or satisfying a funder. It is about taking seriously the responsibility that comes with building tools that touch people’s lives and shape public systems. It is about insisting that good intentions are not enough. I believe we can do better, and I believe we must.