How to Know If a Civic Tech Project Actually Works: A Principled Framework
There is a specific kind of hope that fills a room when someone demos a new civic tech project. A sleek interface, a dashboard pulsing with real-time data, a tool that promises to connect people to services with a tap. The crowd—funders, government partners, neighborhood advocates—leans in. They can see the potential. But the real story of whether that project works almost never gets told in the demo. It comes out months or years later, in the slow pileup of small evidence, in interviews with users who found the tool baffling, in server logs that show a cliff-like drop-off after week one, and in the messy bureaucratic machinery the software never touched.
I have spent my professional life trying to understand that second story. From early work in electoral politics to my time focused on technology and governance at Google and elsewhere, I keep returning to the gap between what we intend and what actually happens. Civic tech lives in that gap. The field runs on a deep belief that digital tools can make democracy stronger, public services better, and give people a real voice. But belief on its own is not enough. We need a disciplined, principled way to evaluate whether these projects are delivering on their promises—and to learn from them when they are not.
This article lays out that way of thinking. It is not a checklist you knock out in an afternoon. It is a mental frame that should shape the entire life of a project. It borrows from program evaluation, user research, and political theory. Above all, it argues that evaluation needs to be a core practice, not a box to tick at the end.

Start with a Theory of Change—and Get Specific
Every civic tech project sits on a theory of change, whether the team realizes it or not. It is the causal story that links the technology to the outcome. A simple version might be: If we build a mobile app that shows real-time bus locations, then more people will take public transit, and that will cut carbon emissions. The trouble is, most theories of change stay unspoken. And when they stay unspoken, you cannot test them.
The first evaluation step is to spell out that theory in full. Who is the user? What exact behavior change do you expect? What are the intermediate steps? In the transit app example, the chain could be: the app gets downloaded → users open it before walking out the door → they tweak their departure time → they wait less at the stop → they report higher satisfaction → they pick transit for the next trip → overall ridership climbs. Every link in that chain is a hypothesis you can measure.
I picked up this discipline from political campaigns, where the theory of change is blunt and concrete: If we knock on 10,000 doors in this precinct and deliver this message, we will boost turnout by 3 percentage points. Civic tech projects rarely feel that same electoral heat, but the need for precision is just as high. A fuzzy goal like “increase civic engagement” falls apart under any real scrutiny. Engagement how? Measured by what? Over what stretch of time? For which people?
Without a clear theory of change, evaluation turns into a fishing trip. You gather data and pray something looks good. With one, you know exactly what you are testing and why.
Define Success Before You Write a Line of Code
This sounds painfully obvious, yet projects skip it with startling regularity. The urge to build is strong. There is funding to lock down, a prototype to show off, a partner eager to see forward motion. But defining success after the tool goes live is a straight path to confirmation bias. You will find metrics that make the project look good because everyone around the table wants it to look good.
I push for a short document—two pages max—that spells out: 1) The primary outcome measure. The single number that, if it moves the right way, tells you the project worked. For a tool that lets residents report potholes, it might be the share of reports that turn into a repair within 14 days. 2) The secondary measures. These catch the wider ripples, the good and the bad. Does the tool reduce total calls to the city hotline? Does it shift the demographic mix of who reports problems? 3) The guardrail metrics. What must not get worse? For example, the tool must not push the workload on frontline staff to the point of burnout. 4) The qualitative indicators. What do users actually say about their experience? What do agency staff say?
Share this document with everyone—funders, engineering teams, government partners, community representatives—and argue about it. If you cannot agree on what winning looks like, you are not ready to start building.

Measure What Happens, Not Just What People Click
Digital tools make it dead simple to measure engagement. Page views, session length, button clicks—these numbers are clean and seductive. But in civic tech, engagement is rarely the real destination. The destination is a change out in the world: a cleaner park, a government that listens better, a more informed public. Clicks are a weak proxy at best, and a dangerous distraction at worst.
I once reviewed a project built to boost public participation in city budget decisions. The team proudly reported thousands of visits to the interactive budget tool. But when we dug past the surface, we found the median time on page was under 30 seconds, and almost nobody touched the features that allowed for meaningful trade-offs. The shiny engagement numbers had masked a failure to reach the actual outcome.
Solid evaluation traces the path from a digital action to a real-world effect. That often means connecting data sets that do not naturally speak to each other: app usage logs stitched to administrative records of service delivery; survey responses matched against voter files; social media sentiment read alongside city council meeting minutes. This is grubby, difficult work. It takes data-sharing agreements, privacy protections, and a lot of analytical patience. But it is the only way to learn if the technology changed anything that actually matters.
Look for the Distribution of Impact
Averages lie. A civic tech project can post positive average results while serving some groups well and others poorly, or even hurting them. Equity has to sit at the center of the evaluation, not get bolted on as a separate exercise.
Start by asking: Who is using the tool, and who is not? Compare the demographics of your users to the demographics of the population you meant to serve. A 311 app adopted mostly by wealthier, whiter neighborhoods may speed up service for those residents while widening the gap for communities that depend on phone calls or in-person visits. If you do not measure that disparity, you cannot begin to fix it.
Go past demographics and look at power. Does the tool shift influence toward people who already hold it? A participatory budgeting platform that demands extensive online research and structured argument may favor those with more education and free time. A tool that makes it frictionless to ping elected officials may simply amplify voices that were already comfortable doing so. Evaluation should ask not only does it work? but for whom does it work, and at whose expense?
This means collecting data that many projects avoid. It means asking about race, income, education, and comfort with digital tools. It means doing outreach and interviews in communities that do not naturally show up in your user base. It means a willingness to report findings that make you squirm. I have a lot of respect for the organizations that publish evaluations showing uneven impact—they are treating evaluation as a learning tool, not a fundraising prop.
Account for the Institutional Ecosystem
Civic tech never floats in a void. Every tool sits inside a tangle of laws, regulations, bureaucratic habits, and political crosswinds. Ignoring that ecosystem is one of the most common reasons projects stumble—and one of the hardest things to capture in an evaluation.
Take a tool meant to smooth the process of applying for housing assistance. The software could be elegant, the user flow intuitive. But if the underlying eligibility rules are tangled and poorly explained, if caseworkers never get trained on the new system, if the legal framework demands paper signatures the tool cannot capture, then the technology will fall flat. An evaluation that stares only at the software will miss every one of those barriers.
I recommend a mixed-methods approach here. Quantitative data can show you where users drop out of the process. Qualitative research—interviews, observations, shadowing—can tell you why. When I worked on voter information tools, we spent hours parked beside local election officials, watching them process registrations. We learned things about their workflow, their quiet fears, and their unofficial workarounds that no dataset would have ever revealed. Those insights were essential to understanding whether our tools had a prayer of succeeding.
Evaluation also has to account for the political clock. A project that shows promising early signals might get defunded after the next election. A tool that threatens settled interests may meet quiet sabotage. These are not external irritants to wave away; they are part of the real texture of civic tech. A full evaluation names them plainly.

Build Feedback Loops That Outlast the Grant
A single evaluation dropped at the end of a project is a snapshot. It tells you whether something worked, but it rarely tells you why, and it lands too late to change direction. The alternative is continuous feedback—systems that produce data throughout the project and feed it back into real decisions.
This can take plenty of forms. A monthly dashboard of key metrics shared with the whole team. A quarterly qualitative check-in with a rotating panel of users. An advisory group of frontline staff who review the data and raise red flags. The aim is to make evaluation a rhythm, not a one-off event.
I am especially drawn to approaches that give users a direct voice in the feedback loop. After a resident uses a tool to report a problem, can they rate what happened? Can they see what came of it? Closing the loop with the user is not just decent customer service; it is a source of data about whether the system actually functions. If people report that nothing happened after they submitted a request, that is an outcome measure every bit as real as any administrative statistic.
Funding structures often fight against this. Grants pay for the build and the launch, maybe a final report. They rarely cover the ongoing slog of monitoring and iteration. But if we mean it about impact, we need to push for funding models that support the full life cycle of a project, including the unsexy work of maintenance and evaluation.
Resist the Tyranny of the Case Study
The civic tech field adores a good case study. One city where a tool transformed engagement. A tight narrative with glowing quotes from grateful users. These stories pack a punch. They inspire. They also warp the picture.
A case study is not an evaluation. It is a curated highlight, often written by the same organization that built the tool, with every incentive to spotlight success. For every celebrated case, there are many more projects that produced thin results, zero results, or negative results—and we almost never hear about them. This publication bias makes the whole field a little dumber.
I am not arguing against storytelling. Stories carry weight. They communicate values and build movements. But they need to sit alongside honest, systematic assessment. If you are evaluating a project, ask yourself: Would I publish these findings if they were disappointing? If the honest answer is no, you might be doing advocacy, not evaluation.
One practical move is to commit to a public evaluation plan before you know the results. Pre-register your measures and methods. This builds accountability and shrinks the temptation to cherry-pick. It is a habit borrowed from academic research, and it belongs in civic tech.
Embrace Humility and Negative Results
This is the hard part. The people who build civic tech are deeply invested. They work long hours, often for less money than they could earn elsewhere, because the mission matters to them. Hearing that their project did not work can feel like a moral verdict. But a project that misses its primary outcome is not a failure if it produces knowledge. It becomes a failure only if that knowledge gets buried.
I have been part of projects that did not work. A voter education tool that almost nobody used. A community forum that pulled in the same ten people every time. Looking back, the most valuable thing we did was to write blunt post-mortems, share them internally, and talk openly about what we learned. Those conversations reshaped how I think about design, about user research, about the quiet assumptions we carry into our work.
An evaluation culture that punishes negative results is a culture that will keep repeating its mistakes. We need funders who reward learning, not just glossy wins. We need conference panels where people lay out their flops. We need to treat evaluation not as a verdict, but as a practice of institutional humility.
Frequently Asked Questions
How do you evaluate a civic tech project when the outcome is slippery to measure, like building trust in government?
Trust is tricky to measure head-on, but you can spot observable behaviors that signal it. If a tool is designed to make government more transparent, you might track whether users come back to the tool over time, whether they pass along information from it, or whether they later show up to other civic activities like a public meeting. You can also use validated survey instruments that measure trust in institutions, though those work best when given to a representative sample, not just the people who used the tool. The trick is to pick a plausible behavioral stand-in and be upfront about its limits.
What if the project team is too small or stretched thin to run a full evaluation?
Start with the smallest viable evaluation. Pick one primary outcome measure that matters most and figure out how to track it, even in a rough way. Conduct five user interviews—you will learn more than you expect. Write a one-page theory of change and revisit it every month. Evaluation does not need to be expensive or academic to be useful. The discipline of asking how would we know if this is working? and gathering even imperfect data is miles better than running on assumptions. If you genuinely have zero resources for any evaluation, that is a signal the project itself may not be set up for learning and should be rethought.
How can we tell the difference between a project that just needs more time to show impact and one that is fundamentally broken?
This is a judgment call, but it should lean on evidence. Look at the early steps in your theory of change. If the very first link—say, user adoption—is not clicking despite reasonable effort, the problem may be foundational. If adoption is solid but the downstream outcome has not yet appeared, you may need more time and should inspect the intermediate steps for bottlenecks. It also helps to compare your trajectory to similar projects. Did they see impact at a comparable stage? Finally, talk honestly to users and stakeholders. Their explanations for what is happening are often more revealing than the numbers alone. If the core trouble is that the tool doesn’t meet a real need or fits awkwardly into people’s lives, no amount of extra time will fix that.
Should community members be involved in designing the evaluation itself?
Yes, whenever it is possible. Community members bring knowledge about which outcomes matter, what questions to ask, and how to make sense of findings in context. Participatory evaluation approaches—where residents help define success measures, gather data, and interpret results—can produce evaluations that are both more accurate and more trusted. This does not replace the need for methodological discipline, but it rounds it out. At the bare minimum, evaluation plans should get reviewed by people the tool is supposed to benefit, not just by the people who built it.