How to Evaluate Whether a Civic Tech Project Actually Works
We pour hope, time, and public money into civic technology with a quiet article of faith: that a well-designed app or platform will make government more responsive, services fairer, and communities more knit together. That faith isn’t misplaced, but it ought to be tested with the same seriousness we’d bring to any public program. Far too often, evaluation shrinks down to counting downloads, tracking page views, or collecting anecdata from the loudest users. Those signals tell us something about reach. They tell us almost nothing about whether a project shifts power, reduces harm, or changes a concrete outcome for the people it claims to serve.
I’ve spent years inside the guts of civic tech—building tools, advising teams, and sometimes watching projects that sparkled on demo day quietly flop in the field. The reasons shift, but a shared thread is a thin theory of change paired with an even thinner appetite for honest measurement. The framework I’m laying out here comes straight from that experience. It’s for funders who need to place limited dollars responsibly, for product teams who sense their dashboards are lying to them, and for public servants who are asked to adopt a tool and want to know what they’re really signing up for.
Start with the Theory of Change, Not the Feature List
Most civic tech projects can describe what they build. Far fewer can lay out the causal chain that links their software to a real-world improvement. A theory of change demands that articulation. It should answer a deceptively simple question: if this tool works perfectly, who will be better off, how exactly, and through what mechanism?

A classic stumble is confusing outputs with outcomes. A portal that publishes open budget data produces an output: the data sits online. The hoped-for outcome might be that community organizations use that data to push for more equitable allocations. The mechanism isn’t the PDF. It’s the mix of accessibility, interpretability, and trust that turns data into something actionable. If the data lives in a format nobody can parse, or if the portal demands institutional knowledge no resident has, the chain snaps long before any outcome can show up.
When I evaluate a project, I ask the team to map this chain on a single page. I hunt for gaps where the logic leans on magical thinking—places where a behavioral shift is assumed with zero supporting evidence or design intent. A strong theory of change is specific enough to be wrong. That’s its value: it hands you something testable.
Test the Assumptions Behind Each Link
Once the chain is visible, you can start tugging on the links. Say a project aims to reduce evictions by helping tenants navigate legal aid. The assumptions might include: tenants know the tool exists, they trust it enough to share sensitive housing information, the information provided is jurisdictionally accurate, legal aid organizations have the capacity to respond, and courts accept the outputs. Each of these can be poked at with modest research—user interviews, usability tests with people under housing stress, secret shopper calls to legal aid intake lines, a review of court filing requirements.
I’ve watched projects stall because the team tested the interface but not the institutional context. A beautiful intake form is dead weight if the agency on the receiving end ignores submissions that arrive through a nonstandard channel. Evaluation means staring at the system the tool sits inside, not just the tool itself.
Choose Methods That Match the Maturity of the Project
Not every evaluation needs a randomized controlled trial, and demanding one too early can become a kind of avoidance—collecting rigorous data about a feature that’s still fundamentally broken. On the flip side, a project that’s been running for three years on soft metrics shouldn’t get to claim impact indefinitely without a more demanding design.
I find it helpful to think in three stages, though the boundaries leak a little.
Formative Evaluation: Is It Solving a Real Problem for Real People?
In the early stage, the central question is whether the problem is understood correctly and whether the proposed solution is even plausible. Methods here are qualitative and exploratory: ethnographic observation, diary studies, cognitive walkthroughs with people who have lived experience of the issue. The standard for evidence isn’t statistical significance; it’s whether the team can describe the problem in words that match what people actually live through.

I once reviewed a project designed to help small business owners navigate permitting. The team had built a wizard that asked a series of branching questions. In testing, we discovered that many business owners couldn’t get past the first question because they didn’t know the zoning classification of their property. The tool assumed a piece of knowledge that simply wasn’t there. That was a formative finding—not a failure, but a bright signal that the solution needed to address a different step in the user’s journey, maybe by integrating parcel data lookup before asking anything else.
Process Evaluation: Is It Reaching the Right People in the Right Way?
Once a project is operational, you need to understand implementation fidelity and reach. This is where quantitative data gets more central, but it has to be paired with qualitative insight or you’ll misread the numbers. Key questions: who is using the tool, who is dropping off and where, what pathways do people take, and are the most vulnerable or least-served populations represented in the user base?
Disaggregation matters enormously here. An overall completion rate of 70% can hide the fact that completion among users with limited English proficiency is 20%. If the tool is meant to reduce inequities, process evaluation has to surface those gaps, not smooth them over with averages. I also look for evidence that the team is actively managing the referral or onboarding channels. A tool that depends on outreach workers to sign people up will succeed or fail based on the quality of those relationships, not just the UX.
Outcome and Impact Evaluation: Did Anything Actually Change?
This is the hardest stage, and it’s where many civic tech projects quietly pivot to a different conversation about “ecosystem building” or “capacity strengthening.” Those are legitimate goals, but they’re not the same as demonstrating that a tool reduced hunger, prevented a wrongful eviction, or increased access to benefits. If a project claims an outcome, it owes the public a credible attempt to measure it.
Credible measurement doesn’t always mean a control group, though quasi-experimental designs are undervalued in this space. Sometimes a well-constructed pre-post comparison with a clearly defined cohort is enough to inform decisions, so long as the limitations are stated plainly. What counts is that the evaluation design is proportionate to the claim being made and that the team is transparent about confounding factors. An honest evaluation report will say: we saw a 15 percentage point increase in SNAP enrollment among users, but we can’t rule out the effect of concurrent policy changes that simplified eligibility during the same period.

Indicators That Cut Through Self-Deception
Over the years, I’ve collected a set of signals that help me distinguish projects with a genuine commitment to evaluation from those that are just performing rigor. None are dispositive on their own, but together they’re surprisingly revealing.
1. The team can name a specific, measurable negative outcome they’re watching for. Every intervention carries a risk of harm. A navigation tool for public benefits might inadvertently steer people toward options that trigger an overpayment notice. A platform for reporting infrastructure problems might concentrate city resources in neighborhoods that already report aggressively, widening disparities. If a team can’t articulate what failure or harm would look like, I doubt their ability to detect it.
2. User research includes people who stopped using the tool. It’s easy and pleasant to interview enthusiastic users. The people who signed up once and never returned, or who tried and gave up, hold far more valuable information. Their absence is a dataset. A project that only studies its successes isn’t evaluating; it’s marketing to itself.
3. Metrics are tied to a decision. Every metric on a dashboard should have an owner and a trigger. If the answer to “what would you do differently if this number dropped by half?” is “we would look into it,” the metric is decorative. Evaluation is useful only when it changes behavior, and that requires pre-committing to action thresholds.
4. The team can explain the limits of their data. No dataset is complete. Administrative data misses people who never interact with the system. Survey data carries response bias. Platform analytics capture clicks, not intent. A trustworthy evaluation names these gaps and discusses what they imply, rather than presenting the available numbers as the full picture.
The Hard Question of Sustained Use
Civic tech projects often enjoy a burst of adoption driven by a launch event, a partnership with a trusted community organization, or a moment of crisis. Sustaining use over time is a different beast, and it reveals whether a tool has worked its way into people’s actual routines or remains a novelty they try once.
I look for evidence of passive retention: do people come back on their own, without a reminder or an incentive? If the project serves a need that’s episodic—like disaster recovery or annual tax filing—sustained use may mean returning at the next episode, not daily engagement. The evaluation question is whether the tool is remembered and trusted when the need recurs. That can be measured through cohort retention curves, qualitative follow-up at intervals, or partnerships that allow tracking across service touchpoints.
One pattern that troubles me is the project that survives only because a single funder keeps paying for it. Sustained use in that context may reflect organizational inertia rather than user demand. A strong evaluation will distinguish between adoption that is bought and adoption that is chosen.
Institutional Fit: The Overlooked Dimension
A civic tech project can be beautifully designed and still fail because the institution it plugs into isn’t ready to receive it. Evaluation has to look at the receiving environment: the workflows, policies, and incentives of the government agency or nonprofit expected to act on the tool’s outputs.
I’ve seen a well-crafted public comment analysis tool go unused because the agency staff responsible for summarizing comments had no authority to change their process. The tool added a step rather than replacing one, and the evaluation—had it been done—would have surfaced that mismatch early. Questions for this dimension include: does the tool integrate with existing case management systems, does it require staff to learn a new skill without dedicated time for training, and is there a person with authority who has publicly committed to using the results?
Sometimes the most important outcome of an evaluation isn’t a number but a clear, evidence-backed recommendation that the tool should not be adopted in its current form, or that the institutional prerequisites aren’t yet in place. That’s not failure; it’s responsible stewardship.
Building Evaluation into the Funding Cycle
Funders shape what gets measured by what they ask for. I’ve watched too many grant reports shrink evaluation to a section at the end where grantees list activities and assert impact without evidence. Changing that dynamic requires funders to invest in evaluation capacity, to accept negative or ambiguous findings without penalizing the grantee, and to allow enough time for meaningful learning.
One practical shift is to fund an independent evaluator as a separate line item, not as part of the project team’s own budget. Independence matters, not because project teams are dishonest, but because they’re human. They have sunk cost, confirmation bias, and a natural desire to show their work in the best light. An external evaluator can ask the uncomfortable questions and report findings without the same conflict of interest.
Another shift is to require a learning plan at the proposal stage, not an evaluation plan at the end. A learning plan specifies what the team needs to learn, when, and how they’ll use that information to adjust course. It treats evaluation as an ongoing function rather than a summative judgment. That orientation is more honest and more useful, especially for projects operating in complex systems where the path to impact is uncertain.
FAQ
How do I evaluate a civic tech project if I don’t have a research budget?
Start with the data you already have, but treat it critically. Platform analytics can tell you where users drop off, which is a starting point for hypothesis generation. Reach out to a handful of users—and former users—for short, structured conversations. Many insights come from just five to eight well-chosen interviews. Partner with a university program that needs real-world projects for students in public policy or human-computer interaction. The constraint is real, but it often pushes teams toward the most essential questions, and that’s not a bad thing.
What’s the difference between evaluating a civic tech project and evaluating a traditional government program?
The core logic is the same: define the outcome, map the causal chain, collect evidence, and assess. The differences lie in the speed of iteration and the nature of the data. Civic tech projects can often generate behavioral data more quickly and at a finer grain than traditional programs, which creates opportunities for faster learning but also risks over-interpreting shallow metrics. Additionally, civic tech projects frequently sit at the boundary between government and community, so evaluation must account for trust dynamics, digital literacy, and the informal networks that shape adoption. The institutional context may be less stable than a mature government program, which makes baseline measurement trickier.
How do I know if a project’s metrics are misleading?
A red flag is a dashboard dominated by vanity metrics: total users, page views, time on site, social media shares. These measure attention, not effect. Another red flag is the absence of disaggregation—if the team can’t show you outcomes broken down by race, income, language, or geography, they probably don’t understand their own distribution of impact. Ask to see the metric that the team is most worried about, not the one they’re proudest of. If they can’t name one, the evaluation culture isn’t honest. Finally, compare what is measured to the project’s stated mission. If the mission is about equity but the metrics are about efficiency, something is misaligned.
Closing Thoughts
Evaluating civic tech isn’t a technical exercise grafted onto a project at the end. It’s a set of habits and commitments that shape how a team thinks from the start. The projects I trust most are the ones that talk openly about what they don’t know, that treat their own data with skepticism, and that can point to specific ways their work has changed because of what they learned. That kind of intellectual honesty is rare, and it’s the single best predictor I’ve found for whether a project will actually make a dent in the lives it touches.