← Back to Writing

Data Should Deepen Judgment, Not Replace It

Five disciplines for research, evaluation, and learning inside community institutions

For more than a decade, nonprofit leaders have been urged to become data-informed rather than data-driven. Beth Kanter helped popularize that distinction, and the fields of strategic learning and equitable evaluation have pushed the argument further. They have challenged organizations to use measurement for learning, place evaluation upstream in strategy, and examine who gets to define evidence and success.

I agree with that body of work. I also think community institutions need a more demanding standard.

That difference is more than a turn of phrase. Information can inform a decision and still leave the assumptions, power relationships, and consequences surrounding that decision untouched. Data deepens judgment when it makes leaders more disciplined about what they know, more honest about what they do not know, more attentive to whose knowledge is missing, and more accountable for what they decide to do next.

This is the standard I arrived at after more than two decades of building and using research and evaluation inside a national community-based organization, while also working with researchers, practitioners, families, funders, and public systems. The value of evaluation was never simply that it could produce a finding. Its value was that it could improve the quality of the questions we asked, the choices we made, and the institution we were becoming.

Organizations can be rich in data and poor in insight

The nonprofit and public sectors have become much better at collecting information. Organizations now have dashboards, performance indicators, participant surveys, implementation databases, quality measures, and external studies. Yet the presence of data does not guarantee the presence of learning.

A dashboard can tell a leadership team that participation declined. It cannot, by itself, explain whether the decline reflects staffing instability, transportation barriers, distrust, a change in eligibility, an implementation problem, or a program that no longer fits what families need. An outcome measure can show variation across sites. It cannot determine whether the variation represents quality, local context, differences in who is being served, or limitations in the measure itself.

Numbers carry authority because they compress complexity. That is part of their usefulness and part of their danger. A clean chart can make a contested reality look settled. A statistically significant result can obscure questions about who was included, what was measured, what was missed, and whether the result matters in practice. A target can focus an organization, but it can also redirect attention toward what is easiest to count rather than what is most important to understand.

Measurement is never neutral. Someone decides what success means, which outcomes matter, what time horizon counts, which people are included, and what level of variation is acceptable. Those choices reflect a theory of change, a set of values, and a distribution of power. The goal is not to eliminate judgment from those choices. That is impossible. The goal is to make judgment visible, evidence-based, contestable, and capable of improving.

What this looked like in practice

At ParentChild+, where I spent much of my career, we worked through community-based organizations to support families and young children. As we built systems for research, program quality, and continuous improvement across a diverse network, no single source of evidence could tell us how the model was operating.

Child and family outcome measures could show patterns over time. Structured observations could help us examine the quality of practice. Participation and implementation data could tell us about reach, consistency, and dosage. Practitioner reflection could reveal where the model encountered the realities of staffing, supervision, language, family schedules, or local service systems. Families could tell us whether the support felt useful, respectful, relevant, and worthy of their trust.

The most important work happened when those sources of knowledge were brought together. A difference between sites was not simply a ranking. It was an invitation to investigate. What might explain the variation? Was the program being implemented differently? Did staff need a different kind of support? Were local conditions shaping participation? Were families describing progress that our measures were not designed to see? Was the measure itself functioning differently across languages or contexts?

These conversations did not make evaluation softer. They made it more rigorous. They forced us to test assumptions, consider alternative explanations, and connect evidence to decisions about training, supervision, program design, measurement, and organizational support. The purpose was not to defend the model or to find a story that made the numbers more comfortable. The purpose was to understand reality well enough to improve the work.

The Five Disciplines of Deep Judgment

Over time, I have come to think of this work as five disciplines. They are not a new evaluation methodology, and they do not replace good research design. They describe the institutional practices that allow evidence to become learning and learning to become better judgment.

1. Co-frame the question and the meaning of success

Evaluation often begins too late. Leaders have already chosen the strategy, designed the program, and defined the outcomes. Researchers are then asked to determine whether the plan worked. By that point, some of the most consequential judgments have already been made.

Deep judgment begins upstream. It asks who helped define the problem, whose aspirations shaped the outcome, and what evidence would genuinely challenge the strategy. It distinguishes between what a funder needs to report, what a board needs to govern, what practitioners need to improve, and what families and communities need the institution to understand.

There is a profound difference between asking people to respond to questions and inviting them to help determine which questions deserve to be asked. When practitioners, families, and community partners help frame the inquiry, they are not merely increasing participation. They are improving the validity of the work by bringing knowledge that organizational leaders and external researchers may not possess.

2. Triangulate across multiple ways of knowing

The familiar debate between numbers and stories is a false choice. Quantitative methods can reveal patterns, differences, and changes that individual experience cannot establish. Qualitative methods can show how a program is understood, why implementation varies, how trust is built or lost, and what a standardized instrument was never designed to capture.

Mixed methods are not a diplomatic compromise. They are often the only honest response to a complex social intervention. Outcome data may tell us where to look. Interviews and observation can help explain what we are seeing. Implementation data can show whether the intended experience was actually delivered. Practitioner knowledge can distinguish a meaningful signal from noise. Community knowledge can reveal a question the formal study never considered.

Triangulation also protects against the seduction of a single compelling result. When different sources converge, confidence increases. When they conflict, the conflict is not an inconvenience to be averaged away. It is information. It tells us that our explanation is incomplete and that more careful inquiry is needed.

3. Interpret evidence with the people closest to the work

Proximity produces knowledge. Practitioners know where a model bends under real conditions. Families know which supports help, which requirements create burden, and whether an institution has earned trust. Community organizations understand how language, housing instability, transportation, immigration concerns, neighborhood history, and public systems shape participation.

This does not mean that every perspective is automatically correct, or that anecdote should substitute for analysis. It means that people closest to the work should be treated as interpreters of reality, not merely as sources of data. Their role is not only to provide a quote for the final report. They should be able to examine findings, challenge explanations, identify what is missing, and help determine what the organization learns.

Collective interpretation is especially important when results are unexpected. Leaders who are far from implementation can quickly produce explanations that preserve the existing strategy. Practitioners and participants can expose where those explanations fail. The aim is not consensus for its own sake. It is a more accurate account of what is happening and why.

4. Make power visible

Every evaluation has a power structure, whether it acknowledges it or not. Someone controls the funding. Someone defines success. Someone asks others to provide information. Someone decides which findings are credible, who sees them, and what consequences follow.

A serious evaluation practice therefore asks: Who selected the questions? Who carries the burden of data collection? Who owns the data? Who has access to the findings? Who is allowed to explain variation? Who benefits when the organization demonstrates success? Who bears the consequences when a target is missed?

These are not side questions about process. They shape the evidence itself. Families may answer differently when they believe honest criticism could affect a service they depend on. Staff may withhold implementation problems when disappointing results are treated as individual failure. Organizations may select measures that satisfy a funder while overlooking outcomes the community considers essential.

An equitable approach is not achieved simply by adding demographic breakdowns or a listening session. It requires redistributing some influence over what counts as knowledge and what happens because of it. People should know why information is being collected, how it will be used, what was learned, and what the institution changed. Otherwise, participation can become extraction dressed as engagement.

5. Close the loop between insight and action

Too many evaluation findings arrive after the important decisions have been made. A study is commissioned, a report is completed, and a presentation is delivered months later. The work may be technically excellent, yet remain organizationally inert.

Deep judgment requires shorter learning loops. Teams need regular opportunities to ask what the evidence suggests now, what remains uncertain, what decision is in front of them, and what they will try next. The resulting action should be explicit. Did the organization change its training, supervision, staffing, outreach, program design, budget, partnership strategy, or measurement approach? What will it watch to determine whether the change helped?

Closing the loop also means reporting back to the people who contributed their time and knowledge. Organizations routinely ask families, staff, and community partners to complete surveys or participate in interviews, then never tell them what was learned. That erodes trust and wastes an opportunity for accountability. A learning organization shows its work. It explains what it heard, what it decided, what it could not change, and what it will revisit.

Accountability and learning are not competing purposes

Funders, boards, public agencies, and communities have a right to expect evidence. Organizations should be able to show what they did, whom they reached, how resources were used, whether the work was implemented well, and whether people benefited. Accountability is part of the obligation that comes with public trust.

But accountability without learning produces performance theater. It rewards certainty, encourages organizations to manage the numbers, and makes bad news dangerous. Learning without accountability has a different weakness. It can become endless reflection without standards, decisions, or consequences.

The two purposes need each other. Accountability makes learning consequential. Learning makes accountability credible. Together, they create an environment in which organizations can name variation, examine failure, test assumptions, and still remain responsible for results.

This requires leadership. Staff will not surface problems if leaders punish candor. Evaluation teams cannot support strategy if they are positioned as compliance units. Communities will not trust listening processes if their knowledge repeatedly disappears into reports that leave institutional choices untouched.

Research and evaluation belong inside strategy

The research and evaluation function is often placed too far downstream. Strategy is set, programs are designed, funding is secured, and then evaluators are asked to measure the result. That arrangement limits evaluation to verification when it should also support inquiry, design, adaptation, and decision-making.

Research and evaluation should help leaders clarify the problem they are trying to solve, make assumptions visible, understand variation, test whether implementation matches intention, and determine what evidence would justify continuing, changing, scaling, or stopping an approach. These are not technical questions reserved for an evaluation department. They are core questions of strategy and governance.

An organization is not truly data-driven because it can populate a dashboard. It is evidence-minded when insights can travel. Information moves from frontline practice to senior leadership, from families to program design, from research findings to budgets and staffing decisions, and from external studies back into daily operations. When data stops inside a database, a research unit, or a board deck, the organization may be measuring without becoming more intelligent.

Judgment is not the enemy of rigor

No dataset eliminates the need for judgment. Leaders must still decide how much evidence is enough, how to weigh competing outcomes, how to respond when findings conflict, and how to account for values that cannot be reduced to a score.

The alternative to false objectivity is not intuition without discipline. It is judgment that can explain itself. What do we know? How do we know it? What are the plausible alternative explanations? Whose knowledge is absent? What values are shaping the decision? What are the risks of acting, and the risks of waiting? Who will experience the consequences if we are wrong?

Data deepens judgment when it makes those questions harder to avoid. Community knowledge deepens judgment when it reveals realities that decision-makers do not experience themselves. Practitioner expertise deepens judgment when it grounds strategy in implementation. Research deepens judgment when it tests belief rather than merely decorating it.

That is the role of research and evaluation inside a community institution. Not to replace human decision-making with measurement, and not to validate whatever leaders already want to do. Its role is to help the institution see more fully, reason more honestly, decide more transparently, and learn quickly enough to become better.

What becoming data-driven should actually mean

A truly evidence-minded community institution is not the organization with the most indicators. It is the organization that has built the culture, capacity, and routines to turn information into insight and insight into action.

It frames questions with the people affected by the answers. It combines methods because complex realities rarely reveal themselves through a single lens. It treats practitioners, families, and community members as knowledge holders. It examines how power shapes the production and use of evidence. It connects findings to actual decisions and reports back on what changed.

Most importantly, it understands that data is not the opposite of humanity. Used well, data can help an institution see people more clearly. It can make inequities visible, challenge comfortable assumptions, identify who is being left out, and hold an organization responsible for the distance between its intentions and its impact.

But data can do that only when we resist the temptation to treat it as certainty. The goal is not to remove judgment from our institutions. The goal is to deepen it with rigor, humility, accountability, and the knowledge of the people closest to the work.

Selected influences and further reading

  • Beth Kanter, "Data-Informed, Not Data-Driven"
  • The Bridgespan Group, "A Practical Guide to Nonprofit Measurement, Evaluation, and Learning"
  • Equitable Evaluation Initiative, "Equitable Evaluation Framework"
  • Center for Evaluation Innovation, "Evaluation for Strategic Learning: Assessing Readiness and Results"