top of page

Beyond the theory of change: A measurement guide for modern social impact leaders

6 days ago
13 min read
dials

Corporate social impact teams have grown up quickly over the past decade. But the way organizations measure their impact has not always kept pace with the scale of investment—or with the scrutiny those investments now receive. Most organizations can explain what they intend to do and why, often through a theory of change. Far fewer can show, with credible evidence, what actually changed as a result. And increasingly, that is what stakeholders expect from any organization claiming to create positive impact.


The theory of change remains a useful and respected planning tool. But it was never meant to function as a full measurement system. When organizations treat it that way, they risk creating gaps between the impact they hope to achieve and the evidence they can actually provide. This guide offers a practical path forward. We look at how social impact teams can move from theory of change into stronger outcome measurement frameworks, including Impact Frontiers’ five dimensions of impact, Social Return on Investment (SROI), and IRIS+ metrics. We also cover the data systems and organizational habits needed to use those frameworks credibly.


The guide closes with a roadmap for teams ready to strengthen their measurement practice: assess current maturity, choose frameworks that fit both capacity and reporting needs, build minimum viable data systems before investing in dashboards, and create feedback loops that help programs improve based on evidence. Throughout the guide, you’ll also find case studies. These examples are based on our team’s decades of experience designing, managing, and measuring social impact programs. They are illustrative, not drawn from specific companies.


The theory of change is a foundation, not a finish line

A theory of change maps a causal logic: inputs and resources lead to activities, activities produce outputs, outputs are expected to lead to outcomes, and outcomes are expected to contribute to longer-term impact. Originating in program evaluation practice in the 1990s and popularized in the nonprofit and philanthropic sector by organizations such as the Aspen Institute’s Roundtable on Community Change, the theory of change has become close to a default planning artifact across CSR, ESG, and philanthropic functions.


Its strengths are real. A well-constructed theory of change forces an organization to articulate assumptions that often go unexamined, such as why a given activity should be expected to produce a given outcome. It creates a shared frame of reference across internal teams, grantees, partners, and funders. It is an effective communication and fundraising tool, because it tells a clear and logical story. We’ve used some form of a theory of change in many client projects to help clients understand how their inputs and investments can produce short- and long-term impacts. In this sense, it is largely a teaching and strategic tool—not an actual measurement system.


Where it falls short

The theory of change describes an intended causal chain. It does not, on its own, verify whether that chain held up in practice. Three limitations are common:


  • First, it is typically built once, during program design, and revisited infrequently. It captures a hypothesis at a point in time rather than an evolving, evidence-tested understanding.

  • Second, it assumes a largely linear relationship between outputs and outcomes. In most social contexts, that relationship is contingent on factors the program does not control. Think: economic conditions, parallel interventions, or individual circumstances. A theory of change rarely specifies how the organization will separate its own contribution from these other factors.

  • Third (and most consequentially for impact and environmental teams), a theory of change has no inherent metrics, baselines, counterfactuals, or thresholds for success. It tells you what to look for in principle. It does not tell you what you actually found, how confident you should be in the finding, or how it compares to a credible alternative scenario.


In short: the theory of change answers “what do we believe will happen, and why?” Measurement frameworks exist to answer “what actually happened, how much, for whom, and how do we know?” The theory of change is a vital starting point and strategic tool, but it’s not the best tool for ongoing, long-term measurement. 


Case study: Workforce readiness program

GoodBuild Corp. launched a workforce readiness program to help build its own pipeline of contractors and construction workers, while helping anyone prepare for jobs in the construction industry. GoodBuild's program leader, Carey, developed a theory of change to track completion rates and post-program employment within 90 days. Carey then published both figures publicly as evidence of success. 


At face value, this seems like a successful program—and it is. But over time, Carey and GoodBuild will find that the program might stagnate instead of scale. Funders might invest a little in year one’s pilot, but won’t increase their giving over time. Numbers will remain about the same, making program reporting less newsworthy over time. The program will continue to do good work, but it will fade into the background and ultimately provide little additional value to the business.


At year three, Carey takes a course on impact measurement and begins implementing a framework for enhanced measurement. She adds a comparison group of similar job seekers who did not participate, tracks employment retention at 6 and 12 months rather than only initial placement, and uses an IRIS+ employment metric to allow comparison with similar programs reported by other organizations. With this data, Carey can not only improve the program itself, but GoodBuild has new lenses through which to talk about impact, resulting in new funder interest, unique local media stories, and the ability to position GoodBuild executives as thought leaders on relevant industry stages. The frameworks-enhanced measurement system can support a contribution claim with reasonable confidence. The theory of change-only version could not distinguish program effect from a strong local hiring market. 


From planning to proof: Moving into measurement frameworks

If the theory of change is your starting point, where do you go next? And how are your inputs and expected outcomes or impacts translated into a longer-term measurement system? Your next step depends, in part, on what your organization is measuring and how robust you need your measurement function to be. This section reviews the most widely used frameworks for outcome and impact measurement and where each is best applied.


The Impact Management Project / Impact Frontiers five dimensions: Originally developed through a multi-stakeholder collaboration convened as the Impact Management Project and now stewarded by Impact Frontiers, this framework organizes impact measurement around five dimensions: What (the outcome, its significance to stakeholders, and how it compares to a counterfactual), Who (which stakeholders experience the outcome and how underserved they are), How Much (scale, depth, and duration of the outcome), Contribution (whether the outcome would likely have happened anyway), and Risk (the likelihood that anticipated impact does not occur as expected).


Its primary value for CSR/ESG teams is as a common analytical structure: it does not mandate a single metric, but it provides a consistent way to interrogate any program’s claimed impact, including ones with very different subject matter (education, environment, health, financial inclusion). It is widely referenced by impact investors and increasingly by corporate impact teams as a due-diligence lens.


Social Return on Investment (SROI): SROI assigns monetary values to social, environmental, and economic outcomes in order to express a program’s value relative to its cost, typically as a ratio (for example, “$4.20 of social value for every $1 invested”). It draws on cost-benefit analysis and welfare economics, using techniques such as financial proxies for non-market outcomes.


SROI’s appeal is comparability: a single ratio is easy to communicate to executives and boards. Its risks are equally well known. Valuation choices materially affect the ratio, comparing SROI ratios across different programs or organizations can be misleading unless methodology is closely aligned, and the technique can create incentives to select outcomes that are easy to monetize favorably rather than outcomes that matter most. Used carefully, with transparent assumptions and sensitivity analysis, SROI is a legitimate tool for a single program’s internal business case. Used as a headline external claim without that transparency, it is a common source of credibility risk.


IRIS+ and standardized metrics: IRIS+, maintained by the Global Impact Investing Network (GIIN), provides a catalog of standardized, pre-tested metrics organized by theme (such as financial inclusion, clean energy, or affordable housing) and aligned to the IMP’s five dimensions and to the UN Sustainable Development Goals. Rather than requiring an organization to invent its own metrics, IRIS+ allows it to select from metrics that are already defined, tested, and in many cases comparable across organizations using the same standard.


IRIS+ is particularly useful at the metric-selection stage, once a company has decided which outcomes matter and needs a credible, comparable way to measure them. Adoption among sophisticated impact investors is already significant: the GIIN’s 2024 State of the Market survey found that 70% of investors surveyed use some form of generally accepted impact metric such as IRIS+, a useful benchmark for where corporate ESG teams reporting against SDG-aligned commitments are likely headed as expectations converge between the investor and corporate sides of the impact ecosystem.


Worth mentioning is the B Impact Assessment and related certification tools: The B Impact Assessment, administered by B Lab, evaluates a company’s overall social and environmental performance across governance, workers, community, customers, and environment, and underpins B Corp certification. It is less a measurement framework for a specific program’s outcomes and more a holistic operational assessment. It is useful for benchmarking overall organizational practice but is not a substitute for measuring the outcomes of individual initiatives.


Choosing among frameworks

These tools are not mutually exclusive and serve different purposes. A practical rule of thumb:


  • Use the IMP/Impact Frontiers five dimensions as the analytical lens for any program, regardless of size.

  • Use IRIS+ to select specific, comparable metrics once outcomes of interest are identified.

  • Reserve SROI for cases where a monetized comparison is specifically needed (such as a capital allocation decision) and where the organization is prepared to disclose its valuation methodology.

  • Treat B Impact Assessment as a separate, complementary exercise in organizational benchmarking rather than program measurement.


Case study: Grow Green

Grow Green is a local nonprofit organization restoring land through a tree-planting program. Beyond helping to restore and beautify local green spaces for the benefit of the community, Grow Green derives around 60% of its funding through a carbon offset program. Local companies fund tree plantings and receive verifiable carbon offsets in return. This revenue stream continues to grow, in large part because Grow Green launched with a robust measurement system in place.


When Grow Green’s founder, Greg, was designing the organization’s model in a local impact incubator program, he began with a theory of change. This reported acres restored and trees planted. Greg was happy with his year one projections and presented them to the incubator’s board, composed of community leaders and businesses, to obtain seed funding. Nigel, the owner of a local business, suggested that Greg go back to the drawing board and apply IMP’s five dimensions instead of relying solely on the theory of change. After undergoing IMP training, Greg returned to the board with new projections and measurement plans. Now, Grow Green would assess:


  • Tree survival rates at 12 and 24 months (How Much)

  • The number of people impacted by restored areas (Who)

  • The ecological significance of the specific land restored relative to regional conservation priorities (What)

  • Whether the land would likely have been restored by other actors regardless of Grow Green’s involvement (Contributions)

  • Risk that the planted vegetation fails to establish in a changing climate (Risk)


The added dimensions highlighted that Greg’s initial planting figures, while accurate, substantially overstated ecological impact once survival and contribution were assessed. Now, Greg and his seed funders could move ahead with a higher level of confidence in what was realistically achievable, while being better prepared for the “what ifs” that would surface in years to come.


The measurement maturity curve

Organizations rarely move from “no measurement” to “full outcome attribution” in one step. Many organizations progress through a measurement maturity curve.


Stage 1: Anecdote and output counting. The organization can describe what it did (people served, funds disbursed, events held) and may have compelling individual stories, but has no systematic outcome data. This is the most common starting point. It reflects an early stage of program design where the priority is launching activities. The risk lies in treating this stage’s outputs as evidence of impact in external communications.


Stage 2: Outcome tracking. The organization defines specific outcomes tied to its theory of change (for example, changes in test scores, employment status, or health indicators) and collects data on those outcomes for program participants, typically using pre/post measurement. This stage answers “did the outcome occur,” but generally without a comparison group, so it cannot yet answer “would it have occurred anyway?”


Stage 3: Comparative and monetized measurement. The organization introduces a counterfactual or comparison basis (a comparison group, a credible benchmark, or a recognized estimation method) and may apply a framework such as SROI to express value, or align metrics to IRIS+ for comparability. This stage can credibly answer questions of contribution and relative value, subject to the quality of the comparison method used.


Stage 4: Portfolio-level impact management. The organization manages a portfolio of initiatives using consistent frameworks (typically the IMP five dimensions and IRIS+ metrics) across programs, enabling comparison and prioritization across the portfolio, and feeds measurement results back into capital allocation and program design decisions on an ongoing basis, not just at year-end reporting.


Many corporate social impact functions sit at Stage 1 or early Stage 2 today, even when their public reporting language suggests otherwise. Moving up this curve is incremental and should be matched to organizational capacity and reporting obligations rather than attempted all at once.


Data infrastructure: The measurement backbone

Framework choice is rarely the binding constraint on measurement quality. Data infrastructure is. Here’s what we mean. 


Baselines and counterfactuals

Outcome measurement requires knowing the starting point (baseline) and having a credible basis for what would have happened without the intervention (counterfactual or comparison). Many programs collect post-intervention data only, which makes it impossible to demonstrate change, let alone attribute it! Establishing baselines requires building data collection into program design from the outset, which is far cheaper than retrofitting it later.


Attribution versus contribution

Few corporate social impact initiatives operate in isolation from other forces that affect an outcome. For example, a workforce training program’s participants may also benefit from a strong local labor market, and a literacy program’s results may be influenced by school-level factors outside the program’s control. Rigorous measurement distinguishes attribution (meaning the outcome is caused specifically by this intervention) from contribution (meaning this intervention plausibly contributed to the outcome, alongside other factors), and is explicit about which claim is being made. Overclaiming attribution where only contribution can be defended is a common source of credibility risk.


Build versus buy/partner

Organizations face a choice between building internal data systems, adopting third-party impact measurement platforms, or partnering with evaluation specialists. Building makes sense when an organization has a small number of large, long-running programs where a custom system pays for itself. Buying or partnering makes sense for organizations managing many smaller grants or programs where standardization across grantees matters more than customization, and where third-party platforms can also bring methodological credibility that an internal build cannot easily offer on its own.


Right-sizing measurement cost

Measurement is not free, and a frequent (and legitimate) objection from program teams is that evaluation costs can consume a disproportionate share of a smaller program’s budget. There is no single correct ratio, but a useful discipline is to scale measurement rigor to decision stakes: programs being scaled, replicated, or used as the basis for significant external claims warrant materially more rigorous (and costly) measurement than small, exploratory, or clearly bounded initiatives.


Good to know: pitfalls and pushback

  • Over-monetization. Not every outcome needs to be translated into dollars. When organizations focus too heavily on monetary value, they may prioritize outcomes that are easier to price instead of the ones that matter most to stakeholders. It can also lead to headline ratios that look impressive but change significantly when valuation assumptions shift.

  • Metric gaming and Goodhart’s Law. Once a metric becomes a target, particularly one tied to incentives or external reporting, there is a tendency for behavior to optimize the metric rather than the underlying outcome it was meant to represent. This is Goodhart’s Law. Frequent metric review and triangulation across data sources help mitigate this.

  • Measurement fatigue among grantees and partners. Nonprofit partners and grantees, who often have far smaller budgets and staff than the corporate funders requesting data from them, can be burdened by inconsistent or duplicative reporting requirements. Standardizing on widely used metrics (such as those in IRIS+) reduces this burden by allowing one data collection effort to serve multiple funders.

  • Cost relative to program size. As noted above, measurement investment should be proportionate to program scale and to the stakes of the claims being made about it, not applied uniformly regardless of size.

  • Internal resistance. If measurement reveals that a program is not achieving its intended outcomes, this can be perceived as a threat by the teams who designed and run it. Building a culture where measurement is understood as a tool for improving programs, not solely for judging them or their owners, is a precondition for getting honest data rather than data shaped to avoid uncomfortable findings.


Case study: Community grants portfolio

Corner Kettle is a growing regional quick service restaurant chain with a commitment to increasing access to healthy foods. Its CSR program director, Pat, helped build and launch its CK Cares program, which uses store locations as distribution points for fresh local produce and dry goods to help families in need, especially in food deserts. Beyond using stores as physical distribution points, the program was designed to be largely hands-off, relying instead on grantees to deliver services and programs. Those grantees, which include nonprofits and community organizations, help source and transport the food, communicate with beneficiaries, and provide supplementary services like nutrition counseling. 


When Pat and her team launched the program, they created a theory of change that reported total grant dollars distributed and the number of grantees funded, organization by organization, with no aggregation. In year 2, a social impact consulting firm that had been working with Corner Kettle suggested that Pat look into implementing an IRIS+ framework. By year 3, CK Cares was standardizing a small set of IRIS+ metrics across grantees with related missions, enabling portfolio-level reporting on outcomes rather than only inputs. Pat then used that aggregated data to inform which types of grantee programs would receive renewed or expanded funding the following cycle, ensuring that Corner Kettle’s investment was making the biggest impact, where it was needed most. Further, since Corner Kettle relied on IRIS+ reporting, grantees did not have to submit duplicative reports to all their corporate funders—they simply had to feed data into the IRIS+ system, freeing them up to focus on mission-critical work.


A practical roadmap for communications teams


  1. Audit current state. Before selecting frameworks, map existing programs against the maturity curve in Section 5. Most organizations will find a mix of stages across their portfolio; this is normal and the starting point for prioritization.

  2. Select frameworks deliberately, not by default. Use the IMP five dimensions as a default analytical lens across all programs. Layer IRIS+ metrics where standardized data collection is feasible and comparability matters. Reserve SROI for specific decisions that call for a monetized comparison, and commit to disclosing methodology and assumptions whenever an SROI figure is shared externally.

  3. Build minimum viable data infrastructure before dashboards. Establish baselines and a defensible comparison basis before investing in reporting visualization. A dashboard built on weak underlying data communicates false confidence.

  4. Match measurement rigor to decision stakes and claim size. Reserve the most rigorous (and costly) measurement, including comparison groups and third-party validation where feasible, for programs being scaled, replicated, or cited prominently in external disclosures.

  5. Build feedback loops, not just reporting cycles. Measurement results should reach the people who can change program design, on a cadence that allows actual adjustment, not only an annual report that arrives after decisions for the next cycle have already been made.

  6. Calibrate external claims to evidence strength. Distinguish, in external communication, between outputs, outcomes with contribution evidence, and outcomes with attribution evidence, and avoid implying a stronger evidentiary basis than the underlying measurement supports.

  7. Reduce burden on grantees and partners. Where the organization funds or partners with third parties, standardize on widely recognized metrics where possible rather than imposing bespoke reporting requirements that duplicate what partners already report to other funders.


A theory of change is still a strong starting point for designing social impact and purpose programs. The problem is not using it; the problem is not going beyond it. As regulators, investors, and the public look more closely at corporate impact claims, the strongest organizations will be the ones that move from planning to real measurement. That means choosing frameworks that fit their maturity and decision-making needs, investing in reliable data systems, and using evidence to improve programs, not just to defend them. When measurement is done this way, it is not just another compliance burden. It helps social impact work improve over time.



 
 
 

Comments


What's the power of your purpose?

  • LinkedIn
  • Youtube

33 Irving Place

New York, NY 10003

​

617-840-4789

Subscribe to Purposeful Connections

Clean-Creatives-Seal[OFF-WHITE].png

2025 by Carol Cone ON PURPOSE LLC

bottom of page