A Framework for Comparing Global Problems in Terms of Expected Impact

Core claim

The problem you choose may be the largest determinant of your career’s social impact. A useful first approximation is to compare problems by scale, neglectedness, and solvability, then adjust for personal fit. Quantifying these factors can expose very large differences and hidden assumptions, but the resulting scores are rough comparisons, not precise measurements or automatic answers.

What the framework is trying to estimate

  • The target quantity is the expected good accomplished by one additional unit of resources devoted to a problem. The additional resource could be a person’s year of work, a dollar of funding, or another relevant input.

  • Estimating that quantity directly is usually too difficult, so the framework decomposes it into three factors:

    1. Scale: if the problem were solved, how much better would the world become?
    2. Solvability: if direct resources devoted to it were doubled, what fraction of the problem would we expect to solve?
    3. Neglectedness: how few resources are already devoted to solving it?
  • In multiplicative form, the rough idea is:

    marginal impact ≈ scale × solvability × neglectedness

  • For a personal career decision, personal fit can be added as a separate adjustment: how likely is this particular person, given their skills and motivation, to excel at useful work in the area?

  • The framework originated at GiveWell and was further developed by 80,000 Hours with staff at Oxford’s former Future of Humanity Institute. It is presented as one tool within a broader problem-comparison process.

Start by defining the problem consistently

  • Before assigning scores, specify exactly what is included. For “global health,” for example, decide which diseases and countries count.
  • A narrow problem can look artificially attractive next to a broad one. “Combating malaria” selects a particularly promising part of global health, while “global health” averages over many interventions; likewise, improving health in one carefully selected country may look better than improving health everywhere.
  • This is not necessarily an error—the narrower result may be genuinely useful—but inconsistent scope makes comparisons misleading and creates opportunities to manipulate scores.
  • Scale and neglectedness must refer to the same unit of analysis. Measuring the scale of a broad category but the resources devoted only to a narrow subproblem will exaggerate its priority.

Why the scores are logarithmic

  • Differences between causes can span several orders of magnitude. The article contrasts roughly $300 billion in annual global-health spending with under $100 million directed at factory farming at the time—a difference of more than 1,000× in measured neglectedness.
  • Each two-point increase on the framework’s 0–16 scales represents a factor of ten. A score of 6 therefore means roughly 10× more on that factor than a score of 4.
  • Because logarithms turn multiplication into addition, the three component scores can be added instead of multiplied.
  • Only differences between scores carry meaning. The absolute numbers are a convenient representation, not direct measurements of moral value.

1. Scale

Definition

  • Scale asks: If the entire problem were solved, by how much would the world improve?
  • A problem can be larger because it affects more beings, affects each one more intensely, lasts longer, or has important indirect effects.
  • 80,000 Hours uses a broad conception of wellbeing—including health, happiness, meaning, and relationships—and, in this article, a longtermist perspective that gives substantial weight to effects on future generations.
  • The framework does not require those values. Someone with different moral commitments can substitute a different definition of scale, but should make the choice explicit.
  • Information gained, movement-building, spillovers into other problems, and coordination benefits can also contribute to scale when they are genuinely instrumental to future impact.

Yardsticks

  • Some within-domain comparisons are comparatively concrete: the article uses the share of global health loss, measured in QALYs, to compare cancer and malaria.
  • Long-run and cross-domain comparisons need yardsticks—measurable proxies expected to correlate with social value. Examples include:
    • healthy life-years saved per year;
    • income gains among the world’s poorest people;
    • increases in global economic output;
    • reductions in existential risk or increases in the expected value of the future.
  • On the article’s illustrative scale, every two points corresponds to 10×:
Scale scoreIllustrative health equivalentIllustrative existential-risk reductionExample given
1610%Eliminate the risk from both nuclear war and pandemics
141 billion QALYs1%Eliminate extreme poverty
12100 million QALYs0.1%Cure cancer
1010 million QALYs0.01%Increase aid by one-third and spend it on cash transfers
81 million QALYs0.001%Eliminate land-use restrictions in major US cities
6100,000 QALYs0.0001%Remove five minutes of needless daily red tape for US teachers
410,000 QALYs0.00001%Identify all risky asteroids
21,000 QALYs0.000001%Turn 10,000 people vegan
0100 QALYs0.0000001%Save three lives
  • The table is a judgment-laden calibration device, not an empirically established exchange rate. Comparisons are most robust within a single yardstick; trading QALYs, income, growth, and existential-risk reductions against one another depends heavily on values and worldviews.
  • Because the scale is logarithmic, if an intervention has effects in multiple columns, its largest effect usually dominates rather than the columns being naively added together.
  • The authors explicitly note that the yardsticks and their particular tradeoffs were not fully up to date even at the time of the page’s last revision, though they still endorsed the broad approach.

2. Neglectedness

Definition and rationale

  • Neglectedness asks how many people or dollars are already directly dedicated to the problem.
  • It matters because of diminishing returns: the best opportunities tend to be taken first, so an additional worker or dollar often accomplishes less in a crowded field than in an overlooked one.
  • A successful intervention may therefore cease to be the best destination for new funding once governments, foundations, or markets have already scaled it. The article gives mass childhood immunisation as an intervention that is highly effective but also heavily pursued.
  • New areas can offer value of information: trying an untested approach teaches the wider community whether it is more solvable than expected.
  • Neglectedness is only a proxy. A problem may be neglected for a good reason—because it is intractable—or for a bad reason—because the affected beings are politically invisible or morally undervalued. The other factors are needed to distinguish these cases.

Crowdedness rubric

Neglectedness scoreDirect annual spendingFull-time staffActive supporters
12$100,000 or less1 or fewer1,000 or fewer
10$1 million1010,000
8$10 million100100,000
6$100 million1,0001 million
4$1 billion10,00010 million
2$10 billion100,000100 million
0$100 billion1 million1 billion
  • When inputs imply different scores, use the lowest score, corresponding to the resource that makes the field most crowded.
  • Avoid extreme neglectedness scores without a comprehensive search. Obscure work may exist outside the researcher’s awareness; 80,000 Hours therefore generally assumes at least $1 million is already directed at a problem unless shown otherwise.

Direct, indirect, and future effort

  • Direct effort comes from actors explicitly trying to solve the problem; indirect effort comes from adjacent research, commercial incentives, or other goals that nevertheless produce relevant progress.
  • Anti-ageing research illustrates the difficulty: little work may explicitly aim to stop ageing, while enormous biomedical research programs advance many of the underlying capabilities.
  • The framework usually counts only direct effort under neglectedness and lets solvability absorb the indirect-effort correction. If adjacent actors have already tried the promising ideas, doubling direct effort will solve less of the remaining problem.
  • Future crowding also matters. A currently obscure topic may be poised to attract substantial funding, reducing the marginal value of entering it. The article offers no general solution beyond considering resource trends and avoiding implausibly high neglectedness estimates.

Diagnostic questions

  • Is there a reason markets, governments, or other altruistically motivated people will fail to solve the problem?
  • Is the relevant research field new or located between established disciplines, where academia may overlook it?
  • If you do not enter, how likely is someone else to take your place?
  • Will working on the problem generate information useful for comparing it with other causes?

3. Solvability

Definition

  • Solvability asks: If direct effort doubled, what fraction of the remaining problem would we expect to solve?
  • A problem can be enormous and neglected yet still be a poor priority if additional effort has almost no chance of making progress. Ageing is used as the example: it causes a large fraction of ill health and receives relatively little direct prevention research, but may be exceptionally difficult to change.

Rubric

Solvability scoreExpected fraction solved by doubling direct effort
8100%
610%
41%
20.1%
00.01%
  • Look for interventions with rigorous evidence, promising ideas that can be tested cheaply, theoretical reasons for optimism, or low-probability approaches with very large upside.
  • Evaluate both the size of the upside and its probability, updating from a skeptical prior as evidence accumulates.
  • Incremental and “all at once” strategies can be compared in expected-value terms: a 10% chance of solving the whole problem receives the same solvability score as certainty of solving 10%, as a simplifying approximation.
  • This is normally the hardest component because it predicts the future. Established fields may have usable cost-effectiveness records; novel research and advocacy often require expert judgment and explicit uncertainty.

Interpreting the combined score

  • Adding the component scores produces a rough estimate of marginal impact. The article’s sanity-check table suggests:
Total scoreIllustrative impact of one extra person
281 million QALYs saved per year, or existential risk reduced by 0.001%
2410,000 QALYs saved per year, or existential risk reduced by 0.00001%
20100 QALYs saved per year (about two lives), or existential risk reduced by 0.0000001%
  • These absolute conversions are explicitly described as extremely approximate. Relative comparisons are more defensible than claims that a worker will literally produce the stated result.
  • Internally, 80,000 Hours treated a four-point gap as enough for reasonable confidence that one problem was more pressing; a three-point gap or less looked like a close call.
  • Raw calculations sometimes imply 10,000× differences between causes. The authors do not take those magnitudes literally because mistakes, omitted effects, model uncertainty, and the unpredictable indirect value of work impose a ceiling on confidence.

Personal fit

  • Personal fit asks how likely a person is to excel in an area given their skills, resources, knowledge, relationships, and motivation.
  • It matters because top performers in some fields may produce 10–100× the impact of the median performer, and a person who cannot remain motivated may contribute little even to a highly ranked problem.
  • Suggested questions:
    • Which forms of career capital are most valuable and where are they relevant?
    • How motivated would you remain in the area?
    • Which concrete roles are available, and are you likely to excel in them?
Personal-fit scoreInterpretation
4Exceptionally suited and potentially capable of becoming a field leader
2Reasonable fit, with relevant skills and motivation
0Actively poor fit because of missing skills or motivation
  • Personal fit matters especially in high-variance roles such as research and entrepreneurship. It matters less for earning to give, where the contribution is money rather than unusually specialised labour.
  • Existing experience can be overweighted because of sunk-cost effects, while people’s capacity to develop new interests is often underestimated.
  • A cause contains many possible roles. Poor fit for fieldwork need not imply poor fit for the cause if research, policy, operations, communications, or another route is available.

What the framework leaves out

  • A complete career decision also considers:
    • the influence available in a specific role;
    • career capital gained;
    • the value of information from trying the option.
  • Individual and group decisions differ. One person may concentrate on one or two causes, while a coordinated community should probably use a portfolio approach:
    1. estimate an ideal allocation of people and resources across issues;
    2. identify which way the current allocation should move;
    3. place each person where they have comparative advantage.
  • Coordination can also justify compromise or moral trade between people with different values.

Comparison with conventional cost-effectiveness analysis

  • Standard cost-effectiveness analysis compares observed interventions using a common outcome, such as QALYs per dollar. When reliable evidence and a common unit exist, it is often the more direct method.
  • Cross-domain comparisons require controversial conversion rates—for example, how many QALYs equal one tonne of carbon avoided—or monetisation through cost-benefit analysis.
  • The problem framework is especially useful when direct cost-effectiveness evidence is unavailable, as in:
    • political advocacy with shifting circumstances;
    • original research with unknown timelines;
    • fields with no established or well-studied interventions.

Strengths and weaknesses of quantification

Advantages

  • Makes scope and orders-of-magnitude differences harder to ignore.
  • Forces assumptions into the open and tests whether the analyst actually understands the problem.
  • Produces reasoning that other people can inspect, challenge, and improve.
  • Offers a common language for comparing otherwise very different opportunities.

Main danger

  • The estimates are deeply uncertain and often non-robust: small changes in assumptions can reverse the ranking.
  • A tidy numerical output can conceal an incomplete model and create false confidence.
  • 80,000 Hours therefore combines scores with qualitative evidence, “cluster thinking,” common-sense checks, and sensitivity to worldview assumptions rather than mechanically selecting the top number.

Practical takeaway

  • Use the framework to structure thought, expose disagreements, and notice very large differences—not to manufacture precision.
  • Define causes at comparable levels, estimate each component independently, record uncertainty, test sensitivity, and scrutinise surprisingly high results.
  • Treat the final ranking as one input into an all-things-considered decision that also includes personal fit, role-specific leverage, downside risks, learning value, and coordination with others.

Notes & Observations

  • “Importance, neglectedness, and tractability” is often presented as a simple checklist, but this article’s version is more disciplined: neglectedness and solvability jointly estimate the return to a marginal doubling of direct effort, reducing the risk of double counting.
  • The framework’s most valuable product may be the decomposition rather than the score. Two people who disagree can locate whether the disagreement concerns values, empirical scale, future effort, tractability, or personal fit.
  • Its central tension is productive but unresolved: quantification is most needed where intuition performs badly, yet cross-cause comparisons are also where the underlying quantities are least measurable.
  • Figures and calibration examples reflect a page last substantively updated in 2019 and should not be reused as current estimates without fresh research.