Moving Beyond the Hype: The 2026 Generative AI Evaluation Playbook for Global Development

Generative AI is expanding rapidly across low- and middle-income countries, powering everything from interactive mathematics tutors to specialized agricultural advisory tools for farmers. While early data points to clear development gains, unverified outputs also bring a real risk of misinformation and systemic harm.

To date, evaluating these tools has been highly fragmented. Technical teams tend to focus strictly on system performance, often ignoring long-term human impacts. Meanwhile, social impact evaluators focus heavily on human outcomes but frequently neglect the underlying nuances of the technology.

To bridge this deep divide, the Center for Global Development (CGD), alongside The Agency Fund and IDinsight, convened 30 cross-disciplinary experts spanning computer science, economics, gender studies, and international development. The result of this collaboration is the Generative AI Evaluation Playbook, a framework designed to establish a standard set of evaluation practices for AI in the social sector.

The playbook breaks the evaluation process down into a comprehensive, four-tier framework. Here is a look at what needs to be evaluated at each level, why it matters, and the absolute baseline required for a Minimum Viable Evaluation (MVE).

The Four Levels of Generative AI Evaluation

  [L4] IMPACT EVALUATION       --> Does the product improve development outcomes?
         ▲
  [L3] USER EVALUATION         --> Does it impact users' thoughts, feelings, and behaviors?
         ▲
  [L2] PRODUCT EVALUATION      --> Does the overall product engage and retain users?
         ▲
  [L1] MODEL EVALUATION        --> Does the AI system perform as intended?

Level 1: Model (System) Evaluation

  • The Core Question: Does the AI system perform as intended?

  • What is Evaluated: The full collection of data, underlying models, and software processing inputs to produce predictions or actionable advice.

  • Why it Matters: Generative AI tools can sound remarkably fluent and persuasive while delivering completely inaccurate or harmful information. In high-stakes fields like healthcare or education, unverified outputs can cause direct, real-world harm. Evaluating at this stage prevents costly misalignment further down the road.

  • Who Leads: AI and Machine Learning Engineers, supported by domain experts and user researchers.

  • Minimum Viable Evaluation (MVE):

    • Establish 2 to 3 performance rubrics with at least one robust safety or guardrail metric.

    • Set a clear success criteria or threshold that must be passed prior to deployment.

    • Build a Golden Dataset containing 30 to 50 items representing diverse, realistic user interactions to test the system against.

    • Create an expert review process to check system responses as configurations are updated.

Level 2: Product Evaluation

  • The Core Question: Does the overall product engage and retain users?

  • What is Evaluated: Actual user uptake and user journey metrics, such as a patient’s regularity of interaction with a health chatbot.

  • Why it Matters: An AI model that produces perfectly accurate responses is entirely useless if it fails to engage its target audience. Teams must track behavioral signals like activation, engagement, and retention to ensure the tool fits into the user’s daily workflow.

  • Who Leads: Product Managers, supported by Data Scientists.

  • Minimum Viable Evaluation (MVE):

    • Instrument the digital product to capture user events automatically.

    • Produce two distinct metrics from this data: activation (using the tool once) and retention (using it repeatedly).

    • Speak directly with users and analyze user data to find drop-off points or friction.

    • Run simple A/B tests to evaluate if product updates successfully improve these baseline metrics.

Level 3: User Evaluation

  • The Core Question: Does the product impact users’ thoughts, feelings, and behavior towards the development outcome?

  • What is Evaluated: Shifts in user knowledge, attitudes, decision-making, and self-reported behaviors.

  • Why it Matters: Strong product engagement at Level 2 does not automatically mean a user’s life is improving. A student might use a tutoring app heavily without actually absorbing the material, or an unhealthy eater might chat with a bot daily without altering their diet. Level 3 tracks intermediate indicators, allowing teams to iterate rapidly and confirm the tool is on the right path before funding an expensive, full-scale impact assessment.

  • Who Leads: User Researchers and Behavioral Scientists.

  • Minimum Viable Evaluation (MVE):

    • Target the most decision-relevant cognitive or behavioral outcome from the project’s Theory of Change, and track at least one early-warning indicator of harm.

    • Pair one automated behavioral or trace metric with a brief, self-reported user survey.

    • Conduct a minimal external check to verify that on-platform actions genuinely link to real-world outcomes.

Level 4: Impact Evaluation

  • The Core Question: Does the product improve development outcomes?

  • What is Evaluated: Net changes in objective, real-world welfare metrics such as verified learning, household income, productivity, or morbidity rates.

  • Why it Matters: For international funders and policymakers, anecdotal success or high engagement data is not enough. Credible, rigorous evidence proving that a product moves the needle on development outcomes is essential for making scaling and investment decisions. Level 4 evaluations isolate the true causal effect by comparing users against a rigorous counterfactual.

  • Who Leads: Policy Researchers, Economists, and Social Scientists, ideally working alongside independent, external evaluators.

  • Minimum Viable Evaluation (MVE):

    • Execute an impact evaluation utilizing a clear counterfactual and a large enough sample size to reliably measure outcomes, including across sub-populations like gender and geography.

    • Maintain strict version control throughout the evaluation, testing either a single frozen product version or a limited, highly controlled number of versions to prevent continuous product updates from breaking the research design.

    • Fully transparently account for all data collection costs.

Interconnected Tools for Continuous Assessment

The playbook emphasizes that these four levels do not exist in isolation. Teams should continuously deploy cross-cutting methodologies to ensure structural integrity across the entire lifecycle:

  • Process Evaluations: These check if the right operational actors are doing the right things at the right time, ensuring the intervention is being implemented exactly as intended.

  • User Research: Systematic user studies—including qualitative interviews, workflow observations, and cognitive interviewing—should be used to build Level 1 golden datasets and design Level 3 surveys.

  • Risk Assessment and Mitigation: From addressing algorithmic hallucinations and user dependency to safeguarding sensitive personal data, teams must embed explicit safety guardrails across all four levels.

A Living Framework for the Tech Ecosystem

Because artificial intelligence evolves at a breakneck pace, the Generative AI Evaluation Playbook is designed as a living document. Co-chaired by Han Sheng Chia and Markus Goldstein of the Center for Global Development, alongside Temina Madon of The Agency Fund, the working group explicitly encourages development practitioners, data scientists, and engineers to contribute ongoing amendments and real-world case studies to the framework.

By adopting this standardized, multi-tiered approach, the global development community can move past tech-optimism hype and build responsible, evidence-based AI tools that create lasting human impact.

Recommended Posts