Universität Potsdam
Forecasting the success of environmental and sustainability activities in international development using language models
Abstract
dc:description.abstractInternational aid and cooperation improve the lives of the poorest in developing countries and help safeguard the environment and promote sustainability. However, international aid activities sometimes fail to achieve their objectives. Few attempts have been made in the literature to create models that forecast the success of international aid activities, and none focus solely on environmental outcomes. This thesis produces a forecasting system for various metrics measuring the success of international aid activities at the time of evaluation using data from the International Aid Transparency Initiative (IATI) database, combining classical statistical methods with modern language model techniques. Novel techniques were applied to improve forecasting skill, including using the reasoning and information-gathering abilities of large language models (LLMs) to improve forecasts, introducing LLM summaries of various dimensions of activity documents, and defining a novel narrative similarity grade benchmark. Forecasting success ratings on a scale from 1 to 6, narrative forecasts, cost-effectiveness forecasts, and finally true/false outcome tag forecasts were assessed. While no methods showed consistently high accuracy in forecasting overall success, statistical models could reliably rank above-chance which activities were more likely to succeed for activities with the same reporting organization and start year (“within-group pairwise ranking”). The full forecasting system outperformed what could be extrapolated from the stated risks in activity documents alone for overall evaluations in tests against a held-out test set of 200 latest-starting evaluated activities in a dataset of 1181 environmental and sustainability-improving activities restricted to 4 reporting organizations. Compared to the baseline of 50.7% pairwise ranking for the risks extrapolation, the chosen statistical model with all input features reached 59.7% [95% CI: 51 %, 66 %] on the test set. Across 14 binary outcome tags, statistical models achieved an average pairwise ranking of 60 % (range: 48 %–77 %) for the 45% of pairs with differing outcomes averaging two chosen methods on the test set. This work lays the foundation to improve decision making wherever there is a large collection of pre-intervention description documents and matched post-activity evaluation documents.
Degree
thesis:*- Level thesis:degree_level
- master
- Grantor dc:publisher
- Universität Potsdam
- Year
- 2026
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Rivers, Morgan
- Contributors dc:contributor
-
- Kuhlicke, Christian
- Kuznetsov, Ivan
Subjects
dc:subject × 11Rights
dc:rights- Statement dc:rights
-
- CC-BY - Namensnennung 4.0 International
Identifiers
dc:identifier.*- Repository record source_url
- https://publishup.uni-potsdam.de/frontdoor/index/index/docId/70771
- OAI identifier oai:identifier
- oai:kobv.de-opus4-uni-potsdam:70771