Nonprofit and Global Operations in an AI-Driven World
Table of Contents
Two numbers tell the real story of nonprofits and AI, better than any “AI for good” slogan. A late-2025 study of 346 organizations found that 92 percent use AI in some way. Only 7 percent say it made a real difference to what they can actually get done. The humanitarian field shows the same split from a much bigger group: a survey of 2,539 workers across 144 countries found 93 percent using AI tools, and only 8 percent working somewhere AI is truly built into how the organization runs.
So adoption is not the question anymore. The real question is the gap between having the tools and getting results from them. Most of the sector is stuck in that gap right now.
This gap is familiar. It’s the same one I spend most of my time on in other fields, just wearing different clothes. The tools work. The person using them is capable. The organization still doesn’t improve, because the improvement has to cross a line the tool never touches. In global programs, that line sits in a specific place. Pointing to it is what this note is about.
The pattern is not a nonprofit problem
It helps to say first that the sector is not behind. Business AI adoption ran from about 55 percent in 2023 to nearly 80 percent in 2025, and the plateau shows up there too. A widely cited McKinsey figure has 88 percent of organizations regularly using AI in at least one function, with only about a third past experimentation into genuine scale. Everyone is stuck in roughly the same place.
So the interesting question is not why nonprofits lag, because they do not. It is why near-universal adoption produces so little organizational change, in any sector. My answer, from years of watching it, is that most AI gets added inside a single function, and most of the value people expect from it lives between functions. The tool speeds up a step. The problem was never the speed of the step. It was the transfer to the next one.
Global and humanitarian programs are worth studying closely because they make this general failure unusually visible, and unusually expensive.
Where the boundary sits in a global program
A global campaign or a multi-country response is not one organization doing one thing. It is a center that plans and funds, and a set of regional or local actors who deliver in contexts the center does not directly see. The work that matters most, and fails most often, is the transfer between those two levels.
I saw this directly while advising the WASH (water, sanitation, and hygiene) communications team on multi-region campaign strategy, working with their Comms Director on how campaigns ran across fifteen countries. The recurring finding was not that weaker regions needed better assets. It was that regional teams were already making good local adaptations, and those adaptations had no route back into central planning. The center could not see what the field knew. Performance gaps between regions were usually broken feedback paths, not weak local execution. Once a return path existed, engagement and reach rose, and the regional differences that were working stayed intact rather than being standardized away.
That is the boundary AI is now being added around, and the direction it is being added from matters enormously.
AI is being installed on the wrong side of the gap
Here is the structural risk the adoption statistics hide. In the humanitarian data, the same survey that found 93 percent individual use also found that three quarters of respondents were in the Global South, and that adoption is not following a Global North to South diffusion pattern. Local practitioners are adopting fast, from the bottom up, mostly with commercial tools like ChatGPT, Copilot, and Claude, and mostly without organizational support. Sixty-four percent of organizations provide little to no AI training. Only around a quarter have any formal AI policy.
Meanwhile the money and the systems remain centralized. Only 3.6 percent of humanitarian funding went directly to local and national organizations in 2024, against a target of 25 percent set years earlier and missed. Larger organizations, with budgets above a million dollars, adopt formal AI at nearly twice the rate of smaller ones. The infrastructure, the governance, and the funded tooling accrete at the center. The contextual knowledge lives at the edge.
Put those two facts together and the risk is clear. When a large organization does invest in AI infrastructure, it naturally builds it where its systems and budget already are, at the center. That hardens the exact boundary that was already the problem. The center gets faster at planning, reporting, and processing from the data it can already see. The field’s local knowledge is no better represented in that data than it was before. You have automated the powerful half of a broken handoff and left the other half exactly where it was, now facing a counterpart that moves faster and asks for more.
This is the specific way AI can make a global program worse while every dashboard says it is improving. Central throughput goes up. The feedback path does not get built, because a feedback path is unglamorous plumbing and AI budget flows toward visible capability. The gap between what the center believes and what the field knows widens, and it widens faster.
Why the dashboards will not warn you
The measurement problem here is the one I find in most sectors, sharpened by distance. Each level reports on what it can see. The center’s metrics, throughput, cost per output, campaigns shipped, will look better after a central AI investment, because those are exactly the quantities central automation improves. The field’s reality is not in those metrics. It never was. That is why the handoff was invisible in the first place.
Before diagnosing any performance gap in a program like this, it is worth confirming the gap is even real, because a meaningful share of apparent variance between regions turns out to be definitional. Two country teams count “reach” or “an active case” differently and call them by the same name. AI does not fix this. It industrializes it. A model trained or prompted on the center’s definitions will confidently apply them to field data that meant something else, and produce fast, clean, wrong aggregates that are harder to question precisely because they arrive with the authority of a system.
The second failure: the skill leaves the building
There is a second way AI degrades an operation, and it is slower, quieter, and harder to reverse than a widened boundary. It is what happens to the people.
The distinction that matters is between using AI to build a better system and using AI to do the work directly. The first leaves a system behind that the team understands and can maintain. The second leaves nothing behind except a dependency, and it slowly removes the team’s ability to do the thing at all. Most organizations, under time and funding pressure, drift toward the second without deciding to.
The clearest evidence now comes from medicine, because the outcome is measurable. A 2025 study in The Lancet Gastroenterology and Hepatology followed nineteen experienced endoscopists, each with more than two thousand procedures behind them, across four centers that introduced AI polyp detection. In the three months before AI, their unaided detection rate was 28.4 percent. Three months after routine AI use, their unaided rate had fallen to 22.4 percent, a six-point absolute drop, roughly a fifth of their skill, gone after a single quarter of leaning on the tool. These were not novices. The researchers called it the first real-world evidence of automation-induced deskilling linked to patient outcomes.
This is the GPS effect, and most people recognize it from their own lives. Turn-by-turn navigation is genuinely useful, and after a few years of it, many people can no longer find their way across a city they have lived in for a decade. The skill did not fail loudly. It was never exercised, so it quietly left. The danger is specific: you do not notice the loss until the moment the tool is unavailable or wrong, which is exactly the moment you needed the skill.
The pattern is not confined to medicine or driving. A study of an accounting firm, titled The Vicious Circles of Skill Erosion, found that reliance on automation bred complacency and eroded staff competence to the point that, when the system was removed, the firm discovered its employees could no longer do the work unaided. A controlled study of 666 participants found a significant negative relationship between frequent AI use and critical thinking, with cognitive offloading as the mechanism. An MIT Media Lab study reported the same direction. The underlying finding across this literature is consistent and uncomfortable: AI often automates the very tasks through which people built and maintained the skill in the first place, so the tool erodes the capability it was meant to augment.
Two automation biases make this worse, and both are well documented. The first is automation bias: people accept the AI’s output without independent checking because it arrives looking polished and authoritative, even when it is wrong. The second is automation complacency: as the tool proves reliable, people relax their vigilance, monitor less, and are slower to catch the errors that do occur. In a context that demands neutrality and accountability, which humanitarian work does, a confident wrong answer that nobody is any longer equipped to question is a serious operational risk, not a convenience.
The other edge: the expert who replaces the team that would have built it
The double-edged part is this. At the same time AI is quietly deskilling the operational staff, it is making a single expert dramatically more capable, capable enough to do alone what used to require a team. One experienced person with good tools can now stand up an analysis, a workflow, a model that would previously have justified hiring several people to build and run.
That sounds like pure efficiency, and in the short term it often is. The risk is what it does to the organization’s depth. If the expert builds the system and the team merely operates it through the AI, the organization has concentrated its real capability in one person and a tool, and hollowed out the layer of people who would otherwise have learned the craft by doing it. When the expert leaves, or the tool changes, or the vendor’s model drifts, there is no bench. Nobody rebuilt the skill on the way up, because the AI did the parts that used to teach it.
So the sword cuts both ways at once. Naive automation deskills the people who run the work, while expert-plus-AI removes the reason to employ and train the people who would build it. An organization can end up faster and more fragile in the same motion, dependent on a capability it can no longer produce or even fully evaluate from the inside. None of this shows up on a performance dashboard until something breaks, which is the signature of every failure worth catching before it does.
What actually helps
None of this is an argument against AI in global programs. It is an argument about sequence and location, and it is the same argument I would make about automating a clinic or a warehouse, with the stakes raised.
Build the feedback path before the central capability. The highest-value AI application in a multi-level program is usually not faster central planning. It is making the field’s local knowledge legible to the center without flattening it: structured capture of what regional teams are actually deciding and why, summarization that preserves the local reasoning instead of averaging it out. That is a genuinely good use of language models, and almost nobody is funding it, because it does not look like transformation on a slide.
Make the data trustworthy before you make it fast. In any context that requires neutrality or accountability, and humanitarian work requires both, automating a lookup against data you have not verified produces confident, wrong answers that are worse than slow correct ones. Reconcile the definitions across regions first. Then automate. Reversing that order is the most common and most expensive mistake I see.
Fund the edge, not only the center. The adoption data already shows local actors moving faster than their institutions. The governance gap, 84 percent of nonprofits saying they need funding to develop AI while only 17 percent report their funders have engaged them on it, is most acute exactly where the contextual knowledge lives. An organization serious about not widening the boundary will put AI capability, training, and policy where the field is, not only where the budget already sits.
Use AI to build systems, not to replace the doing of the work. This is the guard against deskilling, and it is a design choice made at the start, not a policy bolted on later. Aim the tool at constructing something the team then owns and understands, rather than at quietly performing the task in the team’s place. Keep some deliberate unaided practice in the workflow, the way the deskilling research recommends, so the underlying skill stays exercised and the staff retain the ability to judge whether the AI is right. An operation that can still function, and still evaluate its tools, on the day the tool is unavailable is worth more than one that is slightly faster until that day arrives.
Limitations
The obvious risk in an argument like this is that it makes every AI project in the sector look like a centralization problem. Many are not. Plenty of nonprofit AI use is a development officer drafting a grant narrative faster, which is a genuine and uncomplicated gain with no handoff involved. This note is about a specific configuration: multi-level programs where a center and a set of local actors are separated by a boundary that already carried the failures. Where that structure is absent, reach for a simpler explanation first.
The evidence here is of two different grades, and it is worth being clear about which is which. The adoption figures come from sector surveys, self-reported and collected under difficult conditions. They are directionally strong and they agree across independent sources, which is why I trust the shape they describe, but they are not precise, and anyone citing 92 versus 93 percent as if the decimal mattered has missed the point. The deskilling figures are firmer: the colonoscopy result is a peer-reviewed observational study with a measured patient-linked outcome, and the accounting-firm and critical-thinking findings are published research. Even so, the deskilling studies come from medicine, accounting, and driving, not from humanitarian programs specifically. The mechanism, a skill decaying through disuse, is general enough that I am confident it transfers, but the exact magnitude in a nonprofit setting has not been measured, and I am not going to pretend otherwise.
The larger point survives both caveats. Adoption is near total and impact is rare. In global programs the reasons are structural and locatable: the tool went in on the side of the boundary that already held the power, and the way it was used slowly removed the organization’s own ability to do, judge, and rebuild the work. Both are the kind of failure that stays invisible on every dashboard until the day it is expensive.