The Fundraising
Judgment Benchmark

363 advancement and nonprofit professionals judged fundraising recommendations without knowing who wrote them: human or AI. Their verdicts map exactly where AI can help, where human judgment must lead—and everything in between.

By Cherian Koshy, ACFRE, CFRE, CAP
Executive Vice President, Market Insights, Kindsight

2027 Fundraising Judgment Benchmark

AI has arrived in the nonprofit sector as quickly as it has everywhere else, and many practitioners are putting it to work before knowing how good its advice is or how much it should be allowed to do. Where does AI genuinely help? Where does it fall short? And when its recommendation is strong, should it have permission to act autonomously? The Fundraising Judgment Benchmark put these questions to 363 nonprofit professionals. Across 1,452 blinded comparisons, each pairing a practitioner’s recommendation with an AI-generated one, they judged which was stronger, which was riskier, and how much authority AI should have. The results were enlightening.

The study at a glance

363

Nonprofit professional respondents

1452

Blinded recommended comparisons

51.4% vs. 48.6%

Human vs. AI preference overall

98.3%

Judgments requiring human oversight

RESEARCH STATUS

This is an industry benchmark, not a peer-reviewed academic article. Its methods support careful conclusions about how this sample judged these specific recommendations. They do not establish universal superiority for humans, AI, or any vendor.

Five discoveries that should change how fundraising organizations govern AI

The core conclusion of this research is that AI may inform far more decisions than it should execute. The evidence supports variable authority by workflow, not a single organization-wide rule. The central management question is no longer “Should we use AI?” It is “What authority should AI have in this workflow, under these conditions, and who owns the outcome?” Let’s dive deeper into five critical discoveries for fundraising organizations:

  1. Neither human nor AI recommendations were preferred overall; the workflow decided. Across all comparisons, the split was effectively even. But the scenario results told a different story: after correction for multiple testing, eight of the 16 scenarios showed a statistically reliable preference, four for AI and four for the practitioner. All eight remained significant in the sensitivity analysis that removed SurveyMonkey quality-flagged responses. The other eight showed no reliable preference.
  2. AI recommendations were preferred when they paired scale with added safeguards. The strongest AI results considered more variables, verified their claims, added privacy gates, and fell back on conservative options when data were weak.
  3. Practitioner recommendations were preferred when the decision required interpreting ambiguity or owning a consequential outcome. Human judgment was favored in scenarios involving donor intent, relational uncertainty, high-stakes qualification, and deployment decisions.
  4. Preferring an AI recommendation did not mean granting AI permission to act. Only 1.7% of judgments supported AI execution without individual human review. Even when respondents chose the AI recommendation as stronger, they still did not want to give AI the authority to act without human oversight.
  5. Respondents treated risk as part of a recommendation’s quality. In comparable decisions, the rejected recommendation was judged riskier 50.5% of the time. The selected one was judged riskier only 5.9% of the time. This suggests perceived risk was one of the reasons a recommendation was judged weaker.

Why this study is different

Nonprofit AI research is growing quickly, but much of it answers a different question. Adoption benchmarks ask who is using AI, where it is being used, and whether organizations are seeing value. Donation-appeal studies test whether AI-written messages persuade donors. Targeting studies evaluate efficiency and response. Broader human-AI research asks when people and models perform better together.

Those are important questions. But they do not tell a chief development officer, advancement services leader, prospect researcher, or gift officer what authority an AI system should have inside a live fundraising workflow.

Recent work illustrates the distinction. Caffier, Stavrova, and Kleinberg (2026) found that large language model donation appeals could outperform human-authored appeals in two preregistered experiments. Vaccaro, Almaatouq, and Malone (2024), reviewing 106 experiments across domains, found that human-AI combinations were not automatically superior and that decision tasks were especially difficult. The 2026 Nonprofit AI Adoption Report from Virtuous and Fundraising.AI benchmarked adoption, governance, readiness, and impact across 346 organizations. This benchmark sits alongside those streams rather than duplicating them. Because the source was hidden and the recommendations were standardized in presentation, respondents judged the substance before learning whether a person or AI produced it.

As of September 2026, our review did not identify another published fundraising study that combines blinded human-versus-AI recommendations, practitioner judges, multiple fundraising functions, explicit risk judgments, and a direct measure of appropriate AI authority. That is a statement about this combination of design features, not a claim that no related research exists.

Looking for more?

Download the Fundraising Judgment Benchmark Executive Summary for 5 critical insights that every nonprofit and advancement leader should know!

Methodology, study design, sample, and validity

Respondents first selected the fundraising function they were best qualified to evaluate. They then reviewed four fixed scenarios from that function. Each scenario presented two unlabeled recommendations: one supplied by an experienced nonprofit practitioner and one generated using AI. Human and AI positions were counterbalanced within every role branch, with two human-first and two AI-first scenarios.

The practitioners behind the human recommendations

The benchmark depended on experienced practitioners being willing to put real professional judgment on the page. We are grateful to the following contributors, who supplied human recommendations and helped ground the scenarios in the realities of fundraising work. One human benchmark recommendation was fielded per scenario; not every contributed response appeared in the fielded set. Affiliations are listed for identification and do not imply institutional endorsement of the report.

Analytic sample

The raw export contained 539 response records. The prespecified inclusion rule required completion of all 16 core items across all four scenarios in the respondent’s selected function. This produced 363 included respondents and 176 exclusions. Every excluded record had completed zero core items in the selected branch, so the analytic sample contained no partially completed branches.

Of the respondents, 62% had 11 or more years in fundraising, advancement, or nonprofit leadership. 70% were at least somewhat familiar with AI tools, while 15% reported using AI regularly.

Statistical approach

Scenario preferences were tested against a 50/50 split using two-sided exact binomial tests with Wilson 95% confidence intervals. Holm correction controlled the family-wise error rate across all 16 scenario comparisons. The overall preference interval was generated by bootstrapping at the respondent level so each person, not each isolated judgment, was the resampling unit. Exploratory models used scenario fixed effects with respondent clustering.

The primary analysis retained the 14 SurveyMonkey quality-flagged respondents, consistent with the prespecified analytic rule. A sensitivity analysis removed those respondents. All eight primary scenario findings remained statistically reliable in that sensitivity analysis.

Study Findings

Cherian Koshy

Click to read a note from author Cherian Koshy, ACFRE, CFRE, CAP, on interpreting the findings in this report

A quick note on interpreting the findings in this report:

The fastest way to misread this study is to ask who won. Across 1,452 blinded comparisons, practitioners selected the human recommendation 51.4% of the time and the AI recommendation 48.6% of the time. The 95% respondent-cluster bootstrap interval for human preference ran from 49.0% to 53.9%, crossing an even split. There was no overall winner.

But the aggregate masks the important result. Preference moved from 12.7% human in trip planning to 81.5% human in restricted-gift ambiguity. In other words, the workflow mattered far more than the label. AI was strongly preferred in some contexts and human practitioners were strongly preferred in others.

The second result is even more consequential. Recommendation quality did not determine decision authority. Only 1.7% of all judgments supported AI execution without individual human review. Even when respondents selected the AI recommendation as stronger, they usually retained human approval or human-led decision-making.

Taken together, these findings point to a simple discipline. The useful question is not whether AI is better than people, but where in the work it adds value and who should own the decision once it does. Practitioners in this study drew that line clearly: they welcomed AI’s help where the task rewarded speed and structure, leaned on human judgment where context, relationships and donor intent were at stake, and kept a person accountable for the outcome almost every time. Read the rest of this report with that lens, workflow by workflow, and you will find a practical map for using AI where it helps while keeping the trust that fundraising depends on firmly in human hands.

Cherian Koshy, ACFRE, CFRE, CAP

Across 1,452 blinded recommendation comparisons, respondents selected the human recommendation 747 times (51.4%) and the AI recommendation 705 times (48.6%). The 95% respondent-cluster bootstrap interval for human preference was 49.0% to 53.9%, which includes 50%. In other words, there was no overall winner.

The aggregate result does not support “AI beat fundraisers” or “fundraisers beat AI.” It supports a more useful conclusion: recommendation quality was conditional on the decision being made.

The workflow decided

Eight scenarios showed a statistically reliable preference after multiple-comparison correction. Four favored AI and four favored the practitioner. The remaining eight did not support a reliable source advantage. The workflow decided: human preference ranged from 12.7% to 81.5%.

AI won by adding constraints, not simply by moving faster

In the following scenarios, AI was chosen as the stronger recommendation.

  1. 87.3% AI—Trip planning with too many prospects. The stronger recommendation widened the trip-planning logic beyond capacity and proximity to include engagement, responsiveness, and the likelihood of securing a meeting.
  2. 79.3% AI—Grateful patient prospect. The stronger recommendation put privacy, timing, care-team, and information-access gates ahead of grateful-patient outreach.
  3. 72.8% AI—Automated stewardship cadence. The stronger recommendation segmented stewardship by data confidence and used broader, accurate language when program tags were stale or incomplete.
  4. 66.3% AI—AI-generated profile from public sources. The stronger recommendation treated the AI-generated profile as preliminary, verified consequential claims, removed unsupported details, and flagged uncertainty.

The pattern is important. AI was not rewarded for being aggressive. It was rewarded when it combined scale with discipline: more relevant variables, explicit safeguards, verification, and a fallback when the evidence was weak.

Human practitioners won where intent, uncertainty, and ownership dominated

In the following scenarios, human practitioners were chosen as the stronger recommendation.

  1. 81.5% human—Restricted gift ambiguity. The stronger recommendation treated “student success” as a question of donor meaning, not merely gift coding.
  2. 79.2% human—Unmanaged portfolio automation. The stronger recommendation proposed a bounded pilot with evidence review, disclosure, escalation, named ownership, and comparative measurement rather than activating automation across the entire unmanaged pool.
  3. 78.0% human—Donor expresses doubt before a planned ask. The stronger recommendation kept the conversation but changed its purpose from proposal delivery to listening, timing, and fit.
  4. 70.7% human—Predictive model recommends upgrade. The stronger recommendation treated a predictive score as a reason for personal qualification, not as permission to automate a $2,500 appeal.

A useful principle for practitioners: Prediction is not permission. A signal can start a governed process without deciding how, when, or whether a fundraiser should act.

The Authority Gap

AI advice was often accepted, but AI authority remained constrained. The table below shows responses about appropriate AI involvement across all 1,452 scenario judgments.

Only 24 of 1,452 judgments (1.7%) supported AI execution without individual human review. In contrast, 77.5% required human approval, human-led decision making, or no AI involvement at all.

This restriction did not meaningfully relax when AI produced the stronger recommendation. When respondents preferred AI, 78.0% still chose approval or a more restrictive authority level. When they preferred the human recommendation, the figure was 77.1%. A scenario-adjusted model found no reliable association between source preference and governance strictness (p = 0.159).

A practical distinction: Accuracy is not accountability. A professional can believe the AI has the better answer and still believe a named person must own the decision.

Risk was already factored into the quality judgment

Across 1,334 comparable judgments, the recommendation a respondent rejected was judged riskier 50.5% of the time. The selected recommendation was judged riskier only 5.9% of the time. The association was strong (Cramér’s V = 0.59, p < .001). The table below shows the relationship between preferred recommendation and perceived risk. The first frontline risk item is excluded because its response options differed.

In practice, “stronger” did not mean merely cleverer or more efficient. Respondents appear to have folded trust, ethics, compliance, operational downside, and donor consequences into their assessment of recommendation quality.

Practitioners wanted controls that let them intervene

When forced to prioritize up to three controls, respondents chose human approval workflow (71.6%), the ability to edit or override the recommendation (64.0%), and source citations or evidence display (40.3%) far more often than confidence scores (6.8%) or audit logs (6.8%). This does not make technical controls unimportant. It shows that practitioners first want mechanisms that let them see, challenge, and stop the system.

Two role differences remained reliable after correcting for multiple tests. Prospect-development respondents were especially likely to prioritize source citations and evidence display, while frontline respondents were more likely to prioritize communication-preference and opt-out controls. The implication is not that each function needs a different ethics policy; it is that effective governance should expose different controls where the workflow risk actually occurs.

Caution was not explained by unfamiliarity

Exploratory scenario-adjusted models found no reliable association between AI familiarity and choosing the human recommendation (p = 0.695), nor between AI familiarity and requiring human approval or a more restrictive authority level (p = 0.136). Years of experience likewise did not reliably predict source preference or governance strictness.

That matters because it is easy to dismiss governance concerns as resistance from people who have not used AI. These results do not support that explanation.

What practitioners said in their own words

The final open-text question received 196 responses, of which 193 contained enough content for descriptive thematic review. Themes overlapped, so the qualitative review should not be read as population prevalence. The recurring ideas were relationships and trust, ask timing, strategic context, donor intent and privacy, data verification, and final human accountability.

“Humans give to and build relationships with humans.” — Anonymous respondent

“Human judgment should remain most central in reading relationship sentiment and timing: knowing when a donor or prospect is ready to be asked, how they’re really feeling about the organization, and what’s happening beneath the surface of a conversation that no CRM field captures.” — Anonymous respondent

How fundraising organizations can use these findings

The study does not produce a universal list of tasks that should be automated and tasks that should never be automated. It supports a better operating principle: assign authority to the workflow, then make that authority visible before deployment.

  1. Define the decision before you debate the technology. “Should we use AI?” is too broad to govern. “Should AI rank a trip list, send a stewardship message, interpret donor intent, or code an ambiguous restriction?” can be answered.
  2. Separate recommendation quality from authority. A model may be excellent at generating options without being entitled to execute them. Treat quality and permission as two different governance decisions.
  3. Put a name on the line. “Human in the loop” is not an accountable operating model. Name the person or role who approves, can override, can pause, and can stop each consequential workflow.
  4. Make approval and override part of the work itself. The top two requested controls were human approval and edit/override authority. If exercising those controls requires leaving the workflow, opening a policy binder, or escalating to a technical team, they are less likely to function when needed.
  5. Show the evidence at the point of decision. Profiles, predictions, and next actions should expose sources, recency, and uncertainty where a fundraiser decides what to do next.
  6. Treat prediction as a lead, not permission. A model flag can trigger research, qualification, or a policy check. It should not automatically become permission to contact, solicit, code, or communicate.
  7. Design a conservative fallback for weak data. The stewardship result is instructive: broader accurate language is better than confident personalization based on stale tags. Every workflow should define what the system does when the evidence is incomplete.
  8. Pilot consequential automation with boundaries. Start with a bounded population, explicit rules, escalation, a named owner, a comparison or holdout group where feasible, and predefined stop conditions.
  9. Measure trust alongside efficiency and revenue. Time saved and dollars raised matter. So do complaints, opt-outs, corrections, overrides, policy triggers, and relationship recovery.
  10. Train staff to challenge both sources. The study surfaced weak human recommendations as readily as weak AI recommendations. Critical review should be triggered by the stakes, not by whether the recommendation came from a person or a model.

Where to begin

Important caution: These are starting postures, not blanket conclusions from one scenario. The recommendation, the quality of the data, the available controls, the donor context, and the organization’s risk tolerance still matter.

A practical framework for AI authority

The most useful governance question is not whether AI is “on” or “off.” It is how much authority the system has in this workflow. The five modes below are designed to make that decision explicit.

Replace a generic loop with named accountability: For every consequential workflow, identify the person or role whose name is on the decision. That person must have the information and the actual power to approve, override, pause, or stop the system. A name without stop authority is not accountability.

Seven questions that assign the authority level

  1. Does the action communicate directly with a donor or prospect? Increase review as personalization and relational stakes rise.
  2. Does it interpret donor intent, sentiment, readiness, or capacity to commit? Keep the judgment human-led unless intent is explicit and consequence is low.
  3. How reliable and current are the data? Require verification or a conservative fallback when evidence is weak.
  4. What happens if the recommendation is wrong? Use more restrictive authority when harm is difficult to reverse.
  5. Are privacy, legal, ethical, gift-policy, or vulnerable-person rules involved? Apply a policy gate before action.
  6. Who is the named accountable person? Do not deploy consequential automation without an owner who can stop it.
  7. Can the organization observe and learn from the result? Define performance, trust, and stop metrics before launch.

Know your starting point: An AI workflow is only as reliable as the operation underneath it. Before assigning authority modes, it helps to know how consistently your team logs research, tracks pipeline movement, and removes stale prospects. The Prospect Development Maturity Index self-assessment shows where your organization sits on those practices relative to the field.

The Authority Card

A policy becomes operational when a team can answer these questions for each workflow. One simple approach is to create an Authority Card before a pilot begins.

Wondering what happens next?

Download the Fundraising Judgment Benchmark Executive Summary for a 90-day implementation plan!

Final thoughts

The Fundraising Judgment Benchmark began with a simple question: who gives the better fundraising advice, a practitioner or AI? The data made that question less interesting.

Across 363 nonprofit professionals and 1,452 blinded comparisons, neither source won overall. AI could be dramatically stronger in one workflow and clearly weaker in another. That is not a contradiction. It is a reminder that “AI” is not a single job and fundraising is not a single decision.

The more important discovery was the gap between quality and authority. Practitioners were willing to recognize a strong AI recommendation. They were far less willing to transfer accountability along with it. Only 1.7% of judgments supported execution without individual human review.

For nonprofit leaders, the implication is practical. Stop asking whether the organization trusts AI. Decide where AI can inform, draft, recommend, execute within rules, or must stop and escalate. Put a named person on consequential workflows. Make evidence and intervention visible. Pilot with boundaries. Measure donor trust alongside performance. And revisit the authority level as the evidence changes.

The governing belief: The source never decided quality. The workflow did. And quality never decided authority. Accountability did.

AI may eventually participate in far more fundraising decisions than it does today. This study suggests that the organizations best prepared for that future will not be the ones that automate fastest. They will be the ones that know exactly where judgment belongs, who owns it, and what evidence is required before authority expands.

Appendix A: All scenario results

Appendix B: Technical notes, limitations, and responsible claims

Statistical methods

  • Scenario preference: two-sided exact binomial test against 50%, with Wilson 95% confidence intervals.
  • Multiple comparisons: Holm family-wise correction across all 16 scenario tests.
  • Overall preference: mean of four judgments per respondent, with a respondent-cluster bootstrap interval.
  • Sensitivity analysis: all primary tests were rerun after excluding the 14 platform-flagged responses; all eight primary adjusted-significant scenario findings remained statistically reliable.
  • Governance and subgroup analysis: exploratory generalized estimating-equation models with scenario fixed effects and respondent clustering.
  • Risk relationship: cross-tabulation and chi-square; the first frontline scenario is excluded because its risk response options differed.
  • Rationale analysis: the first frontline scenario is excluded because its reason question allowed multiple selections where later items were single choice.
  • Controls: valid-response denominator of 278; respondents could select up to three.
  • Qualitative review: 196 open-text responses; 193 included in descriptive thematic review. Themes overlap and should not be read as prevalence estimates.

Important limitations

  • The sample was a convenience sample and should not be treated as statistically representative of the entire nonprofit sector.
  • The scenarios were hypothetical. The study did not observe donor behavior, revenue, long-term trust, or downstream outcomes.
  • Each scenario compared one selected practitioner recommendation with one AI recommendation. This is not a consensus panel of experts or a benchmark of all AI systems.
  • Source position was fixed within each scenario, although human and AI positions were counterbalanced two-and-two within every role branch.
  • Role groups evaluated different scenarios, so role-level preference percentages should not be interpreted as general attitudes toward AI.
  • The controls question was optional, producing 278 valid responses rather than 363.
  • This is industry research and has not undergone peer review.

What the evidence supports—and what it does not

Acknowledgments

This benchmark was possible because practitioners shared their judgment before anyone knew how the comparison would turn out. We are grateful to Brett Loney, J.D., CFRE, ACFRE, Adopt America Network; Leah Eustace, MPhil, ACFRE, University of Ottawa; Lindsey Nadeau, CFRE, UNICEF USA; Marco Corona, Southern Environmental Law Center; Sarah Marcotte, SickKids Foundation; and Katherine Scott, The Princess Margaret Cancer Foundation for contributing human recommendations and helping ground the work in real fundraising practice.

We are equally grateful to the nonprofit professionals who completed the benchmark. Their willingness to judge difficult scenarios, identify risk, and articulate where human judgment should remain central is what turns this from a technology comparison into a practical governance resource for the field.

Affiliations are provided for identification only. Participation as a practitioner contributor does not imply institutional endorsement of the study, its conclusions, or any commercial product.

References

  • Caffier, J., Stavrova, O., & Kleinberg, B. (2026). Prosocial persuasion at scale? Large language models outperform humans in donation appeals across levels of personalization. arXiv:2604.03202. https://doi.org/10.48550/arXiv.2604.03202
  • Castelo, N., Bos, M. W., & Lehmann, D. R. (2019). Task-dependent algorithm aversion. Journal of Marketing Research, 56(5), 809-825. https://doi.org/10.1177/0022243719851788
  • Dietvorst, B. J., Simmons, J. P., & Massey, C. (2015). Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General, 144(1), 114-126. https://doi.org/10.1037/xge0000033
  • Green, B. (2022). The flaws of policies requiring human oversight of government algorithms. Computer Law & Security Review, 45, 105681. https://doi.org/10.1016/j.clsr.2022.105681
  • Logg, J. M., Minson, J. A., & Moore, D. A. (2019). Algorithm appreciation: People prefer algorithmic to human judgment. Organizational Behavior and Human Decision Processes, 151, 90-103. https://doi.org/10.1016/j.obhdp.2018.12.005
  • Vaccaro, M., Almaatouq, A., & Malone, T. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8, 2293-2303. https://doi.org/10.1038/s41562-024-02024-1
  • Virtuous & Fundraising.AI. (2026). The 2026 Nonprofit AI Adoption Report: The state of AI adoption and transformation for nonprofits. Benchmark study of 346 organizations.

SUGGESTED CITATION FOR THIS WORK: Kindsight Research. (2026). The Fundraising Judgment Benchmark: The Authority Gap—Where AI Helps, Where Human Judgment Must Lead. Kindsight.

Get the insights that will inform your AI strategy!

Our executive summary highlights the findings that are most consequential for fundraisers, advancement leaders, and the organizations deciding how much authority AI should have:

  • 5 essential insights every nonprofit and advancement leader should know
  • A practical framework for assigning AI authority
  • A 90-day path from AI inventory to informed decision
  • And more!

Download your copy of the Fundraising Judgment Benchmark executive summary today!

Download the executive summary today!