Back to Articles
    AI Implementation

    Building an Evaluation Set for Your AI Workflows: How to Know It Still Works

    The AI workflow you built in March is not the same workflow in September, even if nobody touched it. The model underneath it has been updated, possibly more than once, possibly without an announcement. Most nonprofits have no way to tell whether the summaries, drafts, and extractions they rely on are as good as they were six months ago. An evaluation set is the unglamorous fix: twenty to fifty real test cases with known-good answers, rerun on a schedule, checked by a person with a rubric.

    Published: September 21, 2026•17 min read•AI Implementation
    A nonprofit staff member reviewing AI output quality against a set of recorded test cases

    There is a particular kind of failure that almost never gets caught, and it is not the dramatic one. Nobody misses the day the tool stops working, returns an error, or produces something visibly absurd. What slips through is the quieter version: the grant summary that used to include the funder's stated priorities and now sometimes does not, the thank-you draft that used to name the specific program the donor supported and now says "your generous gift" in a way that reads like a form letter, the case note extraction that used to pull the follow-up date reliably and now pulls it about four times in five. Nothing broke. The output still looks fine at a glance. It is just worse, and by the time someone notices, it has been worse for months.

    This happens because the thing your workflow sits on top of is not stable. Providers ship new model versions on a cadence that has accelerated considerably, retire older ones, and sometimes change behavior behind a version label that did not change. Industry writing on this has started calling it silent model drift, and the honest summary is that deprecation is codified while behavior change is not. There is no standard notification when a model stays in place and starts answering differently. Add your own changes on top of that, a tweaked prompt here, a new instruction added to handle one awkward case, a different document format being fed in, and the surface area for quiet degradation is large.

    Software teams solved a version of this problem decades ago with regression tests. You keep a fixed set of inputs and the outputs you expect, you rerun them after every change, and if something that used to pass now fails, you know immediately and you know roughly where to look. The same logic applies to AI workflows, with one complication: the output is language, not a number, so "did it pass" requires a bit more thought. That complication is solvable, and this article is about solving it with a spreadsheet and an hour a month rather than with an engineering team and a testing platform.

    The stakes here are not abstract for nonprofits. The sector has adopted AI quickly and governed it slowly. Reporting on the 2026 State of Nonprofit AI research from NTEN and The Bridgespan Group found that 58 percent of nonprofits report no AI roadmap exists, with fewer than half having written guidance on responsible use and data handling. Organizations in that position are not running quality checks either. They are trusting outputs because the outputs looked right the first time they saw them, which is a reasonable thing to do once and an unreasonable thing to do indefinitely.

    To be clear about scope: this article is not about adversarial testing, prompt injection, or probing a system for ways it can be made to misbehave. That is a different and also worthwhile discipline, covered separately in our pieces on adversarial prompts against nonprofit chatbots and pre-launch red team checklists. This is the mundane counterpart: routine, boring, scheduled confirmation that normal work on normal inputs still comes out the way it should. It is less interesting and more frequently the thing that actually catches the problem you have.

    What an Evaluation Set Actually Is

    An evaluation set is a fixed collection of inputs and the outputs you have already agreed are good, stored somewhere stable, run through your AI workflow on a schedule and after any change, and graded against a written standard. That is the whole concept. Practitioners sometimes call it a gold set or a golden dataset, which makes it sound more formal than it needs to be. For a nonprofit with one or two AI workflows in daily use, it is a spreadsheet with a few dozen rows.

    The defining property is that it does not change. This trips people up, because the instinct when a test fails is to fix the test. If your test cases drift along with your prompts, you lose the only fixed reference point you had, and you can no longer answer the question the set exists to answer, which is whether today's output is worse than the output you approved. You can add new cases as new situations arise. You should essentially never edit an existing case just because it stopped passing. Changing a case is a deliberate decision made because the correct answer genuinely changed, for example because your organization renamed a program or updated a required disclosure, and it should be recorded as such with a date and a reason.

    It is worth being precise about what an evaluation set is not. It is not a measure of whether AI is good at your task in some absolute sense. It is not a benchmark you can compare against other organizations. It is not an accuracy percentage you can put in a board report without heavy caveats. It is a tripwire. Its only job is to tell you that something changed, and to tell you soon enough that you can investigate before the change reaches a donor's inbox or a funder's desk. Holding that modest goal in mind keeps the exercise proportionate. You are not building a research instrument. You are building a smoke detector.

    The second useful property is that an evaluation set makes changes cheap to try. Most organizations are cautious about editing a working prompt because they have no way to tell whether the edit helped, so prompts calcify and accumulate awkward workarounds nobody dares remove. With a set of cases you can rerun in twenty minutes, prompt editing becomes a normal activity rather than a risky one. You change something, you rerun, you see whether the score moved, and you keep or revert. This is the same reason software teams write tests, and it is the benefit people underestimate before they have one.

    Finally, the set gives you an institutional record. Staff turn over, the person who built the workflow leaves, and what remains is a prompt nobody fully understands doing something nobody has verified. The evaluation set is the part of the documentation that is executable. It sits naturally alongside the practice of documenting AI workflows, and it answers the question a written description cannot: not what this workflow was supposed to do, but what it did the last time anyone checked.

    The four parts of an evaluation set

    Everything you need, and nothing you do not

    • The input: the exact document, transcript, record, or question the workflow receives
    • The expected output: an approved answer, or a description of what a correct answer must contain
    • The checks: the specific things a grader looks for, written plainly enough for anyone to apply
    • The history: the score from each run, with the date, the model version, and the prompt version

    Harvesting Cases From Work You Have Already Done

    The single best source of test cases is your own archive, and the reason is that invented cases are too clean. When someone sits down to write test inputs from scratch, they produce well-formed, tidy examples that the workflow handles easily, because those are the examples that come to mind. Real inputs are messier. The grant guidelines PDF has a two-column layout and a scanned appendix. The intake form has a free-text field where someone wrote three sentences in Spanish. The case note has an abbreviation only one caseworker uses. Those are the inputs where quality actually degrades, and you cannot invent them convincingly.

    Go back through the last several months of real work the workflow has handled and pull examples. For a grant summarization workflow, that means actual funder documents you summarized. For a donor acknowledgment drafting workflow, it means actual gift records and the letters that went out. For case note extraction, actual notes. For translation of program materials, actual materials. For intake triage, actual intake submissions. In each case you want the input as it arrived, not a cleaned-up version, because the mess is part of the test.

    Aim for somewhere between twenty and fifty cases. Fewer than twenty and a single odd result swings the score enough that you cannot tell signal from noise. More than fifty and the human grading burden becomes real enough that the monthly run quietly stops happening, which is the most common way an evaluation practice dies. Thirty is a good target for a single workflow. If you run several distinct workflows, each gets its own smaller set rather than one giant combined set, because a combined score tells you something is wrong without telling you where.

    Composition matters more than count. A set made entirely of typical cases will keep passing while the workflow quietly fails on everything unusual. A reasonable mix is roughly half routine cases that represent the bulk of the work, a quarter awkward cases that are legitimately harder, and a quarter edge cases including the specific failures you have already seen. That last group is the most valuable and the most often forgotten. Every time your workflow produces something wrong in real use, that input belongs in the evaluation set permanently, with the corrected output beside it. A set that accumulates your actual past failures becomes a genuine institutional memory of the ways this workflow goes wrong.

    Include at least one or two cases where the correct behavior is refusal or escalation. If your intake triage workflow receives a submission that is ambiguous or falls outside your services, the right output is a flag for human review, not a confident categorization. If your summarizer is handed a document that does not contain the information requested, the right output says so rather than inventing something plausible. Models change in their willingness to say "this is not in the document," and that shift is exactly the kind of drift a well-composed set catches. It is closely related to the broader problem of reducing AI hallucinations, and test cases are the mechanism that tells you whether your mitigations are still working.

    Where to find cases, by workflow type

    Pull from the archive, not from imagination

    • Grant summaries: real funder guidelines, including the badly formatted ones and the ones with unusual eligibility rules
    • Donor acknowledgments: gift records spanning a small first gift, a large restricted gift, a memorial gift, and an in-kind donation
    • Case note extraction: notes with abbreviations, missing fields, multiple dates, and one where the requested field genuinely is not present
    • Translation: materials containing program names, idioms, legal phrases, and terms your community uses differently
    • Intake triage: clear referrals, borderline referrals, out-of-scope requests, and at least one that should be escalated to a person

    Writing Expected Outputs and a Rubric Anyone Can Apply

    Here is where most first attempts go wrong. Someone writes an expected output that is a single perfect paragraph, then discovers at grading time that the model produced a different paragraph which is equally good, and now the grader has to decide whether a different-but-fine answer counts as a pass. Do that thirty times and the exercise becomes exhausting and inconsistent. The fix is to stop treating the expected output as a target the model must match and start treating it as a description of what a correct answer must contain.

    In practice this means writing each case as a short list of requirements rather than a model answer. For a grant summary, the requirements might be: names the funder correctly, states the deadline as March 14 2027, identifies the three stated funding priorities, notes the geographic restriction to the tri-county area, does not claim we are eligible or ineligible, and runs under 300 words. Any output meeting all six passes, regardless of how it is phrased. This is far easier to grade, far more consistent between graders, and it forces you to articulate what you actually need from the workflow, which is a useful exercise in its own right.

    Keep a reference answer alongside the requirements anyway, but use it differently. Its purpose is not comparison, it is calibration. When a new person takes over grading, reading a few approved outputs teaches them what "good" looks like in your organization faster than any rubric can. It also gives you something concrete to point at when quality drops, because showing someone the old output next to the new one makes a case that a score change does not.

    Write the rubric in plain language aimed at whoever will actually grade, who is probably a program manager or a development associate rather than anyone technical. Every check should be answerable yes or no by someone reading the output once. "Is the tone appropriate" is not answerable. "Does it avoid clinical jargon and address the reader as you" is. "Is the summary accurate" is not answerable. "Does every fact in the summary appear in the source document" is, though it takes longer. The discipline of converting vague quality attributes into observable checks is most of the work of building a usable rubric, and it is the part worth spending real time on because you write it once and use it for years.

    Grade at the level of the individual check rather than the whole output. A case with six requirements yields six yes or no answers, which is much more informative than one pass or fail. When quality drifts, it almost never drifts everywhere at once. It shows up as one specific check failing across several cases, and that pattern is the thing that tells you what changed. If the "includes the required disclosure" check fails on four cases this month and passed on all of them last month, you have a specific, diagnosable problem rather than a vague sense that things are worse.

    Turning vague quality into checkable requirements

    Every check must be answerable yes or no on one reading

    • Instead of "accurate," write "every date, dollar figure, and name in the output appears in the source"
    • Instead of "complete," list the specific fields that must be present
    • Instead of "appropriate tone," name the things to avoid and the form of address to use
    • Instead of "right length," give a word count range the grader can verify
    • Instead of "did not make things up," write "contains no claim absent from the source document"

    Deterministic Checks and Judgment Checks Are Different Animals

    Sort every check in your rubric into one of two buckets, because they behave differently and they deserve different treatment. Deterministic checks have exactly one right answer and can be verified by anyone, or by a spreadsheet formula, without interpretation. Is the EIN in the output the correct nine digits for this organization? Does the acknowledgment letter include the required statement that no goods or services were provided in exchange for the gift? Is the deadline the date on page two of the guidelines? Does the output contain the phrase we are legally required to include? These are facts. They either hold or they do not.

    Judgment checks require a person to form a view. Is the tone warm without being saccharine? Does the summary emphasize the things a program director would care about? Is the translation natural rather than literal? Does the draft sound like our organization? These are real quality attributes and you cannot simply drop them, because tone and register are frequently where model updates show up first. A new model version rarely gets the EIN wrong. It quite often becomes chattier, or more hedging, or more inclined toward a corporate register that does not fit a community organization.

    The reason the distinction matters operationally is that deterministic checks are cheap and judgment checks are expensive. If a case has eight deterministic checks and two judgment checks, a grader can clear the eight quickly and spend their attention on the two. In a spreadsheet you can even automate several deterministic checks with ordinary formulas, searching for a required phrase, verifying a number appears, or counting words, which removes them from the human workload entirely. Nobody needs a testing platform to run a text search on a cell.

    Weight the two kinds differently when you score. A failed deterministic check on something legally or factually required is not a small deduction, it is a stop. An acknowledgment letter missing the substantiation language required for a charitable contribution receipt is a compliance problem, not a quality nuance, and a single failure should raise an alarm regardless of how well the rest of the set scored. Judgment checks are better treated as an aggregate trend, where one grader thinking one output felt slightly off is noise and six outputs drifting toward a different register is a signal.

    Build your deterministic checks around the things that would embarrass you or expose you. Required disclosures. Legal names and identifiers. Dollar amounts. Dates. Program names spelled the way your organization spells them. Whether the output ever states or implies a determination the workflow is not authorized to make, such as eligibility for a service or approval of a request. That last one is worth a permanent check in any workflow that touches clients, and it pairs with the design principle behind human approval gates in agentic workflows: the system should not be making the call, and your test set should confirm it is not starting to.

    Deterministic checks

    One right answer, verifiable by anyone

    • Required disclosure or substantiation language is present verbatim
    • EIN, legal name, program name, and dollar figures are exactly correct
    • Every required field is populated and no field is invented
    • Output length falls within the stated range
    • No determination is stated that the workflow is not authorized to make

    Judgment checks

    A person forms a view, and trends matter more than single results

    • Tone matches how your organization actually speaks to this audience
    • The summary emphasizes what the intended reader needs, not just what was longest in the source
    • Translation reads naturally to a native speaker rather than word for word
    • The draft would need light editing rather than substantial rewriting
    • Nothing in the output would be uncomfortable to show the person it describes

    Using a Second Model as Grader, and Where That Breaks

    The obvious labor-saving idea is to have a model do the grading. Give a second model the input, the output, and the rubric, and ask it to score each check. This technique is widely used and it genuinely works for parts of the job, particularly for checks that are mechanical but tedious, like confirming that every date in a summary appears in the source document. Used for that kind of work it can cut your human grading time substantially, and for a two-person organization that difference determines whether the monthly run happens at all.

    It also fails in specific and well-documented ways that you need to know about before you rely on it. Models used as graders show self-preference bias, systematically rating output from the same model family more favorably, and research analyzing a broad set of mainstream models has found the counterintuitive pattern that stronger models often show larger self-preference bias, not smaller. They show verbosity bias, rating longer answers higher regardless of quality. They show position bias, favoring whichever candidate answer appeared first or last when comparing two. And their agreement with human judgment degrades notably outside controlled conditions, with reported gaps between controlled-test accuracy and production bias-test performance that should make anyone cautious about treating a model score as ground truth.

    There is a deeper problem for our purposes specifically. The grader model is itself subject to exactly the drift you are trying to detect. If both your production model and your grader model update, and the grader's standards shift in the same direction as your workflow's output, the score can stay flat while the actual quality falls. You have built an instrument made of the same material as the thing it measures. This is not a reason to avoid model grading entirely, but it is a decisive reason never to let a model-graded score be the only thing standing between you and a silent regression.

    The workable arrangement is a division of labor. Let a model handle high-volume mechanical verification and flag the cases it thinks failed. Have a person review every flagged case, plus a fixed random sample of the cases the model passed, which is what catches the grader missing things. Use a grader from a different provider than the one running your workflow, which reduces the self-preference problem, and pin the grader's version so it changes only when you choose. Never let a model grade the judgment checks unreviewed, because tone and appropriateness are precisely where model judgment diverges most from the judgment of someone who knows your community.

    Calibrate the grader before you trust it. Take ten cases you have already graded by hand, run the model grader over them, and compare. If it agrees with you on most checks, you have a useful assistant. If it disagrees frequently, your rubric is probably too vague for a model to apply, which usually means it is also too vague for a new human grader, so the disagreement is telling you something useful either way. Recalibrate whenever you change the grader model. Running graders and production workflows from different providers is one practical argument for a deliberate multi-model strategy rather than standardizing everything on a single vendor.

    Rules for using a model as grader

    Helpful assistant, unreliable authority

    • Use a different provider for grading than for production, to blunt self-preference bias
    • Pin the grader version so your measuring instrument does not drift on its own
    • Have a person review every flagged failure and a random sample of the passes
    • Keep judgment checks with a human who knows the audience and the organization
    • Calibrate against ten hand-graded cases before trusting it, and again after any grader change

    Scoring, Thresholds, and What Counts as a Regression

    Scoring should be boring. Count the checks that passed, divide by the checks attempted, and you have a percentage for the run. Do this per check type as well as overall, so you can see deterministic performance and judgment performance separately, and do it per check where you can, so you can see that check number four failed on five cases this month and one case last month. The per-check view is where diagnosis lives. The overall percentage is mostly for noticing that something needs diagnosing.

    Establish a baseline before you need one. The first time you run the set, after you have fixed anything obviously broken, record that score and call it the baseline. Every subsequent run is compared against it and against the most recent run. Without a recorded baseline, you are left arguing about whether things feel worse than they used to, which is an argument nobody wins. Record the model version and the prompt version alongside each score, because a score with no version attached cannot be investigated six weeks later.

    Set thresholds in advance, in writing, before you have a bad run and an incentive to rationalize it. A defensible starting point for most nonprofit workflows: any failure of a deterministic check that touches legal or financial correctness is investigated immediately regardless of the overall score. Any drop of more than five percentage points from the baseline triggers an investigation. Any drop of more than ten percentage points, or any two consecutive runs showing a decline, means the workflow goes back to human review for its outputs until you have found and fixed the cause. Adjust these numbers to your own risk tolerance, but adjust them deliberately, not in the moment.

    Distinguish a regression from noise. Model outputs vary between runs even with identical inputs, so a case that passed last month and failed this month is not automatically a problem. The patterns that indicate a real regression are consistency and concentration: the same check failing across multiple cases, a failure that reproduces when you rerun that single case a few times, or a decline that persists into the next run. The pattern that usually indicates noise is one case failing one check once. Rerun the individual case two or three times before you escalate anything, because that thirty-second step prevents a lot of unnecessary alarm.

    A note on improvements, which people forget to handle. If the score jumps up sharply, that is also worth ten minutes of attention. Sometimes it means a new model version genuinely handles your task better, which is useful information and may justify updating your baseline. Sometimes it means something changed in how you are grading, or that a grader is being more generous, or that a case got accidentally edited. An unexplained improvement is a measurement problem until you have explained it. This kind of periodic review fits naturally into an existing continuous quality improvement rhythm rather than sitting as a separate technical ritual nobody outside the process understands.

    Thresholds worth writing down in advance

    Decide what counts as a problem before you have one

    • Any legal or financial check failing: investigate immediately, whatever the overall score
    • More than five points below baseline: investigate before the next production use
    • More than ten points below baseline: outputs return to full human review until resolved
    • Two consecutive declining runs: treat as a regression even if each drop was small
    • An unexplained jump upward: verify the measurement before celebrating it

    Cadence, Ownership, and the Spreadsheet That Holds It All

    Three triggers should cause a run, and the first is the calendar. Monthly is the right default for a workflow in regular use. Quarterly is defensible for something used occasionally and low stakes. Weekly is overkill for almost any nonprofit and guarantees the practice will be abandoned within two months. Put the monthly run on a specific date with a specific person's name on it, because a cadence with no owner is a cadence that happens twice and then stops.

    The second trigger is any change you make. A prompt edit, however small, gets a run before and after. Adding a new instruction to fix one problem very often breaks something else, and the evaluation set is how you find out in twenty minutes rather than in six weeks. This applies equally to changing a document template that feeds the workflow, changing the tool the workflow runs in, or changing a setting you did not think was important. Organizations maintaining a shared prompt library should attach the current evaluation score to each stored prompt version, so anyone reusing it can see when it was last verified and against which model.

    The third trigger is any change the provider makes. When you see a notice about a new model version, a deprecation, or a platform update, run the set. This is the trigger most likely to catch a real regression and the one most likely to be missed, because provider announcements arrive in a technical changelog nobody at a small nonprofit is reading. Subscribe someone to the release notes of the tools you depend on, and treat any version change as a reason to run, including the version changes that arrive without an announcement. If your platform lets you pin a specific model version, do that, so upgrades happen when you decide rather than when a provider decides, and understand that pinning buys time rather than permanence, since versions get retired eventually. Our guidance on model selection for nonprofits covers how to weigh that stability question when you choose a tool in the first place.

    Ownership in a small organization should be explicit and singular. Name one person, give them ninety minutes a month, and make the run a line item in their role description rather than a favor. In a ten-person nonprofit this is usually the operations manager or whoever owns the workflow itself. Grading can be shared among two or three people, which is good practice anyway because it surfaces rubric ambiguity, but scheduling and follow-through need one name. Where an organization has cultivated internal AI capability, this is a natural responsibility for that person, and it is one of the more concrete things an AI champion can own that produces visible value.

    The tooling can be a spreadsheet, and for most nonprofits it should be. One tab holds the cases, with columns for a case identifier, the input or a link to it, the requirements list, the reference output, and notes. A second tab holds the results, one row per case per run, with the date, the model version, the prompt version, a yes or no for each check, and space for the grader's comment. A third tab holds the summary: a row per run with the overall score, the deterministic score, the judgment score, and what you concluded. Charting the summary column gives you the single most useful artifact in the whole system, a line that should be flat and that you will notice when it is not. Working directly with AI assistance inside your spreadsheet, as covered in our piece on using AI in spreadsheets, can automate several of the deterministic checks without any additional software.

    What triggers a run

    Three triggers, one owner, one calendar entry

    • The monthly calendar date, with a named owner and time actually blocked
    • Every prompt edit, before and after, including the ones that seem trivial
    • Any announced model version change, deprecation notice, or platform update
    • Any change to the input format, source template, or connected system
    • Any real-world failure someone reports, which also becomes a permanent new case

    Personal Information in Test Cases

    This is the part that stops nonprofits from building evaluation sets at all, and the concern is legitimate. The most valuable test cases are drawn from real work, and real work at a nonprofit frequently involves client records, donor details, health information, immigration status, student records, or the contents of a difficult conversation. Copying that material into a spreadsheet that lives in a shared drive, gets emailed around at grading time, and persists for years is a genuinely bad idea, and the fact that it is for testing does not change what it is. Federal and state obligations follow the data regardless of purpose, whether that is HIPAA for health information, FERPA for student records, state privacy statutes, or the confidentiality terms in your own grant agreements and client consent forms.

    The default answer is to de-identify before a case enters the set. Replace names with consistent placeholders, shift dates by a fixed offset that preserves intervals, change addresses to a different but similar neighborhood, and alter identifying details that are not relevant to what you are testing. The key word is consistent: if the same person appears twice in a case, they need the same placeholder both times, or you will have broken the thing you are testing. Do the de-identification once, carefully, when the case is created, and the set becomes something you can store and share far more freely for years afterward.

    Preserve the structural characteristics that made the case a good test. If the original case was hard because the record contained three different dates and the workflow had to pick the right one, your de-identified version needs three different dates in the same relationship. If it was hard because a name was ambiguous, keep an ambiguous name. De-identification that also sands off the difficulty produces a test set that passes everything and tells you nothing. This is the same reasoning behind careful use of synthetic data for privacy protection, where the value of the substitute depends entirely on whether it preserved the properties that mattered.

    Some cases cannot be de-identified without destroying them, particularly in translation and case note work where the specific language is the thing being tested. For those, restrict rather than exclude. Keep the case in a separate, access-controlled location with named people who may open it, keep it out of the general spreadsheet, give it a retention date, and record in the main set that a restricted case exists with a pointer to where. A small number of restricted cases handled properly is better than either exposing sensitive material or pretending the hard cases do not exist.

    Two final practical points. First, check whether your AI tool retains inputs for training or logging, because an evaluation run sends every case through it, thirty at a time, every month. Configuration that is acceptable for one-off use may not be acceptable for systematic reuse. Second, write down where the set lives, who can open it, how long cases are kept, and what happens to them when a workflow is retired. This belongs in the same retention and access governance as your other sensitive records rather than being managed informally by whoever happens to own the spreadsheet.

    Handling sensitive material in a test set

    De-identify once, carefully, and keep the difficulty intact

    • Replace identifiers with consistent placeholders, and shift dates by a fixed offset
    • Keep the structural features that made the case hard, including ambiguity and clutter
    • Restrict, do not exclude, the few cases that cannot survive de-identification
    • Confirm your tool's retention and training settings before running cases through it monthly
    • Give the set a documented location, access list, and retention schedule

    When the Set Starts Failing: A Triage Order

    A run comes back eight points below baseline. The instinct is to start rewriting the prompt, and it is the wrong first move, because you do not yet know what changed. Work through the possibilities in order of how easy they are to rule out, which happens to be roughly the order of how likely they are to be the real cause.

    Start with the measurement itself. Did the grader change? Is a new person grading, applying the rubric differently? Did someone edit a case or a requirement? Did a spreadsheet formula break when a column moved? A surprising share of apparent regressions are measurement artifacts, and they are the cheapest thing to check. Compare this run's failures against the same cases last month and read the actual outputs rather than the scores. If the outputs look comparable and the scores moved, the problem is in the measurement.

    Next, check what changed on your side. Pull the prompt version and compare it to the one in the last passing run, including the small edits nobody logged. Check whether the input documents changed format, whether a connected system started returning data differently, whether a template was updated. Changes you made are easier to identify and easier to reverse than changes a provider made, so exhaust them before assuming the model is at fault. This is where version discipline pays for itself, and where organizations without it end up guessing.

    Then check the provider. Look at release notes and status pages for the period between your last passing run and this one. Look at whether the model identifier you are using resolves to a different version than it did. If you find a version change, you have your explanation and a decision to make: revise the prompt for the new version, pin back to an older version if that is still available, or move the workflow to a different model. Run the set again after whichever you choose, because the fix needs verifying like anything else.

    Read the failing outputs before you change anything. Scores tell you that something moved, but only the text tells you what. Look for the pattern: is it the same check failing, the same kind of input, the same part of the output? Is the model now hedging where it used to be direct, adding caveats, changing format, or truncating? These patterns point at a fix directly, and a prompt edit made after reading five failing outputs will be better than three edits made from the score alone.

    While you investigate, change how the output is handled. Put the workflow back under full human review, or pause it for the highest-stakes uses, and say so to the staff who depend on it. Then consider the outputs already produced since the last passing run, because if quality dropped after a model change three weeks ago, three weeks of work may need spot-checking. That is the uncomfortable part of catching a regression, and it is also why the monthly cadence matters: the difference between a month of suspect output and a year of it is the whole argument for doing this at all. Workflows that involve outside language or sensitive judgment, such as the ones described in our article on translation quality review, warrant the most conservative handling during an investigation, because the people most affected by a quality drop are the least likely to be in a position to complain about it.

    Triage order when the score drops

    Cheapest to rule out first, and usually the likeliest

    • The measurement: a new grader, an edited case, a broken formula, a changed rubric
    • Your changes: prompt edits, input format shifts, template and integration updates
    • The provider: release notes, deprecations, and version identifiers that quietly resolved elsewhere
    • The text: read the failing outputs for the pattern before editing anything
    • The backlog: spot-check real work produced since the last passing run

    Getting Started This Week Without Overbuilding

    The failure mode for this kind of project is not doing it badly, it is designing it so thoroughly that it never launches. A set of fifteen cases with four checks each, graded by one person in an hour, running every month without fail, is worth more than an elaborate framework that gets built once and abandoned. Start small enough that the first run happens this month.

    Pick the workflow where a quiet quality drop would matter most. Usually that is the one touching people outside the organization, donors, funders, clients, or participants, because internal work gets a second pair of eyes by default and external work often does not. Then spend one sitting pulling fifteen to twenty real examples from the last few months, including every failure anyone remembers. That sitting is the bulk of the work and it is genuinely a single afternoon.

    For each case, write three to six requirements rather than a perfect answer, and mark each as deterministic or judgment. Run all the cases through the workflow as it stands today, grade them, fix anything obviously broken that you discover in the process, and record the resulting score as your baseline with the date and the model version. Most organizations find one or two genuine problems during this first pass, which tends to settle any internal argument about whether the exercise was worth the afternoon.

    Then put the next run on the calendar with a name attached, write the thresholds down before you need them, and tell the people who use the workflow that it is being checked monthly and what will happen if a check fails. That last step is the one that turns a spreadsheet into a practice, because it creates an expectation other people hold you to. Expand later: add cases as new failures appear, add workflows once the first one has run three times cleanly, and consider model-assisted grading only when the human grading burden is actually the constraint. The discipline is what matters, and the discipline is small.

    A first evaluation set in one afternoon

    Small enough to finish, real enough to be useful

    • Choose the one workflow whose quiet failure would reach someone outside the organization
    • Pull fifteen to twenty real past inputs, de-identified, including every remembered failure
    • Write three to six yes-or-no requirements per case, marked deterministic or judgment
    • Run, grade, fix what you find, and record the baseline with date and model version
    • Calendar the next run with an owner, and write the thresholds down in advance

    Conclusion

    Most nonprofits currently find out that an AI workflow has degraded the way you find out about a slow leak, which is when the damage is already done and someone else points at it. A funder mentions that the last two applications read oddly. A donor replies to ask which program their gift supported, because the letter did not say. A caseworker starts quietly redoing the extraction by hand because it stopped being reliable and it was easier to work around than to report. None of those moments produce a clear signal about when things changed or why, and by then the question is not just how to fix it but how much of the past year's output needs reviewing.

    An evaluation set replaces that with a number on a calendar. Twenty to fifty real cases, harvested from work you have already done and de-identified once. Requirements written as yes-or-no checks rather than perfect answers, sorted into the deterministic ones that a formula or a quick look can verify and the judgment ones that need a person who knows your community. A monthly run with a named owner, plus a run on every prompt edit and every model change. Thresholds decided in advance so a bad result produces an investigation rather than a debate. A spreadsheet, a baseline, and a line on a chart that should be flat.

    None of this is sophisticated, and that is the point. The organizations that catch quality problems early are not the ones with the best tooling, they are the ones with a boring habit that happens on schedule whether or not anything seems wrong. A model grader can reduce the labor, and it is worth using for mechanical checks with a different provider and a pinned version, but it cannot be the only thing watching, because it is made of the same material as the system it measures and it drifts too. A person reading a handful of outputs against a written standard remains the backstop.

    The ground under your AI workflows will keep moving. Providers will ship new versions, retire old ones, and change behavior in ways that are not always announced, and your own prompts and inputs will change alongside them. You cannot make that stop, and you do not need to. You need to know when it has happened, and you need to know within weeks rather than within a year. That is the entire job of an evaluation set, and an afternoon of building plus ninety minutes a month is a remarkably small price for being the organization that noticed first.

    Want to Know Your AI Workflows Still Work?

    We help nonprofits build practical evaluation sets, write rubrics staff can actually apply, and set up a monthly rhythm that catches quality drops before donors, funders, or clients do.