Reading Other Nonprofits' 990s With AI
Every nonprofit in the country files a detailed public financial disclosure, and almost nobody reads anyone else's. The Form 990 tells you what your peers pay, how they structure their programs, who their vendors are, and how fragile their revenue is. This guide covers where to find the filings, which sections actually repay attention, how to use AI to read dozens of them without inventing numbers, and how to build a peer set that makes the comparison mean something.

Nonprofit leaders spend a remarkable amount of energy guessing at things that are already published. What does a development director at an organization our size actually earn? Is our overhead unusually high, or does everyone in this field report it the same way? Which consultants do the organizations we admire actually hire? How many of our peers are one government contract away from a crisis? These questions get answered in board meetings with anecdote, memory, and whatever a colleague mentioned at a conference.
The answers exist in public documents. Nearly every tax-exempt organization in the United States files an annual information return, and unlike a corporate tax return, that filing is public by law. It lists the highest-paid staff by name and amount, breaks expenses into program, management, and fundraising columns, names the five highest-compensated independent contractors, itemizes revenue by source, and in many cases discloses grants made, related organizations, and governance practices. It is the single richest comparative dataset the sector has, and it is free.
The reason nobody uses it is friction. A single 990 for a mid-sized organization runs dozens of pages of dense tables and IRS terminology. Reading one carefully takes an hour. Reading fifteen and turning them into a defensible comparison takes a week that no one has. So the document sits there, technically accessible and practically ignored, while organizations continue to set salaries and budgets by feel.
This is precisely the kind of friction language models remove. A capable model can read a 990, pull the specific figures you asked for, and put them in a table alongside fourteen other organizations in the time it takes to make coffee. That capability comes with a serious caveat, though, which shapes most of the advice that follows: models are fluent enough to produce a plausible number when they cannot find the real one. A benchmarking exercise built on invented figures is worse than no benchmarking at all, because it carries the authority of a spreadsheet. The workflow matters more than the prompt.
What a 990 Actually Tells You, and What It Doesn't
Before pointing a model at anything, it helps to know what is genuinely in the document. The Form 990 is an information return rather than a tax computation, which means most of it is disclosure rather than arithmetic. The parts that matter most for comparative work are concentrated in a handful of places, and knowing which ones saves enormous time.
Part I gives you the summary: total revenue, total expenses, net assets at the beginning and end of the year, and headcount for employees and volunteers. Part VII lists officers, directors, trustees, key employees, and the highest-compensated employees, with reportable compensation and estimated other compensation for each, plus average hours worked per week. Schedule J expands on that for individuals above the reporting threshold, splitting compensation into base pay, bonus and incentive, other reportable compensation, deferred compensation, and non-taxable benefits.
Part VIII breaks revenue into contributions, program service revenue, investment income, and other categories, and separates government grants from other contributions. That separation is one of the most useful pieces of comparative information in the entire document, because it tells you how much of a peer's model depends on public funding. Part IX is the statement of functional expenses, allocating each expense line across program services, management and general, and fundraising columns. Part VII also has a second section that names the five highest-compensated independent contractors receiving more than one hundred thousand dollars, which is how you learn who your peers hire for technology, fundraising counsel, and consulting.
The limits deserve equal attention. The 990 is a lagging document. A filing you download today may describe a fiscal year that ended eighteen months ago, and organizations routinely file on extension. It is also self-reported and shaped by judgment. The functional expense allocation in Part IX depends heavily on how an organization chooses to allocate shared costs, and two organizations with identical operations can report meaningfully different program expense ratios based on allocation policy alone. Schedule B, which lists major contributors, is generally redacted in public copies, so you cannot see who funds a peer from the 990 itself. And the narrative sections describing program accomplishments are marketing copy written for an audience of funders and watchdogs, which is worth remembering when a model summarizes them back to you as fact.
What the 990 Reports Reliably
Audited or tightly defined figures
- Total revenue, expenses, and net assets
- Named individual compensation above thresholds
- Government grants separated from other contributions
- Top five contractors paid over one hundred thousand dollars
What It Cannot Tell You
Gaps that misread as facts
- Who the major individual donors are, since Schedule B is redacted
- Whether the program expense ratio reflects real allocation or policy choices
- Anything about the current year, since filings lag by a year or more
- Program quality, since narratives are written to persuade
Where the Filings Live and Which Format to Use
The format you start from determines how much you can trust the result, so this decision comes before the AI decision. There are three practical routes to a peer's filing, and they are not equivalent.
The first is ProPublica's Nonprofit Explorer, which is the fastest way to find any organization by name or EIN, see summary financials across multiple years, and download the underlying filing. It is free, well maintained, and covers the great majority of filers. For most organizations doing occasional benchmarking, this is the right starting point.
The second is the IRS itself, which publishes electronically filed returns as machine-readable XML along with annual index files. This is the highest-fidelity source available, because each figure sits in a labeled field rather than in a table that has to be interpreted visually. If you intend to compare more than a handful of organizations, or to repeat the exercise annually, working from XML is dramatically more reliable than working from scanned pages, and it removes the largest single source of extraction error.
The third is Candid, formed from the merger of GuideStar and Foundation Center, which layers organizational profiles and additional context on top of 990 data. Parts are free with registration and deeper access is paid. It is useful when you need context the 990 does not carry, such as program descriptions or self-reported demographic data.
The practical distinction is between structured and unstructured sources. When a model reads XML, it is retrieving a value from a named field, and errors tend to be obvious. When a model reads a scanned PDF of a paper filing, it is interpreting an image of a table, and errors tend to be subtle: a figure from the wrong column, a prior-year number, a total mistaken for a subtotal. Older filings and small organizations are more likely to exist only as scans. Knowing which situation you are in should change how much verification you do.
Choosing a Source
Match the source to the size of the exercise
Looking up three or four organizations
Use Nonprofit Explorer, download the filings, and read them with a model that can accept documents directly. Verification is cheap at this scale because you can check every figure yourself in a few minutes.
Building a peer set of ten to thirty
Work from IRS XML where it exists. The extraction becomes field lookup rather than interpretation, and you can spot-check a sample instead of verifying everything. This is the sweet spot for an annual compensation or expense benchmarking exercise.
Sector-wide or repeated analysis
At this scale you are building a small data pipeline rather than doing research, and the work shifts from prompting to parsing. Consider whether an existing analytics product already answers your question before building anything, and treat the model as a tool for interpreting results rather than for reading filings one at a time.
Five Questions Worth Asking of a Peer's Filing
Benchmarking fails most often not because the data is bad but because the question is vague. Asking a model to tell you how your organization compares to a peer produces a bland summary that changes nothing. Asking a specific question with a decision attached to it produces something usable. Five questions repay the effort consistently.
The first is compensation. If you are about to post a director-level role, or your board is reviewing the executive's pay, the compensation tables across ten comparable organizations give you a defensible range rather than a guess. Boards setting executive compensation have a specific interest here, since the process for establishing reasonable compensation generally involves reviewing data on comparable positions at similar organizations, and 990 filings are the most accessible form that data takes. The exercise is more useful when you capture average hours per week alongside the dollar figure, because a part-time executive at a peer organization is not a comparison point for a full-time role.
The second is revenue concentration and mix. Pull the revenue breakdown for every organization in your peer set and the picture that emerges is strategic rather than financial. If most of your peers derive a large share of revenue from government grants and you do not, you are either missing an opportunity or deliberately avoiding a dependency, and either way the board should know which. This is also the most reliable early warning available about sector-wide fragility, since a peer set heavily weighted toward one funding stream is a peer set that will struggle together when that stream contracts.
The third is expense structure. The functional expense table shows how peers distribute spending across program, management, and fundraising, and more usefully how they distribute it across specific lines such as salaries, occupancy, professional fees, and technology. Comparing your technology spend as a share of total expenses against ten similar organizations is a far better argument for a budget increase than an assertion that you are underinvested.
The fourth is vendors and contractors. The independent contractor section names the highest-paid firms a peer works with and describes the service provided. This is how organizations discover which fundraising consultants their peers actually use, what a comparable organization pays for accounting or IT support, and which firms have deep sector experience. It is genuinely practical market intelligence that would cost real money to buy elsewhere.
The fifth applies mainly when you are looking at foundations and grantmaking organizations. The grants schedule lists recipients and amounts, which turns a funder's 990 into a map of what they actually fund rather than what their website says they fund. For a development team evaluating whether a foundation is a realistic prospect, the distribution of past grant sizes answers the question faster than any conversation. This complements the work covered in monitoring peer nonprofit performance more broadly.
Compensation Questions
For hiring and board review
- What is the range of executive compensation across peers of similar budget size?
- How many staff does each peer report above the disclosure threshold?
- Which roles appear on peer filings that we do not staff at all?
- How much of peer compensation sits in deferred and non-taxable benefits?
Financial Structure Questions
For strategy and budgeting
- What share of peer revenue comes from government sources?
- How many months of expenses do peers hold in net assets?
- What do peers spend on professional fees and technology as a share of budget?
- Which peers have grown or contracted materially over three years?
Building a Peer Set That Means Something
A benchmarking exercise is only as good as the comparison group, and this is where most efforts quietly go wrong. The instinct is to pick the organizations you admire or compete with for funding. That produces an aspirational list rather than a comparable one, and the resulting numbers tend to make your organization look either heroic or inadequate for reasons that have nothing to do with performance.
A defensible peer set controls for the things that drive financial structure regardless of quality. Budget size matters most, because a five million dollar organization and a five hundred thousand dollar organization have fundamentally different cost structures and neither tells you much about the other. Mission and program model matter next, since a direct service organization and an advocacy organization with the same budget spend money in completely different shapes. Geography matters for anything involving salaries or occupancy, because a comparison that ignores regional cost differences will mislead a compensation committee badly. Age and lifecycle stage matter too, since a twenty-year-old organization with an endowment and a five-year-old organization funded entirely by grants are not comparable on reserves.
A practical approach is to build two sets rather than one. The first is a tight comparability set of eight to fifteen organizations matched on size, mission, and region, which is what you use for compensation and expense ratios. The second is a looser aspirational set of organizations doing work you want to learn from, which is what you use for questions about program model, vendor choice, and revenue diversification. Keeping them separate prevents the common failure where an organization benchmarks its salaries against a national leader with ten times its budget and concludes it is failing.
The classification codes on the filings help here. Organizations are assigned activity classifications that group them by field, and filtering by classification alongside budget range and state gives you a starting list that is neutral rather than selected. From there you can remove organizations that are clearly not comparable and document why, which matters if the analysis is going to a board. An undocumented peer set invites the objection that you chose organizations to support a predetermined conclusion, and that objection is usually unanswerable after the fact.
Comparability Checklist
Document each of these before you present results
- Budget range: peers within roughly half to double your total expenses
- Program model: direct service, advocacy, grantmaking, or membership
- Geography: comparable labor market and cost of living
- Fiscal year: note where peers have different year ends
- Filing year: use the same tax year for every organization, not the most recent available for each
- Exclusions: a written note on any organization you removed and why
Using AI Without Inventing Numbers
This is the section that determines whether the whole exercise is trustworthy. Language models are excellent at reading a document and pulling out requested values, and they are also capable of producing a confident, well-formatted, entirely fabricated figure when the value is not where they expected it. The failure is not random noise that you would notice. It looks exactly like a correct answer.
The single most effective safeguard is to require citation of location. Instead of asking for total program service expenses, ask for total program service expenses along with the part and line number where the figure appears and the exact text surrounding it. A model that has to point at the source is far less likely to invent a value, and more importantly, you can verify a claim in seconds rather than rereading the filing. Any figure that comes back without a location should be treated as absent rather than as a number worth checking.
The second safeguard is to insist on an explicit not-found response. Models default to helpfulness, and helpfulness in this context means guessing. Telling the model directly that returning a null value is the correct and preferred answer when a figure is absent changes behavior substantially. Small organizations filing the short form will not have most of the fields you are asking for, and a workflow that cannot say so will silently fill the gaps.
The third safeguard is arithmetic validation outside the model. Once you have extracted a table, check that revenue components sum to total revenue, that the three functional expense columns sum to total expenses, and that the change in net assets equals revenue minus expenses. These checks cost nothing in a spreadsheet and catch a large share of extraction errors, including the ones a human reviewer would miss. Any row that fails a check goes back for manual review rather than into the analysis.
The fourth is sampling. Pick three organizations at random from the finished table, open their filings, and verify every extracted figure by hand. If all three are clean, your error rate is probably acceptable. If any of them has an error, the workflow needs fixing before the results go anywhere near a board packet. This discipline is the same one that makes any AI-assisted data work defensible, and it applies equally to the data cleanup work most organizations face in their own systems.
One more practical note on document handling. Long filings can exceed what a model handles well in a single pass, and accuracy degrades when a model has to locate a small figure inside a very large document. Splitting the filing and asking for specific sections, or working from the structured XML where the relevant fields can be isolated directly, produces noticeably better results than dropping the entire document in and hoping.
A Verification Loop That Works
Four checks, applied in order, before any figure is used
1. Require a source location for every value
Part, line, and surrounding text. Figures without a location are treated as missing, not as candidates for review.
2. Allow and expect null
State explicitly that not found is a correct answer. Short-form filers and small organizations will legitimately lack most fields.
3. Validate arithmetic in the spreadsheet
Components must sum to totals. Change in net assets must reconcile. Failures go to manual review rather than into the analysis.
4. Hand-verify a random sample
Three filings, every figure, checked against the original. Any error means the workflow gets fixed before the results are used.
The Traps That Ruin Otherwise Good Analysis
Even with clean extraction and a defensible peer set, several interpretive errors show up repeatedly. They are worth naming because each one produces a conclusion that feels rigorous and is wrong.
The most common is treating the program expense ratio as a measure of quality. Watchdog conventions have trained the sector to read a high program percentage as evidence of efficiency, and the number is genuinely easy to compute from the functional expense table. But the allocation between program and administration reflects accounting policy as much as operating reality, and organizations that invest seriously in evaluation, technology, or staff development often report worse ratios precisely because they are doing the things that make programs effective. Using the ratio to rank peers rewards the organizations with the most aggressive allocation policies. It is fine as a descriptive statistic and dangerous as a judgment.
The second is mixing filing years. Because organizations file on different schedules and extensions are routine, the most recent filing available for one peer may cover a year that ended two years before the most recent filing for another. Comparing them produces a table that looks consistent and is not, and the effect is large when the period spans a funding disruption. Pick one tax year, use it for everyone, and accept that this means working with older data than you would like.
The third is over-reading a single year. Nonprofit finances are lumpy. A capital campaign, a bequest, a property sale, or a one-time government award can distort every ratio in a filing. Three years of data for each peer costs little additional effort and prevents the embarrassment of building a strategy around an anomaly. Where a figure looks surprising, the anomaly is usually the explanation.
The fourth is forgetting that this is a two-way mirror. Your own filing is equally public, equally readable by a model, and equally likely to be pulled up by a funder, a journalist, or a prospective executive candidate. Organizations that take peer 990 analysis seriously usually come away with a sharper view of what their own filing communicates, which is a useful reason to invest in the narrative sections of your own return rather than treating the whole document as a compliance chore. The related question of where AI genuinely helps with your own filing is covered in more depth in the guide to what you can and cannot automate on the 990.
The fifth is confusing benchmarking with competitive analysis. They answer different questions and require different peer sets. Benchmarking asks whether your structure is normal for organizations like yours. Competitive analysis asks who else is pursuing the same funders and constituents, which often means looking at organizations of very different sizes. The distinction is developed further in the discussion of how competitor analysis differs from peer benchmarking.
A Starting Workflow You Can Run This Month
The version of this work that actually gets done is small. A first pass with ten organizations and four data points is more valuable than an ambitious sector study that never finishes, and it teaches you where the friction is before you commit to anything larger.
Start by writing down the decision the analysis is meant to inform. If the answer is that a compensation committee meets in October, your data points are compensation, hours, and budget size, and everything else is a distraction. If the answer is that the board wants to understand revenue risk, your data points are the revenue breakdown across three years. A benchmarking project without a decision attached produces a document nobody reads.
Then build the peer list and write down the inclusion rule before looking at any numbers. Ten to fifteen organizations matched on size, mission, and region is enough. Pull each filing for a single common tax year, preferring the structured version where it exists. Extract the specific fields you need with location citations, load them into a spreadsheet, run the arithmetic checks, and hand-verify three filings at random.
Finally, present the result as a range rather than a verdict. Boards and executives respond well to a statement that peer executive compensation for organizations of comparable size and region falls between two figures, with a note on how many organizations were reviewed and which were excluded. They respond badly to a single number presented without provenance, and they should, because a number without provenance is indistinguishable from a guess. The value of this whole exercise is that your guesses become documented ranges, and documented ranges are what allow a board to make a defensible decision.
The Ten Organization First Pass
A complete cycle in a few hours of work
- Write the decision the analysis will inform, in one sentence
- Define the inclusion rule, then build the peer list from it
- Pull one common tax year for every organization
- Extract four fields maximum, each with a source location
- Run arithmetic checks and hand-verify three filings
- Report a range with methodology, not a single number
Conclusion
The nonprofit sector has an unusual transparency regime. Organizations that would never share a budget with a peer are required to publish one every year, in a standard format, for anyone to read. That openness was designed to serve accountability, and it does, but it also creates a shared resource that most organizations simply never use. The barrier was never permission. It was the hour it takes to read one filing carefully.
Removing that barrier changes what a small organization can know. A development director with no research budget can now assemble a defensible compensation range in an afternoon. A finance committee can see how its expense structure compares to fifteen similar organizations rather than to a national average that describes nobody. A board considering a new program can look at what peers pursuing the same model actually spend and how they fund it. None of this requires new data. It requires only that someone finally read what has been published all along.
The discipline that makes it work is unglamorous: a documented peer set, a single common filing year, source citations on every figure, arithmetic checks, and a hand-verified sample. Those steps are what separate a credible analysis from a confident-looking table of numbers that a model partly invented. They cost an hour and they are the entire difference between a benchmarking exercise a board can act on and one that quietly misleads everyone who reads it.
Start narrow. One question, ten organizations, one year, four fields, verified. The first pass will surface something your team did not know, and the mechanics will be familiar enough by the second pass that the annual version takes a fraction of the effort. This is one of the few places where AI makes a genuinely new capability available to organizations that could never have afforded the consultant version, and it is worth claiming.
Turn Public Filings Into Usable Intelligence
We help nonprofits design AI workflows that are accurate enough to bring to a board, with the verification built in from the start.
