Back to Articles
    Communications & Marketing

    Producing a Nonprofit Podcast With AI and a Two-Person Team

    Most nonprofit podcasts die somewhere around episode seven. Not because the interviews were bad or the audience never showed up, but because the two people producing it were also running the newsletter, the social calendar, the annual report, and the emergency press response, and the podcast was the only thing on that list nobody would notice was missing. The question is not whether AI can make podcast production faster. It plainly can. The question is which parts of the work it removes, which parts it quietly degrades, and whether what remains is a workload two people can carry for two years instead of two months.

    Published: September 6, 202614 min readCommunications & Marketing
    A small nonprofit communications team editing podcast audio with AI transcription tools

    Podcasting has stopped being an experiment for nonprofit communications teams and has become a normal channel with normal expectations. Roughly half of American adults now listen to podcasts monthly, and mission-driven organizations have found three distinct ways to use the format: producing their own show, placing staff experts as guests on someone else's, and monitoring what gets said about their issue on shows they do not control. Each of those is a different amount of work, and only one of them requires you to become a producer.

    The economics have changed too. Transcript-based editing, automatic filler word removal, one-click noise reduction, and AI-generated show notes have collapsed the post-production time for a typical interview episode from most of a workday to something closer to an hour. Tools that generate show notes, blog posts, social copy, and newsletter sections from a single upload now do work that a small organization would previously have had to hire out or skip entirely. That shift is real, and it is the reason a two-person team can credibly attempt this in 2026 when they could not in 2019.

    But the shift is uneven. AI has taken large bites out of editing, transcription, and repurposing. It has taken almost nothing out of booking guests, preparing questions, holding a real conversation, chasing consent forms, and deciding whether an episode is worth publishing. Those remaining tasks are the ones that consume a small team's scarcest resource, which is attention rather than hours. A production plan that assumes AI shortened everything proportionally will fail in exactly the same way the old plans failed.

    What follows is a working plan for two people: how to decide whether you should be doing this at all, which formats survive contact with a small team, a realistic weekly workflow, where AI earns its place and where it does not, the recording quality problems no software can repair, the consent and rights questions that catch nonprofits specifically, and how to measure whether the show is doing anything for your mission.

    First, Decide Whether You Should Have a Podcast

    A podcast is a subscription commitment. Unlike a blog post, which sits there and accumulates search traffic whether or not you write another one, a podcast asks people to expect you on a schedule. Missing that schedule is visible in a way that missing a blog post is not, and an abandoned feed is a worse artifact than no feed at all, because it tells anyone who finds it that your organization starts things it does not finish.

    So the first question is not about equipment or software. It is whether you have something that genuinely belongs in audio. Audio is uniquely good at a few things: conversation between people who know something, the sound of a place, the voice of a person telling their own story, and explanation that benefits from a human tone. It is uniquely bad at anything visual, anything data-heavy, and anything that would be faster to read. If your idea is essentially your annual report read aloud, the format is fighting you.

    The second question is who the show is for and what you want them to do. Nonprofit podcasts tend to serve one of four audiences, and they are not interchangeable. A donor and supporter audience wants access, impact, and the feeling of being an insider. A field or peer audience wants practical knowledge and comes for expertise. A beneficiary or community audience wants information they can use. And a general public audience wants a story worth their time. A show that tries to serve all four will bore all four.

    The third question is the honest one about capacity. Two people producing a fortnightly interview show are committing to roughly six to ten hours a month once the workflow is established, and considerably more for the first several episodes while everything is still being invented. If that time is not being taken from something, it does not exist. Deciding in advance what the podcast replaces is more useful than deciding what it adds, and it is a conversation worth having with your executive director before episode one rather than after episode seven.

    There is also a legitimate answer that is not a podcast at all. Getting your executive director booked as a guest on four established shows in your sector will very likely reach more of the right people than twelve episodes of your own show reaching your existing email list. Guesting costs an hour of preparation and an hour of recording, with no production burden whatsoever. For many organizations that is the correct strategy, and the fact that it is less satisfying to own does not make it worse.

    Signs you should not start a podcast yet

    Better to know now than at episode seven

    • Nobody can name the specific audience beyond "our supporters"
    • The idea came from a board member who will not be producing it
    • You cannot list ten guests or topics without straining
    • Your newsletter is already going out late most months
    • The content would work better as writing and you know it
    • No one has agreed what the show replaces on the comms calendar

    Formats a Two-Person Team Can Sustain

    Format is the single biggest determinant of whether the show survives, and it is chosen almost entirely at the start. The differences in production load between formats are enormous, far larger than any difference AI tooling makes within a format. A narrative documentary episode and an interview episode of the same length can differ by a factor of ten in the hours required.

    The interview show. One host, one guest, thirty to forty-five minutes, published fortnightly. This is the workhorse for a reason. Preparation is a page of questions, recording is a single session, and editing is mostly removing false starts and the first ninety seconds where everyone settles down. It scales down gracefully when a week goes badly. Its weakness is that it is the most common format in existence, so the show lives or dies on whether your guests are genuinely worth hearing.

    The short explainer. Eight to twelve minutes, one host, one topic, scripted. Excellent for organizations whose value is expertise rather than access: policy shops, legal aid organizations, health nonprofits, associations serving a professional field. Scripting takes longer than people expect, but recording is fast, editing is minimal, and the episodes have a long useful life because they answer questions people keep asking.

    The internal conversation. Two colleagues discussing what they are actually working on, lightly structured, twenty to thirty minutes. Cheapest format to produce because there is no booking, no external scheduling, and no consent complexity. It works when the two people genuinely have something to say and have some chemistry, and it is dreadful when they do not. Be honest about which situation you are in.

    The limited series. Six or eight episodes on one theme, produced as a season, then finished. This is the format most nonprofits should consider and fewest do. It sidesteps the entire sustainability problem by having a defined end. You can produce a season around a campaign, an anniversary, a program area, or a policy fight, publish it, and then decide with full information whether to do another. Nobody is disappointed when a series that was always eight episodes ends at eight.

    The narrative documentary. Field recordings, multiple voices, scoring, a written and rewritten script. This is the format everyone admires and almost no two-person team can maintain. A single strong documentary episode released once a year as a centerpiece can be worth more than a mediocre monthly show, but it should be treated as a project with a budget, not as a podcast schedule. The same discipline that applies to video storytelling production applies here.

    Whatever you pick, publish on a cadence you can beat rather than one you can barely meet. Fortnightly delivered reliably reads as professional. Weekly delivered erratically reads as struggling. And build a buffer of two finished episodes before you launch, because the first thing that will go wrong is a guest cancelling in a week when someone is on leave.

    The Workflow, Start to Finish

    A sustainable production process is one where every step has a defined owner, a defined trigger, and a defined finish. Ambiguity is what kills small-team projects, because ambiguous work gets postponed rather than done. What follows is a workflow for two people producing a fortnightly interview show, with the AI-assisted steps marked, and it can be compressed or expanded for other formats.

    Booking, ongoing. Maintain a running guest list of at least twelve names with contact details and a one-line note on what makes each one interesting. Book three to four episodes ahead. This single habit prevents more failures than any tool, because scrambling for a guest the week of recording is where the schedule breaks. AI is useful for the background research on a prospective guest, pulling together what they have written and said publicly so your host walks in informed.

    Preparation, one hour. A single page: the guest's background, the three things you want the audience to come away with, eight to ten questions in rough order, and one question you are slightly nervous to ask. AI can produce a first draft of the question list from the guest's public work, which is genuinely useful, but the draft will be generic and safe. The good questions come from the human who has read the material and noticed something. Treat the AI list as the floor, not the plan.

    Recording, one hour. Schedule ninety minutes for a forty-five minute conversation. Record locally on both ends where possible rather than relying on the call platform's audio, since a remote recording platform that captures each speaker's local audio separately is the single largest quality improvement available for the money. Send the consent and release language before the session, not after.

    Post-production, one to two hours. Upload, generate the transcript, and edit against the transcript rather than the waveform. Remove the settling-in period, the tangent that went nowhere, the filler words if they are dense enough to distract, and any section the guest asked to remove. Apply noise reduction and level the two speakers. Resist the urge to over-edit, because a conversation edited to remove every hesitation stops sounding like a conversation.

    Packaging, forty-five minutes. Title, episode description, show notes with timestamps, links mentioned, the corrected transcript for the website, and three to five social assets. This is the step AI has changed most dramatically, since a single upload can now produce drafts of all of it. It is also the step where the drafts most need a human pass, because AI-written show notes tend toward the enthusiastic and the vague.

    Publishing and distribution, thirty minutes. Schedule the episode, post the transcript page, queue the social assets, add the episode to the next newsletter, and send the guest a short note with links and a request to share. That last step takes four minutes and is the most reliable growth mechanism a small show has. Your existing content calendar workflow should absorb the podcast rather than run alongside it.

    A realistic per-episode time budget

    Fortnightly interview show, two people, established workflow

    • Booking and scheduling: 30 minutes, spread across weeks
    • Preparation and research: 60 minutes, partly AI-assisted
    • Recording session: 60 to 90 minutes for both people
    • Editing: 60 to 120 minutes with transcript-based tools
    • Show notes, transcript, and social assets: 45 minutes
    • Publishing and promotion: 30 minutes
    • Total: roughly 5 to 7 hours per episode across two people

    Where AI Genuinely Earns Its Place

    The useful mental model is that AI is very good at anything downstream of a finished recording and mediocre at anything upstream of it. Everything that happens after you stop recording is transformation of material that already exists, which is what these tools do well. Everything before is judgment, relationship, and preparation, which is what they do badly.

    Transcription is solved. Automatic transcription with speaker labels is now accurate enough on clean audio that correcting it takes minutes rather than hours. This matters for editing, for accessibility, for search, and for repurposing, and it is the foundation everything else sits on. Accuracy degrades sharply with poor audio, heavy background noise, overlapping speakers, and unfamiliar names, which is the first of several reasons recording quality pays for itself downstream.

    Transcript-based editing is the biggest single change. Editing by deleting text rather than by manipulating waveforms removes the skill barrier that kept podcast production in the hands of people who knew audio software. A communications coordinator who can edit a document can now edit an episode, and reported time savings in the range of half to two-thirds of previous edit time are consistent with what small teams find in practice. This is the feature that makes the whole enterprise viable for two people.

    Cleanup is reliable and unglamorous. Filler word removal, background noise reduction, room echo reduction, and loudness matching between two speakers recorded in different rooms are all now one-click operations that used to require real expertise. Used with restraint they make an amateur recording sound competent. Used aggressively they make a human conversation sound like a synthesized one, so apply less than the tool offers.

    Repurposing is where the time is actually won. A forty-minute conversation contains a newsletter section, two or three social posts, a blog article, a quote graphic, and often a paragraph for a grant report. Generating first drafts of all of those from the transcript in one pass turns the podcast from a cost center into a content engine, which is the argument that justifies the show internally. Our guide to repurposing content with AI covers the mechanics, and the same discipline described in scaling your storytelling capacity applies directly.

    Where it underperforms. AI-drafted episode titles are consistently bland, because they describe the topic rather than create curiosity. AI-drafted show notes over-promise, using words like powerful and inspiring that no human wrote and no listener believes. AI-generated interview questions are safe by construction and will never produce the moment where a guest says something they had not planned to. And AI cannot tell you whether an episode is good, which is the judgment that matters most and the one a two-person team should reserve entirely for itself.

    Hand to AI

    High volume, low judgment

    • Transcription and speaker labelling
    • Filler word removal and noise cleanup
    • Timestamped chapter markers
    • First drafts of show notes and summaries
    • Pulling quotable moments out of the transcript
    • Background research on prospective guests

    Keep with a person

    Low volume, high judgment

    • Deciding who to invite and why
    • The questions that go beyond the obvious ones
    • The conversation itself, including the follow-ups
    • Judging whether an episode should be published
    • Anything involving a client, patient, or beneficiary
    • The final title and episode description

    The One Thing Software Cannot Fix

    Audience tolerance for mediocre audio is low and falling. People will forgive an amateur host, a rambling structure, and a bad cover image. They will not forgive audio that is fatiguing to listen to, because listening is the entire product. AI cleanup tools have narrowed the gap between good and bad recordings considerably, but they work by removing and reshaping what is there. They cannot add information the microphone never captured.

    The three problems that cannot be repaired afterward are heavy room reverb, clipping from a recording level set too high, and two people captured on a single distant microphone. Reverb removal tools have improved but leave a distinctive hollow artifact. Clipped audio is information that no longer exists. And two voices on one track cannot be balanced independently, which is the difference between an episode that sounds professional and one that sounds like a meeting recording.

    Fixing all three is cheap. A pair of dynamic USB microphones costs less than a single day of consulting time, and a dynamic microphone is far more forgiving of an untreated office than a condenser. Record in the smallest carpeted room available, ideally one with soft furnishings, and avoid the conference room with the glass wall and the hard table, which is the worst space in most nonprofit offices and the one people instinctively book. A closet full of coats genuinely outperforms a boardroom.

    For remote guests, use a platform that records each participant locally at full quality and uploads afterward, rather than capturing the compressed call audio. This is the highest-value single decision in the whole setup, and it costs a modest monthly subscription. Send guests a two-line note beforehand: use wired headphones, sit in a small soft room, and close the window. Most people will do it, and it makes more difference than anything you can do in post.

    On synthetic voices, a brief note. Voice cloning and AI narration have reached the point where they are usable for pickups and corrections, and some teams use a cloned host voice to fix a mispronounced name rather than re-recording. That is a defensible use if the host consented to the clone. Generating substantial narration in a synthetic voice, or worse, in a voice resembling a real person who did not agree to it, is a different matter for a nonprofit whose currency is trust. If you use synthetic voice for anything a listener would reasonably assume was a person, disclose it.

    Consent, Rights, and the Questions Nonprofits Get Wrong

    Podcast production raises consent questions that ordinary communications work does not, and nonprofits face a sharper version of them than commercial producers do because their guests are often people connected to services rather than professionals promoting a book. Getting this wrong is not merely a legal exposure. It is a relationship failure with someone your organization exists to serve.

    Written release, before recording. A short plain-language release covering the recording, editing, publication, distribution on third-party platforms, and use of excerpts in promotion. Send it in advance so nobody is signing under the social pressure of a scheduled session. Keep it readable, because a release a guest does not understand is not meaningful consent even where it is technically enforceable.

    A different standard for people you serve. When the guest is a client, patient, program participant, or family member, the ordinary release is not sufficient. They need to understand that the episode is permanent, publicly searchable, and effectively impossible to fully retract once distributed. They should be offered the option to review before publication, the option to use a first name only, and a genuine ability to say no without any effect on the services they receive. If the person is a minor, involve a guardian and consider whether the story needs their voice at all.

    AI processing and third-party terms. When you upload a recording to a transcription or editing service, you are sending someone's voice and words to a vendor. Read what that vendor does with uploaded content, particularly whether it is used to train models and how long it is retained. For a conversation with a policy expert this is a minor consideration. For a conversation with a domestic violence survivor or an undocumented client it is a serious one, and some recordings should never go through a consumer tool at all. The ownership questions this raises are covered in our piece on who owns the transcript.

    Music and clips. Theme music must be licensed for podcast use, and the licence needs to cover distribution rather than only internal use. Free does not mean cleared. Using a few seconds of a commercial song under a belief that short excerpts are automatically permissible is a common and incorrect assumption. Licensed production music costs very little and removes the problem entirely.

    Retraction requests. Decide in advance what happens when a guest asks you to remove an episode six months later. Having a policy, even a simple one that says you will remove it from your feed and website while acknowledging that copies may persist elsewhere, is far better than improvising under pressure. Write it down before you need it.

    Before you hit record

    A six-item pre-flight check

    • Signed release on file, sent and returned in advance
    • Guest told how the recording will be edited and published
    • Vendor terms checked if the content is sensitive
    • Both ends recording locally, levels checked on a test clip
    • Theme music licensed for public distribution
    • A written policy for later removal requests

    One Recording, a Month of Content

    The strongest internal argument for a podcast is not the podcast. It is that a recorded conversation with a knowledgeable person is a raw material that feeds every other channel you run, and that a two-person comms team producing two episodes a month has effectively created a content pipeline rather than added a channel. Framing it that way changes the budget conversation and the leadership conversation.

    From one forty-minute episode you can reasonably produce a full transcript page that earns search traffic on the guest's name and topic, a newsletter section built around the single best exchange, three short social posts using direct quotes, an audiogram clip of a strong ninety-second moment, a blog article that develops one idea the conversation raised, and often a paragraph of evidence for a grant report or board update. Each of those has an AI-generated first draft available within minutes of the transcript existing.

    Two cautions. First, drafts are drafts. Publishing AI-generated show notes without editing produces the flat promotional register that readers now recognize instantly, and it undermines the credibility the conversation itself built. Second, everything derived from the episode must sound like your organization rather than like a language model, which is a problem worth solving once at the system level rather than repeatedly at the draft level. Our guide to maintaining brand voice across AI-generated content covers how to build that constraint into the workflow.

    The transcript deserves particular attention because it is the most undervalued asset in the whole process. It makes the episode accessible to deaf and hard of hearing audiences, which is both a legal consideration for many organizations and a straightforward matter of who you are willing to exclude. It makes the content searchable, which is the only way anyone will find episode eleven eighteen months from now. And it is the substrate every repurposed asset is generated from. Publishing a clean, corrected transcript alongside every episode is the highest return per minute in the entire workflow.

    What one episode should produce

    Drafted by AI, finished by a person

    • A corrected, published transcript page
    • Show notes with timestamps and every link mentioned
    • A newsletter section built on one strong exchange
    • Three social posts using verbatim quotes
    • One short captioned audiogram or video clip
    • A blog article developing one idea further

    Measuring Whether It Is Working

    Download counts are the metric everyone reports and the one that answers the least. A download tells you a file was requested, not that anyone listened past the introduction. Worse, downloads flatter shows with large existing email lists and punish shows reaching precisely the right two hundred people. A nonprofit podcast reaching four hundred committed listeners in its field can be worth vastly more than one reaching four thousand casual ones, and a metric that cannot express that will lead you to make the show worse.

    Consumption rate, which is how far into an episode the average listener gets, is far more informative. A show where most listeners finish is working, whatever its size. A show where listeners drop out at eight minutes has a structural problem in its opening, which is diagnosable and fixable. Watching that number over a season tells you more about the quality of your format than any amount of feedback from colleagues who feel obliged to be encouraging.

    Beyond the platform numbers, the outcomes that justify the show are usually organizational rather than audience-level. Did a guest introduce you to someone. Did a funder mention an episode. Did a journalist find you through it. Did a new board prospect say the show is how they first understood the work. Did the transcript page bring in search traffic for a term you care about. Did the repurposed content reduce the time spent producing the newsletter. These are trackable if you decide to track them, and mostly invisible if you do not.

    Set a review point before you launch, ideally at ten episodes or at the end of a season, and agree in advance what would cause you to stop. Small teams are extremely bad at ending things, and a podcast that continues on inertia consumes the capacity that a better idea needs. Deciding to conclude a series at eight episodes because it did what it was for is a success, and should be described that way internally. The measurement discipline described in our piece on calculating AI return on investment applies here as well, since the time AI saves in production only counts if you can say what it was spent on instead.

    Conclusion

    The reason a two-person comms team can produce a credible podcast in 2026 is narrow and specific. Transcript-based editing removed the skill barrier, automatic transcription removed the tedium, and AI-drafted show notes and repurposing removed the packaging work that used to consume an afternoon per episode. That is a genuine change, and it moves the total production burden from something only a dedicated producer could carry to something two people with other jobs can fit into a fortnight.

    What has not changed is everything that determines whether the show is any good. Someone still has to know who is worth interviewing, prepare properly, ask the question that was not on the list, notice when a conversation went somewhere unexpected, and decide honestly whether the episode deserves to be published. Those hours are the show. Tools that shorten the other hours are valuable precisely because they protect these ones.

    So choose a format you can sustain rather than one you admire, build a buffer before you launch, record properly because that is the one thing software cannot repair, treat consent as a relationship question rather than a paperwork question, publish the transcript every time, and set the date at which you will decide whether to continue. Do that, and the podcast becomes a content engine that feeds every other channel you run. Skip it, and you have another abandoned feed with seven episodes and a hopeful description.

    Build a Comms Workflow Your Team Can Actually Sustain

    We help small nonprofit communications teams design AI-assisted production workflows that fit the staff they have, not the staff they wish they had.