Short answer: an AI agent can run the mechanics of a community today — drafting, scheduling, reading, summarising, routing, and answering questions that already have a documented answer. It should not be handed decisions whose failures are expensive and hard to reverse, and the reason is not that models are stupid. It is that they are unreliable in a specific measured way, and that the human review step people rely on as a safety net is itself software that can fail. The third part is the one that usually gets left out: "always keep a human in the loop" is not the correction either. The workable question is which decisions, at what confidence, with what appeal path.
Disclosure: Zealy publishes this page and Zealy sells the thing being discussed. Zealy is a community and quest platform, Zealy ships an MCP server that lets an AI agent administer a Zealy community, and Zealy ships an AI reviewer that decides quest claims. All three appear below, including the place where two of them disagree with each other and the two places where a competitor's design is better than ours.
Last verified: 18 August 2026.
What can an AI agent actually do in a community today?
An AI agent can reliably do the mechanics: drafting posts and replies, scheduling them, reading a channel and summarising it, routing a report to the right person, and answering questions that already have a documented answer. Those are jobs where a wrong output is visible, cheap and easy to correct before anyone is harmed.
That list is not a small thing. Most community work is mechanical, and most community teams are behind on it. If an agent drafts your weekly recap and a person spends four minutes fixing it instead of forty writing it, that is a real win with a bounded downside.
We wrote this post because the question is being answered badly. Searching can ai run my community on 18 August 2026, every result we retrieved was a vendor blog or a vendor product page — heightsplatform.com, mightynetworks.com, social.plus, circle.so, socialglow.com, aismartventures.com. Not one of the results we retrieved answered the question with a "no" or a "partly". Each restated it as a feature list. That is an observation about the results we pulled on one day, not a claim about everything ever written.
The category's own vendors are more honest than their landing pages suggest. Sprout Social's guide to building AI agents for social media, by McCall Lanman, published on April 16, 2026, defines the category by the absence of a person: "An AI agent is a software program that uses a large language model (LLM) as its brain to autonomously complete tasks, make decisions and interact with external tools—without a human directing every step." Further down the same page it tells you to put the person back: "Approval workflows: Route sensitive responses to a human manager before they're sent—this is called human-in-the-loop." And plainly: "Without guardrails, even a well-built agent produces off-brand or harmful outputs."
That is not hypocrisy. It is the actual difficulty of the category showing through a marketing page, and it is worth saying kindly. If you want the term itself pinned down properly, our pillar on agentic marketing judges it against the machinery rather than the ambition.
Why the real limit is unreliability, not capability
The problem with handing an agent a judgement call is not that models are stupid. It is that they are inconsistent, and inconsistency is a harder thing to manage than a known weakness. Research on multi-turn conversation puts most of the degradation in variance rather than in raw ability.
The paper is "LLMs Get Lost In Multi-Turn Conversation" by Philippe Laban, Hiroaki Hayashi, Yingbo Zhou and Jennifer Neville, arXiv:2505.06120, submitted 9 May 2025. Its abstract reports that "all the top open- and closed-weight LLMs we test exhibit significantly lower performance in multi-turn conversations than single-turn, with an average drop of 39% across six generation tasks."
The sentence that matters more is the next one: "Analysis of 200,000+ simulated conversations decomposes the performance degradation into two components: a minor loss in aptitude and a significant increase in unreliability." And the failure mode has a shape — "when LLMs take a wrong turn in a conversation, they get lost and do not recover".
Read the caveats before you use that 39% anywhere. These were simulated conversations across six generation tasks, measured on 2025 models, and 39% is an average across those tasks rather than a property of any product you might buy today. Models have improved since. The durable finding is not the number, it is the decomposition.
Here is why the decomposition is the whole argument. A tool that is simply not good enough yet is a scheduling problem: you wait, or you narrow the task. A tool that is right on average and wrong unpredictably is a management problem, and it is a different one. You cannot staff around it by checking the hard cases, because you do not know in advance which case was the hard one. Community work is multi-turn by nature — a member asks, clarifies, pushes back, adds context three messages later — which is exactly the setting the paper measured.
Discord's human in the loop was a piece of software, and it had a bug
Discord published, in the present tense, that a member of its Trust & Safety team always reviews flagged content before any action is taken. A bug banned more than 8,000 accounts anyway. The human in the loop was a code path, and code paths fail like software, not like people.
TechCrunch reported it on 7 July 2026, in a piece by Lauren Forristal headlined "Discord admits AI moderation bug wrongfully banned users over harmless images". The core of it: "Discord has acknowledged that a bug in its AI moderation system mistakenly banned more than 8,000 users over the past two months, after harmless images—including spreadsheets, chessboards, game textures, as well as white and gray transparent backgrounds—were incorrectly flagged as harmful content."
The scale is worth reading precisely rather than rounding up. "The company confirmed that the issue had been affecting accounts since May, with an additional 200 users banned over the weekend before its team identified and fixed the problem. All affected accounts are currently in the process of being restored."
Now the part that changes how you should think about this. Discord's own support account, quoted by TechCrunch from its thread of 7 July 2026, described the design: "Our systems flag content by matching it against known harmful material. This kind of similarity matching can produce false positives, which is why a member of our Trust & Safety team always reviews flagged content before any action is taken."
And TechCrunch's own summary of what happened: "A human moderator reviews the content, but a bug caused the system to immediately ban affected accounts."
Both of those are true at once. The review step was designed, was documented, and was described publicly in the present tense — and it did not run. Discord did not lie, and this post is not the argument that it did. The point is narrower and more uncomfortable: a human-in-the-loop is a piece of software and fails like one. It is a conditional, a queue, a service call, a feature flag. Every one of those can be wrong, and when one is wrong the safety property you were relying on quietly stops existing while the architecture diagram still shows it.
Also notice what the failing inputs were. Spreadsheets. Chessboards. Grey transparent backgrounds. Nothing about that list looks like an edge case anyone would have written a test for, which is the ordinary condition of a similarity-matching system meeting the real world.
If your entire assurance to your members is "a person checks", you have not made a claim about people. You have made a claim about a code path, and you should be able to say when it last ran.
Why "always keep a human in the loop" is not the lesson either
The tidy correction to vendor hype is to say a person must approve everything. The Digital Trust & Safety Partnership, writing for the companies that do this at scale, does not say that. It reports that automated enforcement without human review is common when confidence is high, and that human review is not infallible.
The source is "Best Practices for AI and Automation in Trust & Safety", Digital Trust & Safety Partnership, September 2024. It is the most authoritative document in this field that anyone can download, and it is more comfortable with automation than most people expect.
On whether you can avoid automation at all: "AI and automation are not a panacea—they come with several challenges and limitations and human involvement remains necessary. However, addressing content- and conduct-related abuse at scale is not possible for many digital products and services without some level of AI and automation." Note the scope in that sentence — "for many digital products and services", not for everyone. A community of four hundred people is not in the same position as a platform of four hundred million.
On what the member companies actually do: "all DTSP partner companies interviewed for this report continue to rely on both automated tools and human review and oversight to take appropriate action, especially where a more nuanced approach to assessing content or behavior is required."
And then the two sentences that stop this article from replacing vendor hype with a tidier simplification:
- "It is common for trust and safety teams to enable automated enforcement action without human review when a system is highly confident in its assessment."
- "It is often assumed that a human reviewer will be more accurate and reliable than an AI system; however, human review is not infallible."
The report also names the shape of the tradeoff you are actually making when you tune any of this: "Optimizing for recall (catching more potential violations) can increase both true positives and false positives, potentially leading to over-action. Conversely, optimizing for precision (more accurate detection of violations) can reduce false positives but might also increase false negatives (under-action)."
So the design that the report describes is neither "let the agent decide" nor "a person approves everything". It is confidence-tiered: "When the confidence level is lower, the automated system may take a temporary enforcement action (e.g., temporarily downranking the content or reducing the discoverability of the account), and route the case for human review to approve or modify the action."
That is a genuinely different mental model from the one most community owners carry around, and it is better. The low-confidence path does not stop and wait. It takes the smallest reversible action available and then asks a person. Downranking is not banning. A quest claim held for review is not a claim rejected. If you can invent a middle action, you no longer have to choose between speed and safety on every single case.
Which decisions you can hand to a machine today
Sort the decisions by what happens when the machine is wrong, not by how hard they look. A wrong draft gets edited. A bad summary gets re-read. A wrong ban costs a member their account, and costs you a support thread, a reputation, and possibly a ruling. Reversibility is the axis that matters.
This table is Zealy's judgement, not a finding from any of the sources above. It is what we would tell a community owner who asked us on 18 August 2026.
| Decision | Our answer today | What has to be true before you automate it |
|---|---|---|
| Drafting a post, quest description or reply | Yes | A person reads it before it ships. Nothing else is required. |
| Summarising a channel, a week, or a submission queue | Yes | The raw material stays reachable, so a reader can check the summary against it. |
| Routing a report, question or claim to the right person | Yes | Misrouting is cheap and visible, and the queue is not a black hole. |
| Answering a question that has a documented answer | Yes, with care | The answer exists in writing somewhere you control, and a wrong answer can be corrected publicly and fast. |
| Tagging, ranking or ordering a queue for a human | Yes | The human still sees the items the model ranked last. Ordering is not filtering. |
| Approving or rejecting a submission that pays out | Only with an appeal path | The member is told the reason, a person can reverse the verdict, and someone is watching the reversal rate. |
| Removing content or restricting an account | Only with a reversible middle action | A low-confidence path exists that does something smaller than the full action, as the DTSP report describes. |
| Banning a member outright | No | We cannot name a confidence level that makes this safe without a human, and the Discord incident is why. |
| Spending money or moving rewards | No | Nothing in this article supports automating it, and the failure is not reversible by apologising. |
The pattern down that middle column is not "hard things are for humans". Ranking a queue is a harder cognitive task than banning someone. It is at the top of the list because ranking wrong costs you nothing you cannot undo in a second, and banning wrong costs you a person.
One practical note that does not fit in a table. Almost every community that automates a decision forgets to measure the reversal rate afterwards, which means the appeal path exists but nobody is reading what it produces. If you build one thing from this article, build the counter. If you are sizing the human side of that work, how to delegate community tasks covers the labour cost that never shows up on a pricing page.
Zealy runs two AI paths, and they answer this question differently
Over Zealy's MCP server, nothing an agent changes takes effect until a person says so. Inside the Zealy product, Zealy's AI review decides a quest claim outright: it marks the claim successful or failed and posts its reason to the member. Two AI paths in one company, opposite answers to the same question.
Take the second one first, because it is the one that argues against us. On the Plus and Enterprise plans you can switch on AI review for a quest and write a prompt describing what a valid submission looks like. Zealy's own documentation states the consequence without softening it: the AI decides the claim. It does not leave a recommendation for a human to confirm. It marks the claim successful or failed, and the reason it gives is posted as the review comment the member reads in their inbox. Its verdicts are attributed to a Zealy AI reviewer that shows up alongside your team in the top-reviewers analytics. If you want a person to make the final call, you leave AI review off for that quest. That is all on Zealy's reviews documentation, and it has been published there the whole time.
Measured against the table above, that is an automated approve-or-reject on a decision that awards XP and role rewards — and, since Zealy's documentation says it is the community's responsibility to distribute any other kind of reward, on the decision that determines who the community then owes those too. The mitigations are real — a reviewer can undo any verdict, which resets the claim to pending, and the member is told the reason rather than being silently failed — but the decision is made by the model, and the appeal starts with the member noticing.
Zealy's other AI path took the opposite position. Over the Zealy MCP server, an agent reads freely and cannot change anything without a person approving that exact change first. The mechanics of that gate, tool by tool, are in how to audit an AI agent's access to your community, which enumerates the surface rather than describing it.
We are not going to resolve that contradiction inside a blog post, and pretending it is a coherent single philosophy would be worse than printing it. Two teams, two threat models, two answers. What we can say is that the split is not random: the MCP path is an outside agent acting with an administrator's authority across a whole community, and the review path is a scoped judgement on one submission against a prompt the community wrote itself. Those are different risks. Whether they are different enough to justify opposite defaults is a fair question to put to us, and if you are evaluating Zealy you should put it.
The general point stands past our own product. When a vendor tells you their AI keeps a human in the loop, ask which of their AI features they mean. Companies do not have one policy. They have one policy per team that shipped something.
Where Discord and Cove do this better than Zealy
Discord's AutoMod runs inside Discord itself and blocks a message before anyone sees it; Zealy does not sit in Discord's message path and cannot stop anything before it is posted. Cove ships a Moderator Console built on the premise that its own automation will miss things. Zealy has no equivalent.
Those are two different advantages and both are worth understanding.
Discord's is positional. Discord's own Auto Moderation documentation describes the feature as one "which allows each guild to set up rules that trigger based on some criteria", and its BLOCK_MESSAGE action as one that "blocks a member's message and prevents it from being posted". That is the whole advantage in one clause: the intervention happens before publication rather than after. Zealy connects to Discord, awards roles and verifies tasks; it is not the chat server, and no amount of AI on our side changes where we sit. If your problem is "a bad message must never appear", the tool that solves it is the one already in the pipeline. Our survey of Discord engagement tools covers what else belongs in that layer.
Cove's is architectural, and it is the more interesting one. Cove describes itself in its own documentation as "Cove is a no-code platform for Trust & Safety", and then writes this about its own automation: "Some harmful content will inevitably slip through your Automated Rules, and it'll get reported/flagged by your users. With our Moderator Console, your moderators can review those reports and take action on them directly in our UI."
Read the word "inevitably" again, in a vendor's own product documentation, about the product they are selling you. Cove designed a surface whose entire reason for existing is that the automated layer will be wrong, and built the workflow for catching that as a first-class part of the product rather than a support ticket.
Zealy has an undo on a review, which is not the same thing. Undo is a control for a person who already knows a verdict was wrong. What Cove has is a place where the misses arrive. If you are choosing a platform partly on how it handles its own failures, that distinction is the one to test, and you should test it on us too. We have written a wider comparison of community management software if you are earlier in that decision, and what Zealy is if you want the plain description of our own scope.
Who answers when the agent is wrong
You do. In February 2024 a Canadian tribunal ordered Air Canada to pay damages after its website chatbot gave a grieving passenger wrong information. Air Canada argued, in the tribunal's words, that the chatbot was a separate legal entity responsible for its own actions. The tribunal awarded CA$812.02 in total, CA$650.88 of it damages.
The Register reported it on 15 February 2024, in a piece by Katyanna Quach headlined "Air Canada must pay damages after chatbot lies to grieving passenger about discount". The tribunal member deciding the case was Christopher Rivers, and The Register quotes the ruling characterising the airline's position: "In effect, Air Canada suggests the chatbot is a separate legal entity that is responsible for its own actions."
And the finding, which is the sentence every community owner deploying an agent should have somewhere they can see it: "It should be obvious to Air Canada that it is responsible for all the information on its website. It makes no difference whether the information comes from a static page or a chatbot."
One caveat on how we are citing this. We are quoting The Register's reporting rather than the ruling itself, which we did not read, and we are not printing a case citation.
The amounts are small and the principle is not. An agent speaking in your community is you speaking in your community. Members will hold you to what it said, and at least one tribunal has too. That applies to a refund policy it invents, a rule it states that you never wrote, and a ban it hands out. "The AI did it" was tested as a defence and it lost.
How to set the boundary for your own community
Write down the decisions your agent will make, then write down what happens when each one is wrong and who fixes it. If you cannot name the person who fixes it, the decision is not ready to automate. That single exercise is worth more than any confidence threshold you will pick blind.
A short version of what the evidence above supports, in the order we would do it:
- Automate the mechanics immediately. Drafting, scheduling, summarising, routing, tagging. The downside is a bad draft, and you already know how to handle a bad draft.
- For anything with a consequence, invent the middle action before you invent the threshold. The DTSP report's low-confidence path takes a small reversible step and routes the case onward. A held claim, a downranked post, a restricted-but-not-removed account. If your only two options are "act" and "wait", every threshold you pick will be wrong in one direction.
- Treat your human-review step as software, and test it like software. Discord's existed on paper and did not run. Ask when yours last actually fired, and what would tell you if it stopped.
- Count the reversals. An appeal path nobody measures is a complaint form. The reversal rate is the only honest read you will get on whether the automated decision is good enough, and it costs almost nothing to record.
- Assume you own everything it says. Air Canada tried the other position in front of a tribunal.
The question in the title has a real answer and it is not a shrug. An AI agent can run the mechanics of your community today, well, and you should let it. It cannot yet be the thing that decides who gets removed, who gets paid, or what your policy is — not because it is stupid, but because it is unreliable in a way you cannot predict per case, and because the human safety net you would put behind it is a piece of software with its own failure modes. Design for the case where both are wrong, and you will have built something better than either the hype or the backlash would have got you.
