Short answer: almost no gamification research measures retention, so nobody can honestly tell you which mechanic keeps members. The one exception is a pair of Wikipedia studies by the same authors, and they split by population: badges made the most productive contributors less likely to stop, while the follow-up associated the same rewards with lower retention among less-active ones. Beyond that, points and levels raise how much people do without moving how they feel about it; leaderboards are genuinely contested; and streaks are the thinnest-evidenced mechanic of the lot. Pick a platform for the behaviour you need — Discourse for earned permissions, Circle for optional leaderboards, Guild.xyz for gated access, Zealy for verified off-platform tasks.
Disclosure: Zealy publishes this page and Zealy sells one of the tools on it. That is a reason to read the section listing the five gamification features Zealy does not have, and the two places below where Discourse and Circle do something better than we do.
Method: every Zealy claim here was read out of Zealy's own source code on 17 August 2026, not from Zealy's marketing or its documentation, which we have found wrong before. Every research finding is quoted from the paper or the publisher's own abstract record. Where we could only reach an abstract, we say so and print no sample size or effect size from it. Where a vendor's site refused us, we say that too rather than describing features we did not see.
Last verified: 17 August 2026.
What people actually mean by a gamification platform
The phrase "gamification platform" covers three different markets. Search it and you get e-learning tools, employee-performance software and marketing-campaign widgets — Genially, Playable, SC Training, Centrical, Drimify. Community and web3 platforms like Zealy, Galxe and Layer3 did not appear in the results we ran on 17 August 2026.
So if you typed "gamification platform" wanting something to put points and a leaderboard into your Discord or your forum, the results are not addressing you. They are addressing an L&D manager who wants a quiz to feel like a game show, or a marketer who wants a spin-to-win wheel on a landing page. Those are real products solving real problems. They are not this problem.
This post serves the qualified version of the question — gamification for a community — and it does not try to claim the head term. The rest of it is organised by mechanic rather than by product, because the mechanic is the part that either works or does not, and every platform below is some combination of the same six.
If your question is really "which software should I run my community on", that is a different and wider decision, and our survey of community management software covers the five unrelated jobs hiding inside it, with prices.
What the gamification research actually measures — and what it does not
Almost none of the gamification research measures retention. We grepped the three meta-analyses in our review for retention, churn, attrition and dropout and found none of those words. It was measured in essentially one place — Wikipedia, by one pair of authors, twice — with results that split by population. The literature measures motivation, output volume and learning outcomes instead.
This is worth sitting with, because it is the opposite of what the category's marketing says. "Gamification increases retention by X%" is a sentence you can find on dozens of vendor pages. We could not source a single such figure, and we are not going to invent one.
Here is what the research does measure, and where it was measured.
Learning and motivation, in education. Sailer and Homner's 2020 meta-analysis in Educational Psychology Review reports "significant small effects … on cognitive (g = .49 …, k = 19, N = 1686), motivational (g = .36 …, k = 16, N = 2246), and behavioral learning outcomes (g = .25 …, k = 9, N = 951)". Two caveats travel with that: it is education, and the authors themselves note the motivational and behavioural effects "were less stable" under a stricter subsplit. It is not a licence to say "gamification works".
Psychological need satisfaction, also in education. Li, Hew and Du's 2024 meta-analysis in ETR&D reports an overall g of 0.257 [0.043, 0.471], with autonomy at 0.638 [0.139, 1.136] and relatedness at 1.776 [0.737, 2.814]. That relatedness interval is enormous — the number without the interval would mislead you, which is why both are printed. Competence came in at 0.277 [0.001, 0.553], a lower bound of essentially zero.
Output volume, in the lab. Several online experiments measure how much work people produce. They generally find more of it, and no change in how people feel.
Intrinsic motivation, in the lab. Deci, Koestner and Ryan's 1999 meta-analysis in Psychological Bulletin is the foundation of every "extrinsic rewards backfire" argument. More on its actual numbers below, because they are routinely misquoted.
Retention: two studies, one platform, opposite populations. Restivo and van de Rijt measured it twice on Wikipedia. Their 2012 randomised trial reports that barnstar recipients "exhibited greater sustained productivity and were less likely to discontinue contributing" — but the sample was the top 1% of editors. Their 2014 follow-up, across editors of varying productivity, reports that rewards "were associated with lower retention of less-active contributors". The 2014 paper is paywalled, so we print its direction and nothing else. Two findings, one platform, one pair of authors: that is the whole retention evidence base for this article, and it says recognition holds your best people and may cost you everyone else.
The generalisability caveat, stated plainly. Most of this research is education and laboratory psychology. A handful of studies sit in anything resembling an online community — the Wikipedia experiments, a crowdsourcing experiment, a Stack Overflow analysis, a Duolingo forum study — and one of those points the opposite way to the classroom work. Sailer and Homner even found setting moderates results inside the education literature, for cognitive outcomes: "The effects found in school settings were significantly larger than those found in either higher education settings or informal education settings." That moderator was not significant for the motivational or behavioural outcomes, so do not stretch it further than they did — but if a school-versus-university gap shows up at all, assuming a school-versus-Discord transfer holds is unsupported.
Read every number below with that attached.
Points and XP: evidence for volume, not for motivation
Points and XP have decent evidence for one thing: they increase how much people do. In an online image-annotation experiment, Mekler and colleagues found points, levels and a leaderboard produced significantly more tags, and described the elements as effective only for promoting performance quantity. Motivation itself did not move.
Their wording is worth quoting exactly: "points, and especially levels and leaderboard led to a significantly higher amount of tags generated … effective only for promoting performance quantity." We reached that finding through the publisher's abstract record rather than the full text, so no sample size or effect size appears here. Treat the group's 2013 and 2017 studies as one result from one lab on one kind of task.
The honest summary is "quantity, not quality, not motivation". That is still useful. If what you need is volume — more submissions, more translations, more test reports — points are a defensible tool with real support behind them. If what you need is people who care more, points are not the evidence-backed lever and nobody should sell them to you as one.
There is a second finding that constrains how you use them. Deci, Koestner and Ryan found that "engagement-contingent, completion-contingent, and performance-contingent rewards significantly undermined free-choice intrinsic motivation (d = -0.40, -0.36, and -0.28)". That −0.40 is the number everybody quotes, and quoting it alone is the classic misuse: it is one subgroup. The overall free-choice figure across all rewards in the same paper is "composite effect size was d = -0.24", and on self-reported interest it is "d = 0.04 … rewards did not significantly affect self-reported interest". Even that comes with the scope caveat the authors state themselves — their studies were "all done either in the laboratory or under well-controlled, laboratory-like conditions", with field studies excluded. It is a lab literature about short tasks, not a finding about your community.
Levels: the widest gap between adoption and evidence
Levels are the mechanic with the widest gap between how much software ships them and how much evidence supports them. Mekler and colleagues' image-annotation experiment credits levels, as it does points, with raising output volume. Beyond that, we found no study isolating levels, and none measuring whether reaching one keeps anybody around.
Nearly every platform in this article has levels. MEE6 has a Levels plugin, Tatsu levels members across every server it runs in, Carl-bot sells levels to its Patrons, Layer3 says "XP drives your Level, your spot on the Leaderboard", and Zealy derives a level from XP on every read. The mechanic is close to universal and the evidence for it is one lab experiment about tagging images.
That does not make levels theatre. It makes them an assumption. The interesting question is what a level does when you reach it, and for most implementations the honest answer is "changes a number next to your name". Where a level unlocks a capability — a permission, an area, a right — you are no longer testing levels, you are testing whether the capability was worth wanting. Discourse's trust levels are the strongest example in this article, and they are covered below.
Leaderboards: contested, and the direction depends on the setting
Leaderboards are contested, and the direction of the effect depends on the setting. A classroom study found students at the bottom of an absolute leaderboard reported less motivation than those at the top — but no difference in engagement or learning. A crowdsourcing experiment found low ranks intensified effort instead.
The classroom study is Bai, Hew, Sailer and Jia, 2021, in Computers & Education. Two quasi-experiments, n = 24 and n = 26, at one institution, with overwhelmingly female postgraduate participants. On an absolute leaderboard, the difference in motivation between the top and bottom thirds at post-intervention was significant (U = 11, p = .028). On a relative leaderboard showing only five neighbouring positions, there were "no significant variations among the students at the three position levels".
Now the part that gets dropped when this study is cited, and it cuts both ways. On the absolute leaderboard, the motivation gap came with no difference in course engagement (F(2,17) = 0.55, p = .59) — so "people at the bottom report enjoying it less" is supported, while "people at the bottom disengage" is not, and disengagement is the claim everyone actually repeats. But on the relative leaderboard, where motivation and engagement were equal across positions, the bottom third still scored significantly worse on the post-test (M = 68.75 against the top third's 91.07, F(2,23) = 8.71, p = .002) — which the authors report as supporting their hypothesis that low-ranked learners on a relative leaderboard perform worse. Read together: hiding the exact ranking protected how people felt and did not protect what they achieved.
Against it sits Na and Han's 2023 crowdsourcing experiment in Internet Research, which manipulated ranks across five rounds and found that "High ranks … induced complacent behaviors … while low ranks led the participants to stick to the right process of the task with intensified motivation round after round." We reached this through the publisher's abstract, so no sample size appears. The authors add a line that is usually dropped with it: "neither of the motivations seemed to be of intrinsic nature." And in fairness, it is another image-tagging task — the same objection we raised against the points studies above applies here. It is the study in this set that sits closest to an online community, and it points the other way.
Two practical consequences, and the first comes with a caveat attached. If you are worried about the bottom of your leaderboard, the available mitigations are showing relative position rather than absolute rank, or letting the community owner switch the leaderboard to private, which is what Circle actually offers. Relative ranking is the better-evidenced of the two for how people feel — but in the same study it did not stop the bottom third underperforming, so treat it as protecting morale rather than outcomes. Circle is the only platform we checked that ships the private-leaderboard switch. Second, a leaderboard scoped to a short window rather than to all time changes who is at the bottom, because it resets the accumulated advantage of whoever arrived first.
Badges and recognition: the best community evidence, and the most uncomfortable
Badges have the best community evidence of any mechanic here, and the most uncomfortable. A randomised experiment on Wikipedia found awarding a barnstar increased productivity by 60% among highly active editors. The same authors' follow-up found the rewards were associated with lower retention among less-active contributors.
Restivo and van de Rijt's 2012 PLoS ONE study is a proper randomised controlled trial: a random sample of 200 drawn from the top 1% of 144,120 editors, 100 treated and 100 control, measured over 90 days, reporting that "receiving a barnstar increases productivity by 60%". Note the population, and note that this paper does report on retention as well: recipients "were less likely to discontinue contributing". This is evidence that recognising your most active contributors makes them more active. It is not evidence about anybody else, and the headline outcome is edits, and the retention finding above sits alongside it.
Their 2014 follow-up extended the design across productivity levels and found "rewards yielded increases in work only among already highly-productive editors … rewards were associated with lower retention of less-active contributors". Paywalled, so direction only. It is simultaneously the best evidence in this article that recognition can demotivate the periphery, and the retention finding that points the other way.
Anderson, Huttenlocher, Kleinberg and Leskovec's 2013 WWW paper on Stack Overflow adds the mechanism: "their badges steer behavior in ways closely consistent with the predictions of our model" — users accelerate as they approach a threshold, then relax after crossing it. That is observational, and steering is not retaining. Do not let anyone convert it.
One more thing worth knowing before you reach for badges as the "safe" alternative to points. Deci, Koestner and Ryan's coding scheme is explicit that "Symbolic rewards such as 'good-player certificates' were considered tangible". The vendor line that badges escape the extrinsic-reward problem because they are not money does not survive contact with the paper that defined the problem.
The distinction that does survive is between a reward and a compliment. The same meta-analysis found positive feedback went the other way: "The composite d = 0.33 … indicates that positive feedback did enhance intrinsic motivation", rising to 0.43 for college students. Recognition from a person is not the same stimulus as an automatically-granted icon, and the evidence treats them differently. If you take one design decision from this article, take that one.
Streaks: the thinnest evidence of any mechanic here
Streaks are the thinnest-evidenced mechanic in this whole review. We found no effect size for them in any source we could reach. What exists is qualitative: interviews with recreational runners describing sadness, anger and disappointment when a streak broke, and a Duolingo study describing gamification misuse, where users fixate on the mechanic and get distracted from learning.
Ingalls and colleagues' 2026 PLoS One study interviewed recreational runners about broken run streaks and reported "feelings of sadness, anger, disappointment and relief", described a "grieving process", and noted that "some runners felt streak-related inconveniences and ran with injury to prolong the streak". They also reported that "No long-term negative consequences were reported." That is running, not software, and it is a mechanism illustration rather than a statistic — there are no effect sizes to quote.
Hadi Mogavi and colleagues' 2022 ACM Learning at Scale paper looked at Duolingo forum discussion and interviews, and named the failure mode: "Gamification misuse … occurs when users become too fixated on gamification and get distracted from learning", producing "competitiveness, overindulgence in playfulness, and herding". Qualitative, and read from the arXiv preprint.
So the streak is the mechanic with the strongest folk reputation and the weakest paper trail. It clearly does something to people — the running interviews are vivid about that. What nobody has shown us is that the something is good for the community running it, or that it survives the day the member misses.
Variable rewards, raffles and loot boxes: the weakest transfer in this article
Random rewards are where the evidence gets weakest and the ethics get loudest. The loot-box literature links spending to problem-gambling severity, but its own authors say the direction of causality is unclear. Loot boxes involve real money; a free raffle over quest XP does not, so the transfer is weak.
Zendle and Cairns' 2018 PLoS ONE study reports "n = 7,422 … link (η² = 0.054) between the amount that gamers spent on loot boxes and the severity of their problem gambling". Their own caveat is the important part: "It is unclear … whether buying loot boxes acts as a gateway to problem gambling, or whether spending large amounts … appeals more to problem gamblers." Nobody should write "loot boxes cause gambling problems", including us. The study is cross-sectional and self-selected, and the authors call η² = 0.054 "small-to-medium" and say plainly "This is not a weak or unimportant relationship", so we are not going to file it under negligible either.
We are flagging this as the weakest transfer in the article deliberately. A loot box is bought with a credit card. A raffle over XP that members earned for free is a different stimulus with a different risk profile, and treating the two as equivalent would be sloppy in our favour as well as against it.
There is one finding in the reward literature that genuinely supports surprise over announcement. Deci, Koestner and Ryan report that for unexpected tangible rewards, "average d = 0.01 … unexpected tangible rewards do not affect free-choice behavior". Announced, contingent bounties undermine free-choice motivation in that literature; unexpected ones do not. The k there is small, so this is a mechanism argument rather than a robust result — but it is the best available reason to hand out a surprise reward rather than post a bounty board.
Arkada lists the most chance-shaped features of any platform we checked — a "Wheel of Fortune" and "Pyramids" appear in its core-features documentation alongside its questing product. We could not find a live page for either, so we cannot tell you how they work, only that Arkada advertises them. If you are evaluating Arkada, evaluate those two specifically and ask to see them running.
The two things every gamification article says that did not survive checking
Two claims appear in nearly every gamification article as settled fact, and neither survived checking. "Gamification fades from novelty" runs against the strongest meta-analysis we read, whose duration moderator found larger motivational effects in longer interventions. "Leaderboards demotivate the bottom" rests on small classroom samples and is contradicted, on effort, by the study closest to a community.
Claim one: it fades because it is new. Sailer and Homner's duration moderator ran the other way. Interventions lasting up to half a year showed larger motivational effects than one-day interventions (Q(1) = 4.93, p < .05), and the authors write that their results "can weaken the fear that effects of gamification might not persist in the long run". Their full text mentions "novelty" exactly once, paraphrasing someone else's review.
Be fair to the other side. That moderator is between-study, not within-person — it compares different studies of different lengths, not the same people over time — and the longest bucket is sparse. And there is a survey pointing the other way: Koivisto and Hamari's 2014 study of Fitocracy users in Computers in Human Behavior found that "perceived enjoyment and usefulness of the gamification decline with use, suggesting that users might experience novelty effects". That one is cross-sectional, which means it cannot establish decay within a person either, and it is fitness rather than community.
So the correct statement is: the novelty-decay claim is contested, the better-powered evidence points against it, and anyone asserting it as fact has not read the moderator. In fairness we should also say that Hamari — whose 2014 review we lean on further down for its bluntest sentence — is on the other side of this one, and Sailer and Homner name that review as among the work their duration finding contradicts.
Claim two: the leaderboard crushes the bottom. Covered above. Supported for self-reported motivation, at small classroom samples, post-intervention only, and with no effect on engagement — but on the relative leaderboard in the same paper the bottom third did score worse. Contradicted, on effort, by a crowdsourcing experiment where low ranks intensified effort round after round. "Contested" is the honest word; "settled" is not.
Worth adding a third, quieter correction. Hamari, Koivisto and Sarsa's 2014 HICSS review of the empirical literature is blunter than the field's marketing: "Only two studies found all of the tests positive", and the largest studies "reported that gamification might not be effective in a utilitarian service setting". That last phrase is the most transferable sentence in this whole evidence base, because a community platform is a utilitarian service setting.
Which platforms implement which mechanics, and what they cost
Four of the thirteen publish a price: Guild.xyz, Discourse, Circle and Zealy. MEE6 has one, but only through an API its own page calls at runtime. Seven publish none, and QuestN we could not check. The mechanics differ more than the marketing suggests: Discourse ties badges to permissions, Circle ships a privacy toggle, Layer3 verifies repeated on-chain actions.
Every mechanic below was matched against strings in pages we fetched on 17 August 2026. Every price is quoted as the page printed it. "No public price" states which pages we checked, because an unstated negative is worthless.
| Platform | Mechanics we could verify | Published price, as printed |
|---|---|---|
| Guild.xyz | Points and leaderboard exports, hidden and secret roles, token-gating, role ranking, sybil resistance | Starter $29/mo, Plus $99/mo, Growth $399/mo |
| Discourse | Badges granted automatically or manually, badge groups including Trust Level, trust levels from Basic (Trust Level 1) to Leader (Trust Level 4); admins can disable the entire badge system | Free; Pro $100/mo; Business $500/mo; Enterprise custom |
| Circle | Points, levels and rewards; leaderboards that spotlight top contributors; one point per like received; the leaderboard can be made private | Professional $89/mo and Business $199/mo, which the page's footnote says are "based on annual billing"; its embedded pricing data carries the month-to-month figures, $99 and $219. Circle Plus custom; add-ons including extra admins at $10/mo per admin |
| Zealy | XP from approved quest claims and manual admin grants, a globally fixed level curve, all-time and per-sprint leaderboards, five reward methods including raffle and vote | Free $0 with 1,000 quest claims; Standard $149/mo and Plus $359/mo at the annual rate, excluding VAT. Standard has no monthly option — six months is its shortest commitment. Read from Zealy's own pricing source on 17 August 2026 |
| Galxe | Galxe Quest, rewards from tokens to loyalty points, credential checks via API, REST, GraphQL and contract query, a Galxe Score reputation score | No public price (checked galxe.com, /pricing → 404, /quest, docs.galxe.com, docs …/pricing → 404) |
| Layer3 | Streaks defined as a repeating on-chain action on a given dApp, XP driving Level and leaderboard position, Leagues, Achievements, CUBEs | No public price (checked layer3.xyz, /business → redirects to homepage, docs plus 3 platform pages) |
| MEE6 | XP and Levels listed as premium in its own help-centre matrix, Achievements plugin, Levels plugin, Economy plugin | Priced, but not on the page. mee6.xyz and /premium serve a byte-identical SPA shell, and the price is fetched at runtime from mee6.xyz/api/shop/premium, which returned $11.99/mo (annual $49.99, lifetime $89.99) when we called it. That endpoint is geo-priced and A/B-tested, so treat any single figure as one sample |
| Tatsu | XP earned by chatting, levels across all of Discord, leveled roles granted automatically, a currency buying profile cards, badges, pets and furniture, rankings that persist across servers | No public price readable (/premium, /store and /docs all 404, and the support subdomain redirects off-site to a community wiki) |
| Carl-bot | Levels as a Patron-only feature, a default rate of 15–25 XP randomised per message, an XP curve of 5*(n^2)+50*n+100, a MEE6 import command | No public price on-domain (carl.gg and /premium are byte-identical shells); premium sold via Patreon |
| Arkada | Questing platform, Wheel of Fortune, Performance Hub, Pyramids, leaderboard, referral system and level progression, all as named in its core-features documentation — we found no live page for the Wheel or the Pyramids | No public price (checked arkada.gg, docs.arkada.gg and its core-features page) |
| Influitive | Gamified campaigns, completed challenges | No public price — its pricing page contains zero dollar figures; every call to action is a demo request. Its footer read "© 2024 Influitive" when we fetched it on 17 August 2026 |
| Bunchball / Nitro | Enterprise engagement software, per BI Worldwide, which now hosts the product | No public price. Its own domain is dead: bunchball.com fails TLS then 404s, and nitro.bunchball.com does not resolve |
| QuestN | We could not check it. questn.com and app.questn.com return 403 behind Cloudflare, and its docs and blog subdomains do not resolve in DNS. Its robots.txt is reachable and allows general crawlers, but a Cloudflare-managed block inside it names ours specifically, which we respected. The 403s were the real obstacle. We describe no QuestN features because we saw none | Could not verify |
Two things that table changes.
Discourse's trust levels are the strongest implementation of gamification in this article, and Zealy has nothing like them. A Discourse badge can carry a trust level, and a trust level carries actual moderation capability — the ability to do more in the community, not a picture next to your name. That is the difference between recognition that grants competence and recognition that decorates it, and it is precisely the distinction the research keeps pointing at. If you are running a forum and you want status that means something, Discourse is doing this better than we are.
Circle ships the mitigation the leaderboard evidence implies, and we do not. Circle's own gamification page says you can choose to make your community leaderboard private. Given that the best-supported leaderboard finding is about how people at the bottom feel, an off switch controlled by the community owner is a real design answer. Zealy has no equivalent toggle.
One aside from the fetching, because it says something about the category: Layer3's homepage no longer uses the word "quest" in its visible copy — its vocabulary is Activations, Streaks and CUBEs. The word survives inside the shipped JavaScript, in route names and a QUEST_COMPLETE string, which is arguably better evidence for the point: the rename is in the marketing, and the plumbing still remembers. The category leader dropped the category word. We compare its capabilities against ours in Zealy vs Layer3, and if Galxe is the incumbent you are replacing, best Galxe alternatives and Zealy vs Galxe go through that decision properly.
MEE6 and Tatsu belong to a different shape of problem — chat-activity XP inside a single Discord server — and we wrote about that set separately in Discord engagement and growth tools.
What Zealy's gamification actually does, and the five things it does not have
Zealy awards XP for approved quest claims and manual admin grants, derives a level from a single global formula, and ranks members on two leaderboard scopes: all-time and per-sprint. It has no badges, no streaks, no daily-login bonus, no XP multipliers and no role-by-level automation, and it computes no retention metric.
All of the following was read out of the source code on 17 August 2026. Where our documentation disagrees with the code, the code is what actually runs.
The level curve is one formula and you cannot change it
Cumulative XP for level L is 150(L−1)² − 50(L−1). Level 1 is 0 XP, level 2 is 100, level 3 is 500, level 4 is 1,200, level 5 is 2,200. That curve is identical in every Zealy community and there is no per-community configuration for it. Level is not stored in the generated database models at all; it is derived from XP on every read. Our level calculation documentation matches the code here, which we checked rather than assumed.
Compare that with Carl-bot, whose operators can blacklist channels and roles, set voice XP and rate-limit the whole economy. Its documented curve is 5*(n^2)+50*n+100, which is the XP needed for the next level rather than a running total, so do not read it against Zealy's cumulative form — the point is the dial, not the shape. Zealy gives you no such dial.
Nothing happens when a member levels up
This is the finding we expected to be wrong and was not. Zealy has two notification enums — a narrow one in its database models with three types, and the one its notification API actually returns, with thirteen — plus a webhook event enum with ten. None of the three contains a level event. No notification fires. No webhook fires. No celebration appears. A helper called hasLeveledUp exists in the codebase, is commented out of the export barrel, and is called nowhere.
Level's only mechanical consequence in Zealy is that it can be used as a quest condition — a gate on who can attempt something. That is the whole of it.
Given the section above on levels, this is less damning than it sounds and more honest than the alternative. Levels have almost no evidence behind them, and a level-up celebration is exactly the kind of thing that gets shipped because everyone ships it. But if you are choosing Zealy because you pictured members getting a satisfying ping at level 10, picture it accurately: they get a bigger number.
Two leaderboard scopes, and no weekly or monthly one
Zealy has exactly two leaderboard scopes: all-time, and per-sprint. There is no weekly, monthly or daily leaderboard. If you want a shorter window — and the leaderboard evidence above is a decent argument for one, since a fresh window resets who is at the bottom — you create a sprint, which is a time-boxed second leaderboard scoped to a chosen set of quests.
Two details that get misdescribed, including sometimes by us. A leaderboard reset is a date cutoff plus a soft delete, not an erasure — quest claim history is untouched and members' XP is zeroed going forward. And ties break in favour of whoever reached the total first, except that a manual XP grant strips that tie-break until the member's next quest claim.
Sprint rewards are proportional, not a podium and not a raffle
The default sprint payout is proportional to each member's share of XP among everyone inside the reward zone. It is not top-three, and it is not a draw. The reward zone is an eligibility cutoff applied before distribution, not a payout curve. The alternative setting splits a rank range evenly within each tier. Both paths round down — the proportional one once per currency, the tiered one twice — so a small remainder goes undistributed either way. On the tiered path that is deliberate, and documented: the function's own comment reads "Remainder from rounding stays undistributed."
Sprint prize pools are USDC and zaps only. A sprint cannot award XP, roles, NFTs or tokens; those live on quest rewards, where there are five reward methods, one of which is a raffle. On reward types the count depends on where you look: Zealy's API advertises seven, but the quest editor excludes NFT rewards, so an admin can configure six — which is why our docs page lists six and not seven. Zealy draws raffles automatically in its own database, and the draw selects claims rather than distinct members, so somebody with several successful claims on a recurring quest holds proportionally more entries. That is worth knowing before you design one.
The five things Zealy does not have
Each of these is a claim about tracked source in Zealy's repository, searched on 17 August 2026 across the API, backend, schemas, models, utils, queries, contracts, the web app and the database migrations.
- No badges and no achievements a member can earn. Be careful how you check this, because we got it wrong first time. Searching
badgeacross the codebase returns well over a hundred files, and the biggest group is aBadgecomponent in Zealy's design system — a small label used for statuses and counts throughout the interface, not something anyone earns.medallikewise hits agetMedalhelper that draws first, second and third place on a leaderboard, which is a rendering of rank rather than an award. What does not exist anywhere is a badge or achievement entity: nothing a member is granted, holds, or can be shown a collection of. Discourse and Tatsu both have real badge systems. We have a UI component with the same name. - No streaks.
streak,loginStreak,visitStreakandday streakreturn nothing in tracked source. Layer3 has Streaks and Duolingo built a category around them. Given how thin the streak evidence is, we are not in a hurry, but the absence is the absence. - No daily-login bonus.
dailyLogin,loginBonus,dailyBonus,login reward,lastLoginAt,lastSeenAt,dailyClaim,daysActiveandactiveDaysall return zero. Zealy does run a platform-level Daily Challenge — complete a quest in each of three featured communities, pass humanity verification, claim a lottery ticket — but that is a daily reset with no consecutive-day bonus, and it belongs to Zealy rather than to your community. Recurring quests reset on calendar boundaries too, which is recurrence, not a streak. - No XP multipliers or boosts.
xpMultiplier,boostXp,doubleXpandxpBoostall return zero. Zealy does have aboostfeature, and it is paid promotion of a quest or sprint, not a reward multiplier. Do not plan a double-XP weekend. - No role-by-level or role-by-XP automation.
autoRole,roleByLevel,levelRole,xpThreshold,autoAssignRole,rankRole,tierRoleandroleSyncall return zero. A Discord role in Zealy is a quest reward, granted once when a claim is approved. Tatsu gives leveled roles automatically; Zealy does not.
And the one that matters most for this article's subject:
Zealy computes no retention metric at all. retention, churn, D7, D30, DAU, WAU, MAU, returnRate, repeatVisit, stickiness, dropOff, dailyActive and monthlyActive all return zero in tracked source. Structurally, the analytics interface declares fifteen methods — two running totals that take no date range, one windowed count by claim status, five ranked lists, six series bucketed by interval, and one member count as of a single date. None of them joins a member's earlier activity to their later activity, which is the operation retention requires.
So when a gamification vendor tells you their mechanics improve retention, ask what they measure. We cannot measure it, and we are telling you that in our own article about gamification.
All the usual scope applies: these are statements about the searches listed, in that codebase, on that date. They are not claims about every feature Zealy has ever shipped or will ship.
If you want the general picture of what Zealy is before evaluating the mechanics, what is Zealy covers it.
How to choose a gamification platform for a community
Choose by the behaviour you need, not the mechanic you like. Discourse if earned status should grant real permissions. Circle if leaderboard visibility must be optional. Guild.xyz if the job is token-gated access. A quest platform like Zealy if you need verified off-platform actions rewarded on a schedule.
A few decisions the evidence actually supports:
If you want more output from your most active people, recognise them specifically. That is the Wikipedia barnstar result, and it is the strongest community finding in this article. It also comes with its own warning: the follow-up found the same rewards were associated with lower retention among less-active contributors. Recognise widely and specifically, not narrowly and competitively.
If you want a person to feel it, have a person do it. Positive feedback enhanced intrinsic motivation in the same meta-analysis where tangible rewards undermined it, and a certificate counted as tangible. An automated badge and a moderator saying "this was the best answer in here this week" are not the same product feature, and the cheaper one is the better-evidenced one.
If you run a leaderboard, think about the bottom half before you launch it. Show relative position rather than absolute rank, or let the owner switch the leaderboard to private, or scope it to a short window that resets. Circle's privacy toggle is the cleanest version of this we found; Zealy's sprints are the short-window version.
If you want volume, points are fine and you should not expect more from them. They raise quantity. They do not appear to change how people feel about the work.
Do not buy a platform on a retention promise. Nobody in the literature we could reach has measured it. A vendor quoting a retention percentage either has private data they should describe, or they read it off another vendor's blog.
Two adjacent decisions this post deliberately does not make for you. If the mechanics are for an ambassador programme, the design constraints are different and how to run a crypto ambassador program covers them. If your gamification will generate a review queue — and any task requiring a human to read a submission will — how to delegate community tasks covers the staffing cost that never appears on a pricing page. And if you are choosing gamification as part of a wider go-to-market, our web3 marketing guide puts it next to the other channels and their costs.
The most useful thing we can leave you with is the shape of the uncertainty. Six mechanics, one real community RCT between them, one paywalled retention direction, and a literature that is mostly about students. Anyone selling you certainty here has not looked.
