Postlia Slop Reports
Methodology
This page is the pre-registration for every dataset we publish under /reports. It goes live beforewe collect anything, so the rules can't quietly change after we've seen the results. If a future report deviates from what's written here, the deviation is a bug and we'll say so.
This detector measures style, not authorship.
A high score means the writing carries patterns that make text read as generic or templated — patterns that large language models produce heavily, and that humans produce too. It is not a claim that any specific post was written by an AI, and it cannot be used as one.
The detector
Scoring is done by the same engine that powers the free Slop Detector in the product: a deterministic, rule-based scorer — deliberately not another LLM. Every rule is a fixed pattern (a lexicon match or a structural check), so any run over the same text produces the same score, and every finding can be pointed at: this phrase tripped that rule.
Each finding carries a severity penalty (6 / 10 / 16 points, with diminishing returns for repeat findings of the same rule). Texts shorter than 80 characters are dampened (×0.6) because short posts can't exhibit most patterns. Scores clamp to 0–100 and band at 25 (“a little polished”), 50 (“smells like AI”), and 75 (“full slop”).
The 14 rules, in full
Rendered directly from the production rule set — this table cannot drift from what the engine actually runs.
| Rule | Class | Severity |
|---|---|---|
“X, not Y” framing Contrast constructions like “built for creators, not for lock-in” are a hallmark polished-marketing pattern. Say the one thing you mean. | AI / template | Medium (10 pts) |
“No X, no Y, no Z” mantra Chained negations read like ad copy, not a person. Pick the single promise that matters and make it concrete. | AI / template | High (16 pts) |
Decorative three-part list Three perfectly parallel phrases (“the flows, the copy, the details”) are rhythm decoration. Cut to the one that carries the point. | AI / template | Medium (10 pts) |
Empty intensifier Words like “seamlessly”, “genuinely”, “supercharge” add hype, not information. Replace with a concrete fact or delete. | AI / template | Low (6 pts) |
Hype opener “Exciting news!” / “Thrilled to announce” is the most recognizable template opener on every feed. Start with the actual news. | AI / template | High (16 pts) |
Em-dash overuse Multiple em-dash asides per sentence is a telltale polished-prose rhythm. Break into plain sentences. | AI / template | Low (6 pts) |
Uniform sentence rhythm Every sentence the same length reads machine-smooth. Vary it: one short. Then a longer one that actually explains. | AI / template | Medium (10 pts) |
Emoji burst Clusters like 🚀✨💯 are feed-filler decoration. One emoji that means something beats three that don't. | Broadcast style | Medium (10 pts) |
Hashtag stuffing A wall of hashtags reads as broadcast, not conversation — and on this platform it hurts more than it helps. | Broadcast style | Medium (10 pts) |
Exclamation overload More than one exclamation point per few sentences reads as manufactured enthusiasm. | Broadcast style | Low (6 pts) |
Cliché call-to-action “Let me know in the comments” / “Who's with me?” is the same closer on a million posts. Ask the specific question you actually have. | Broadcast style | Medium (10 pts) |
Essay sign-off “In conclusion” / “At the end of the day” is essay filler. Posts don't need a formal close — end on the point. | AI / template | Low (6 pts) |
AI-favored word Words LLMs reach for far more than people do (delve, tapestry, testament, harness). One is fine; a cluster reads generated — swap for the plain word. | AI / template | Low (6 pts) |
Corporate cliché Stock business phrases (“move the needle”, “in today's fast-paced world”) signal filler, not thought. Say the specific thing. | AI / template | Low (6 pts) |
The two classes are reported separately in every dataset: AI / template rules capture vocabulary and structures that surged with LLM writing; broadcast stylerules (hashtag stuffing, emoji bursts, exclamation overload, cliché CTAs) predate LLMs and measure templated feed behavior generally. Headline findings use the phrase “AI-flavored and templated-writing tells” — never a bare claim of AI authorship. One detail worth stating up front: the em-dash rule only fires on multiple em-dash asides in a single sentence, not on any use of an em-dash.
Corpus: open networks only
Reports are built exclusively from networks whose public firehoses are open by design: Bluesky (the public Jetstream feed) and Mastodon(public timeline APIs). We do not collect from LinkedIn, Instagram, TikTok, or any platform whose terms forbid it — if we can't collect it cleanly and openly, we don't publish it. That constraint is the point: the whole pipeline is reproducible by anyone, with no scraping and no terms-of-service gymnastics.
Consent and privacy
Opt-outs are honored. Mastodon accounts that set discoverable=false are excluded from sampling, and where servers expose an indexable flag we honor that too. Every dataset publishes the count of posts excluded this way.
Raw text is deleted after scoring. We keep only aggregate counts — rule frequencies, score distributions, and phrase tallies. No post text, author handle, or per-account result is ever stored or published. The honest trade-off: this makes the pipeline reproducible but not the sample — nobody (including us) can re-audit the exact posts, only re-run the method on fresh data. We consider that the right side of the trade.
Sampling and filters
Collection windows are stated on each report (the standard run is 7 consecutive days, sampled around the clock). Posts are deduplicated, obvious automation floods are dropped, and only English-language posts are scored: platform language tags are checked against a script-based heuristic, since client defaults over-report English. The residual bias runs one direction and we state it — non-English text mislabelled as English trips almost nothing against an English lexicon, which deflates scores, never inflates them.
Length is analyzed in three buckets — 40–79, 80–300, and 300+ characters — with the boundary at 80 chosen deliberately: it is where the engine's short-text dampening ends, so no bucket straddles a scoring discontinuity. Posts under 40 characters are excluded entirely.
Collector uptime is logged, and any gap in the collection window is published as a gap log with each dataset. Missing hours are a disclosed footnote, never a discovered flaw.
Metrics, pre-committed
The primary metric is trip rate: the share of posts with at least one finding. The secondary metric is slop rate: the share scoring 50+. We commit to this ordering now because the scoring math is strict — several distinct rules must fire in a short text to cross 50 — so slop rates on short-form networks may be very low. A near-zero slop rate is the design working, and it will be reported, not buried.
Cross-platform comparisons are length-normalized: platforms are compared within matched length buckets, never on raw averages, because post-length distributions differ mechanically between networks. The only platform-conditioned rule is the hashtag threshold, and it is identical (more than 3) for both Bluesky and Mastodon. Hour-of-day and day-of-week charts plot trip rate only — band-level cells are too thin to publish honestly.
Where Mastodon's public application field is present, we additionally report scheduler-posted vs. organic-client posts as separate populations. Yes — we are a scheduler company publishing how scheduled posts score. That cut exists because we want the answer too.
The pre-2022 baseline
“Overused” is a comparison, not a vibe. Phrase and rule frequencies are reported against a baseline sample of public Mastodon posts from before ChatGPT's release (Mastodon's public archives reach back far enough; Bluesky, which opened in 2023, has no pre-LLM baseline — a disclosed limitation). Every headline stat is a now vs. then number computed with the same rules and filters on both samples.
Versioning
Every dataset is stamped with the detector version and collection window (e.g. detector v2 · dataset 2026-08). Rule changes bump the detector version; numbers across different detector versions are never mixed in one chart. This page carries a changelog of any methodology amendment, dated.
Changelog: 2026-07-23 — initial methodology published, before any collection.
Questions, holes, corrections
We're two founders, not a research lab, and this methodology is meant to be poked at. If you find a hole in it, tell us and we'll fix it in public: support@postlia.com. Reports and their aggregate data are free to cite with a link.