Postlia Slop Reports

Methodology

This page is the pre-registration for every dataset we publish under /reports. It goes live beforewe collect anything, so the rules can't quietly change after we've seen the results. If a future report deviates from what's written here, the deviation is a bug and we'll say so.

This detector measures style, not authorship.

A high score means the writing carries patterns that make text read as generic or templated — patterns that large language models produce heavily, and that humans produce too. It is not a claim that any specific post was written by an AI, and it cannot be used as one.

The detector

Scoring is done by the same engine that powers the free Slop Detector in the product: a deterministic, rule-based scorer — deliberately not another LLM. Every rule is a fixed pattern (a lexicon match or a structural check), so any run over the same text produces the same score, and every finding can be pointed at: this phrase tripped that rule.

Each finding carries a severity penalty (6 / 10 / 16 points, with diminishing returns for repeat findings of the same rule). Texts shorter than 80 characters are dampened (×0.6) because short posts can't exhibit most patterns. Scores clamp to 0–100 and band at 25 (“a little polished”), 50 (“smells like AI”), and 75 (“full slop”).

The 14 rules, in full

Rendered directly from the production rule set — this table cannot drift from what the engine actually runs.

RuleClassSeverity

“X, not Y” framing

Contrast constructions like “built for creators, not for lock-in” are a hallmark polished-marketing pattern. Say the one thing you mean.

AI / templateMedium (10 pts)

“No X, no Y, no Z” mantra

Chained negations read like ad copy, not a person. Pick the single promise that matters and make it concrete.

AI / templateHigh (16 pts)

Decorative three-part list

Three perfectly parallel phrases (“the flows, the copy, the details”) are rhythm decoration. Cut to the one that carries the point.

AI / templateMedium (10 pts)

Empty intensifier

Words like “seamlessly”, “genuinely”, “supercharge” add hype, not information. Replace with a concrete fact or delete.

AI / templateLow (6 pts)

Hype opener

“Exciting news!” / “Thrilled to announce” is the most recognizable template opener on every feed. Start with the actual news.

AI / templateHigh (16 pts)

Em-dash overuse

Multiple em-dash asides per sentence is a telltale polished-prose rhythm. Break into plain sentences.

AI / templateLow (6 pts)

Uniform sentence rhythm

Every sentence the same length reads machine-smooth. Vary it: one short. Then a longer one that actually explains.

AI / templateMedium (10 pts)

Emoji burst

Clusters like 🚀✨💯 are feed-filler decoration. One emoji that means something beats three that don't.

Broadcast styleMedium (10 pts)

Hashtag stuffing

A wall of hashtags reads as broadcast, not conversation — and on this platform it hurts more than it helps.

Broadcast styleMedium (10 pts)

Exclamation overload

More than one exclamation point per few sentences reads as manufactured enthusiasm.

Broadcast styleLow (6 pts)

Cliché call-to-action

“Let me know in the comments” / “Who's with me?” is the same closer on a million posts. Ask the specific question you actually have.

Broadcast styleMedium (10 pts)

Essay sign-off

“In conclusion” / “At the end of the day” is essay filler. Posts don't need a formal close — end on the point.

AI / templateLow (6 pts)

AI-favored word

Words LLMs reach for far more than people do (delve, tapestry, testament, harness). One is fine; a cluster reads generated — swap for the plain word.

AI / templateLow (6 pts)

Corporate cliché

Stock business phrases (“move the needle”, “in today's fast-paced world”) signal filler, not thought. Say the specific thing.

AI / templateLow (6 pts)

The two classes are reported separately in every dataset: AI / template rules capture vocabulary and structures that surged with LLM writing; broadcast stylerules (hashtag stuffing, emoji bursts, exclamation overload, cliché CTAs) predate LLMs and measure templated feed behavior generally. Headline findings use the phrase “AI-flavored and templated-writing tells” — never a bare claim of AI authorship. One detail worth stating up front: the em-dash rule only fires on multiple em-dash asides in a single sentence, not on any use of an em-dash.

Corpus: open networks only

Reports are built exclusively from networks whose public firehoses are open by design: Bluesky (the public Jetstream feed) and Mastodon(public timeline APIs). We do not collect from LinkedIn, Instagram, TikTok, or any platform whose terms forbid it — if we can't collect it cleanly and openly, we don't publish it. That constraint is the point: the whole pipeline is reproducible by anyone, with no scraping and no terms-of-service gymnastics.

Consent and privacy

Opt-outs are honored. Mastodon accounts that set discoverable=false are excluded from sampling, and where servers expose an indexable flag we honor that too. Every dataset publishes the count of posts excluded this way.

Raw text is deleted after scoring. We keep only aggregate counts — rule frequencies, score distributions, and phrase tallies. No post text, author handle, or per-account result is ever stored or published. The honest trade-off: this makes the pipeline reproducible but not the sample — nobody (including us) can re-audit the exact posts, only re-run the method on fresh data. We consider that the right side of the trade.

Sampling and filters

Collection windows are stated on each report (the standard run is 7 consecutive days, sampled around the clock). Posts are deduplicated, obvious automation floods are dropped, and only English-language posts are scored: platform language tags are checked against a script-based heuristic, since client defaults over-report English. The residual bias runs one direction and we state it — non-English text mislabelled as English trips almost nothing against an English lexicon, which deflates scores, never inflates them.

Length is analyzed in three buckets — 40–79, 80–300, and 300+ characters — with the boundary at 80 chosen deliberately: it is where the engine's short-text dampening ends, so no bucket straddles a scoring discontinuity. Posts under 40 characters are excluded entirely.

Collector uptime is logged, and any gap in the collection window is published as a gap log with each dataset. Missing hours are a disclosed footnote, never a discovered flaw.

Metrics, pre-committed

The primary metric is trip rate: the share of posts with at least one finding. The secondary metric is slop rate: the share scoring 50+. We commit to this ordering now because the scoring math is strict — several distinct rules must fire in a short text to cross 50 — so slop rates on short-form networks may be very low. A near-zero slop rate is the design working, and it will be reported, not buried.

Cross-platform comparisons are length-normalized: platforms are compared within matched length buckets, never on raw averages, because post-length distributions differ mechanically between networks. The only platform-conditioned rule is the hashtag threshold, and it is identical (more than 3) for both Bluesky and Mastodon. Hour-of-day and day-of-week charts plot trip rate only — band-level cells are too thin to publish honestly.

Where Mastodon's public application field is present, we additionally report scheduler-posted vs. organic-client posts as separate populations. Yes — we are a scheduler company publishing how scheduled posts score. That cut exists because we want the answer too.

The pre-2022 baseline

“Overused” is a comparison, not a vibe. Phrase and rule frequencies are reported against a baseline sample of public Mastodon posts from before ChatGPT's release (Mastodon's public archives reach back far enough; Bluesky, which opened in 2023, has no pre-LLM baseline — a disclosed limitation). Every headline stat is a now vs. then number computed with the same rules and filters on both samples.

Versioning

Every dataset is stamped with the detector version and collection window (e.g. detector v2 · dataset 2026-08). Rule changes bump the detector version; numbers across different detector versions are never mixed in one chart. This page carries a changelog of any methodology amendment, dated.

Changelog: 2026-07-23 — initial methodology published, before any collection.

Questions, holes, corrections

We're two founders, not a research lab, and this methodology is meant to be poked at. If you find a hole in it, tell us and we'll fix it in public: support@postlia.com. Reports and their aggregate data are free to cite with a link.