The AI roleplay slop lexicon: 95 phrases, with receipts

A maintained, sourced reference of the 95 phrases that mark AI-generated roleplay prose: shivers down spines, whispers barely above themselves, mischief-sparkling eyes, and the "not X, but Y" construction. Cross-checked from Sukino's Banned Tokens, the Antislop paper (arXiv 2510.15061 — "Elara" runs 85,513x over the human baseline), and EQ-Bench's slop score. Plus two numbers most lists miss: top human-written cards average slop hits too, and flagship models' single-shot output now passes word-list checks clean.

A pinboard of paper strips, most crossed out in red ink, a hand with a red pen marking another

"Are you tired of ministrations sending shivers down your spine? Do you swallow hard every time their eyes sparkle with mischief and they murmur to you barely above a whisper?" That's the opening of Sukino's Banned Tokens list, the most-circulated slop lexicon in the roleplay community, and it manages to commit five entries from its own list in two sentences. Everyone who chats with AI characters knows these phrases by feel. This page is the feel, written down, with sources.

We compiled it for our own card-authoring pipeline (the scorer that enforces it is open source), cross-checking three independent kinds of evidence so no single taste dictates the list: community ban lists (Sukino's), statistical over-representation measured against human baselines (the Antislop paper, ICLR 2026 submission), and benchmark component weights (EQ-Bench's slop score: 60% word frequency, 25% contrast constructions, 15% slop trigrams). 95 phrases total, ten groups. Data note: Foreverse research, 2026-07.

A scale-setter before the tables, because the numbers are the argument: Antislop measured the character name "Elara" at 85,513x the human baseline in one model's creative output, the trigram "heart hammered ribs" at 1,192x, and one mid-size model producing "eyes never leaving" 102 times across 96 test prompts. These aren't vibes. They're measured fixations.

A. Body reactions

The workhorse group. Twelve printed here; three more from this family are NSFW-register and live with group I below.

PhraseNote
shivers down her/his spinewhole family: shiver down/up, sending shivers, sent a shiver
swallowed hard / swallows hard
breath hitchescommunity-reported high-frequency
heart hammered against his ribs1,192x over human baseline (Antislop)
knuckles turning white / whitening
adam's apple bobbing
felt a twinge
felt a chill run
stomach does a flip
takes a deep breath / took a deep breath
felt a strange sense
arched spine

B. Voice

PhraseNote
barely above a whisperthe community's favorite specimen
barely a whisper / voice barely audible
husky voice / husky whispers
voice a low purr / purred
chuckles darkly
seductive purrs
murmurednot banned outright — flagged as a dependency when it carries every line
voice thick with (emotion)
voice steady despite
though her voice lacks any real bite
whispering words of passion
grins wickedly

C. Eyes and expressions

PhraseNote
eyes sparkling with mischiefwhole family: gleam, glint, glow, shine, sparkle, twinkle
eyes never leaving102 occurrences in 96 prompts on one model (Antislop)
half-lidded eyes
a smile that did not reach her eyes
knowing smile
smirk playing on her lips
playfully smirking
crinkle at the corner of his eyes
long lashes
calculating gaze / assessing looknewer-generation tell, reported by RP users of 2025-26 models

D. Touch and motion

PhraseNote
ministrationsthe word that made the genre notorious
tracing a finger / tracing a nail
practiced ease
pushing aside a strand of hair / tucking a strand
fidget with the hem ofall variants
grips like a vice
waggles her eyebrows
towers over

E. Sentence constructions

The highest-value group, because constructions survive vocabulary swaps. If you only police one thing, police the first row.

PatternNote
not X, but Y / It's not just X, it's Ythe #1 LLM rhetorical crutch (NousResearch); a standalone 25% of EQ-Bench's slop score; measured up to 6.3x over human baseline
a mix of X and Y / felt a mix ofthe emotion-cocktail formula
couldn't help but
testament to
despite himself / herself / themselves
torn between
maybe, just maybe
little did she/he/they knownarrator wink; zero tolerance in our house rules
for what felt like an eternity
unbeknownst to them
whether you like it or not
without waiting for a response

F. Stock nouns and similes

PhraseNote
tapestry (of) / rich tapestryshared with general AI-writing lists
symphony of
kaleidoscope
cacophony
siren call / siren's call
like a moth to a flame
like a predator stalking its preynewer-generation tell
a dance as old as time
soothing balm

G. Atmosphere

PhraseNote
dimly lit
casting long shadows
the air is thick with / the air crackles with tension
sun dipped below the horizon
dust motes dancing in the lightall variants
words hung in the air / hung heavy in the air
the atmosphere was charged

H. Endings and outlooks

The fishing-line family. These matter double in character cards because a greeting that ends on one trains the model to end every reply on one.

PhraseNote
they would face it together
was only just beginning
the night is still young
for now, that was enough
ready to face whatever lay ahead
renewed sense of purpose / newfound sense of
the ball is in your court / the choice is yours / what do you sayfishing lines: the reply ends by begging for the next one
. Or something elsethe trailing-option fish

I. The NSFW cluster

Ten phrases, plus the three body-family ones held back from group A. This is a general-audience page, so we won't print them; if you need them for filtering, they're the NSFW_PHRASES constant in the scorer source. On SFW cards our gate treats any hit from this group as an automatic fail rather than a deduction.

J. Naming slop

Models fixate on character names, and the fixations are the most statistically extreme entries on the whole list.

NameNote
Elara85,513x over human baseline — the most famous AI name (Antislop)
Kael
Seraphinadoubly contaminated: also SillyTavern's default example character
Lily / Sarah Chenmodel-family favorites; the specific names vary by model line

The fix isn't a longer ban list — models will fixate on new names next year. It's a naming procedure: anchor in a real language culture (Welsh, Polish, Nigerian, Irish surnames…) or coin something, then search the result to confirm you didn't collide with a franchise.

How to use a slop list without ruining your prose

By cluster density, not word bans. The anti-slop projects themselves warn about this: not every use of these words is wrong, but clusters of them are a giveaway. Humans wrote every phrase here first — that is precisely how the models learned them — so a zealous find-and-delete pass produces prose with a different problem: the sanded, evasive texture of text that's afraid of itself. Our house rules quota the constructions ("not X, but Y" at most once per card, emotion cocktails at zero, fishing endings at zero) and treat the vocabulary groups as rewrite triggers rather than deletions.

Also worth knowing: slop fingerprints cluster by model family. The Antislop data shows each model line has its own favorites, so a list tuned on one model's output will under-detect another's. That's the strongest argument for pattern-level rules (group E) over vocabulary-level ones — constructions transfer across models better than words do.

The two numbers most slop lists don't tell you

First: star counts don't filter slop. We scanned the greetings of the 30 top-starred SFW character cards on chub.ai (sampled 2026-07, human-written, community-validated) with a 13-pattern subset of this lexicon and got 20 hits across the sample — "you feel" in 5 cards, "a mixture of" in 4, plus scattered gleaming eyes and mischievous grins. The most popular human cards in the world average 0.07 lexicon hits per greeting. Slop phrases are not an AI marker; slop density is.

Second: flagship models now pass word-list checks in a single polished greeting. In our calibration set — 9 raw greetings from 5 current model families, generated with a plain prompt and zero anti-slop instructions — the average was 0.22 hits, with the two strongest models scoring zero across the board. The separation direction is correct (0.07 human vs 0.22 machine) but the honest reading is that word lists catch weak models and long multi-turn drift, not a frontier model on its best behavior. Where the slop actually resurfaces is turn forty of a chat, when the context fills with the model's own output and the fixations compound.

Method note

Sources: Sukino's Banned Tokens (community ban list, continuously updated), the Antislop paper (arXiv 2510.15061, statistical over-representation vs human baselines), EQ-Bench slop-score components, NousResearch's anti-slop writing guide, and RP community complaint threads for the 2025-26 additions (calculating gaze, predator similes). Human baseline scan: 30 top-starred SFW cards, first messages, July 2026. Machine baseline: 9 single-shot greetings across 5 model families, plain prompt, no style instructions. The scorer implementing all of it — including the red lines and the construction quotas — is score_card_en.py in character-card-skills, runnable with no dependencies. If a phrase on this list false-flags your human prose, file an issue; reader-flagged false positives are how the list stays calibrated. And if you came here wondering about the em dash: it's deliberately absent, because the data says it isn't a tell in English.

We maintain this lexicon because we run character-card authoring on top of it — the full writing guide shows where the quotas slot into the larger procedure, and the cards it produces are chattable in Foreverse or any SillyTavern-compatible frontend.

FAQ

What are slop words in AI roleplay?

Stock phrases that language models over-produce relative to human writers — "barely above a whisper," "eyes sparkling with mischief," "shivers down her spine," "ministrations." The over-production is measurable: the Antislop paper (arXiv 2510.15061) found the character name "Elara" appearing 85,513 times more often in one model's fiction output than in human writing, and the trigram "heart hammered ribs" at 1,192x. Individually the phrases are ordinary English; it's the density and clustering that reads as machine output.

Is one slop phrase proof that text is AI-generated?

No. Humans wrote every phrase on this list first — that's how the models learned them. In our scan of 30 top-starred, human-written SFW character cards, the phrases still turn up: "you feel a mixture of" appears in 4 of the 30, and the sample averages 0.07 lexicon hits per greeting. Slop judgment works on clusters: one "murmured" is prose, five murmurs and a knowing smile and a breath someone didn't know they were holding in the same scene is a fingerprint.

What is the most reliable AI tell in English roleplay prose?

Not a word — a construction. The "not X, but Y" contrast pattern is the most-cited single marker: EQ-Bench weights contrast patterns as a standalone 25% of its slop score, and NousResearch's anti-slop guide calls it the number-one LLM rhetorical crutch. Behind it: emotion-cocktail formulas ("a mix of anticipation and dread") and fishing-line endings ("the choice is yours"). And a negative result worth knowing: em-dash density does not separate human from AI in English RP prose — we measured that separately.

Do slop word lists work as AI detectors?

As lint, yes; as a lie detector, no. In our calibration, human top-card greetings averaged 0.07 hits while raw single-shot LLM greetings averaged 0.22 — the direction is right but the gap is small, because 2026 flagship models produce near-zero word-list hits in a single polished greeting. The list earns its keep in long multi-turn chats, on weaker models, and as a writing-discipline gate for card authors. Treat a clean scan as "not obviously sloppy," never as "human."

AI Roleplay Slop Words: The 95-Phrase Reference List (2026) · Foreverse · Xinmeng