Foreverse Research · Fiction Bench

How good is GLM-5.2 at writing fiction?

It's the only model of nine in the top three on both genres: 2-4 on xuanhuan, 2-3 on court romance, flat and drift-free throughout. Its one recurring demerit is oddly small — it writes dialogue in half-width quotes (108 pairs on one book, 90 on the other), which keeps breaking Chinese typographic immersion.

Xuanhuan fantasy · Rank 2-4 / 9Court romance · Rank 2-3 / 9None observed$0.34 / 10k

Xuanhuan fantasy

2-4 / 9

Both packs: p303 第2 · p404 第3

Flat and drift-free end to end; half-width quotation marks were its only recurring demerit.

Court romance

2-3 / 9

Both packs: p505 第2 · p606 第3

“Purest limited-POV observational texture”; 90 half-width quote pairs recurred across books — an ingrained generation habit.

What the 20-round chains actually showed

Xuanhuan p303 #2 / p404 #3; romance p505 #2 / p606 #3 — the only model of nine stably in the top three on both genres. Judges' sketch: “flat and drift-free end to end,” “purest limited-POV observational texture.”

Its half-width-quote habit yielded a methodology lesson: our dialogue-rate regex initially only matched full-width quotes and scored it 0% dialogue — when it had actually written 108 quoted exchanges (a real rate of 13.5%). Blind judges independently caught the same habit and docked it for surface immersion breaks. Metric implementation details can fabricate signals; only cross-checking with blind review catches them.

Its failure shape deserves its own note: not drift that worsens over rounds, but a constant deviation present in round 1 and unchanged in round 20. For a reader that means predictability — whatever bothers you in chapter one is exactly what you'll see in chapter twenty, no worse.

Long-run failure mode

None observed

No long-run failure on either chain; its issue is a constant deviation (quote glyphs), not drift.

Structural fingerprint

The only model in the top three on both genres; half-width quotes throughout (108 pairs xuanhuan / 90 romance) — full-width-only detectors read 0% dialogue when the real rate was 13.5%, a documented metric artifact.

Cross-genre profile: The only stable top-three across both genres; half-width quotes are its cross-book stubborn habit.

Test-condition disclosure (hosted models are moving targets)

Model under test: glm-5.2 (released 2026-06-13)

Evaluated: 2026-07-16 · Access channel: Eval gateway (OpenAI-compatible pass-through)

Protocol: one 20-round continuous chain per genre · temperature 0.7 · double-blind full ranking with two shuffled mappings · structural-metric cross-checks

Directive condition: xuanhuan chains = legacy directive / romance chains = corrected D2 directive (full note in the leaderboard's method section)

What 10,000 characters cost

$0.34 list price: in $1.4/M · out $4.4/M (models.dev snapshot 2026-07-24)

Using the app's continuation recipe: one segment ≈ 400 chars = 8k input + 550 output tokens; 10k chars ≈ 25 segments; no cache discount. For between-model comparison only.

Foreverse connects to 60+ providers with your own keys — import your book and keep writing it with GLM-5.2.

Continue your book with it
GLM-5.2 for Fiction Writing — 20-Round Blind-Judged Test (Fiction Bench) · Foreverse · Xinmeng