Foreverse Research · Fiction Bench
How good is Gemini 3.1 Pro at writing fiction?
One thing first: its last-place xuanhuan rank was mostly our own instruction-wording bug — the old phrasing made it treat AI-continued passages as “ignorable reference,” so it rewrote the opening for 20 straight rounds; with corrected wording, 0/20 rewinds. Under the fixed condition it placed a mid-field 4-6 on romance — but its simile-stock density stayed the field's highest. The looping was our bug; the purple prose is its nature.
Xuanhuan fantasy
9 / 9
Both packs: p303 第9 · p404 第9
Under the old directive it rewound to the continuation point 20/20 rounds (rewriting the opening every time, zero progress) — ablation proved this was a bug in our instruction wording; 0/20 after the fix.
Court romance
4-6 / 9
Both packs: p505 第6 · p606 第4
With the fixed directive it climbed from disqualified to mid-field; its simile-stock density of 2.50‰ stayed the field's highest — the purple prose is its nature, the looping was our bug.
What the 20-round chains actually showed
Its xuanhuan chain (old directive) replayed the same opening in all three windows — 20 rounds, zero progress, a unanimous #9 in both packs. A three-way ablation later dissected the disqualification: old wording (“for plot-transition reference only”) → 20/20 rewinds; explanation removed → ~6/8; corrected wording (“established story canon”) → 0/20. Faced with an unexplained source-marker glyph, its default reading was “annotated text isn't canon.”
The fix principle graduated into the product: the explanation sentence declares status only (established canon, not ignorable reference) and never commands actions. DeepSeek-family models were insensitive to all three wordings; Gemini was the easiest to mislead — so instruction robustness gets designed against it.
On the romance track (fixed directive) it competed normally at 4-6 — and its 2.50‰ flavor density stayed the field's highest. What remains after our bug was fixed is the model's own temperament: a constant, non-drifting lean toward ornate simile.
Long-run failure mode
Restart loop
Restart loop (triggered by our old directive): every round returned to the source's last line and rewrote the opening — 20 rounds, zero progress. Three-way ablation: old wording 20/20 rewinds, no explanation ~6/8, fixed wording 0/20. Its default reading of an unexplained marker was “annotated text isn't canon”; the explanation sentence is load-bearing.
Structural fingerprint
Highest simile-stock density in the field (2.47‰ xuanhuan post-fix / 2.50‰ romance vs the 1.0‰ source baseline); constant purple-prose lean.
Cross-genre profile: The directive-wording fix was its watershed: from disqualified to normal competition.
Test-condition disclosure (hosted models are moving targets)
Model under test: gemini-3.1-pro (released 2026-02-19)
Evaluated: 2026-07-16 · Access channel: yunwu aggregator gateway
Protocol: one 20-round continuous chain per genre · temperature 0.7 · double-blind full ranking with two shuffled mappings · structural-metric cross-checks
Directive condition: xuanhuan chains = legacy directive / romance chains = corrected D2 directive (full note in the leaderboard's method section)
What 10,000 characters cost
$0.56 list price: in $2/M · out $12/M (models.dev snapshot 2026-07-24)
Using the app's continuation recipe: one segment ≈ 400 chars = 8k input + 550 output tokens; 10k chars ≈ 25 segments; no cache discount. For between-model comparison only.
Foreverse connects to 60+ providers with your own keys — import your book and keep writing it with Gemini 3.1 Pro.
Continue your book with itHow to cite
Foreverse Research, “How good is Gemini 3.1 Pro at writing fiction (Fiction Bench),” 2026-07. https://foreverse.cn/research/fiction-bench/gemini-3-1-pro