Foreverse Research · Fiction Bench
How good is Gemini 3.6 Flash at writing fiction?
Opposite directions on the two books: it beat predecessor 3.5 Flash 9:3 on xuanhuan but lost romance 1:10 (five judges stable for 3.5 across both mappings). The essence isn't genre preference — it's who crashes uglier: under this 20-round protocol both flash generations fell into plot loops; on xuanhuan 3.5 collapsed to the verbatim level, on romance it was 3.6 that decayed into a verbatim death loop. A generational update isn't a monotonic upgrade — it's a redistribution of failure modes.
Duel scorecard (paired double-blind, 6 judges × flipped mappings)
vs Gemini 3.5 Flash · Same-vendor predecessor · not on the main board
Xuanhuan fantasy
9 : 3 (12 ballots)
Court romance
1 : 10 (11 ballots)
Tested 2026-07-23
Of the 9:3 xuanhuan tally, the mapping-stable core is 3:0 plus six swing ballots (a medium signal); on romance five judges were mapping-stable for 3.5 (a strong signal).
Where it stands against the board
The board's Gemini data point is gemini-3.1-pro: its xuanhuan #9 was proven by ablation to be mostly our own instruction-wording bug (0/20 rewinds after the fix), and it placed 4-6 on romance under the fixed condition. This duel adds the flash-tier data point: both flash generations still loop at scale under the corrected D2 directive — D2 fixed the “AI passages are skippable” misreading; flash-tier looping is a long-run capability problem that prompt wording can't cure.
Three of four chains crashed, and the two generations crash differently: 3.6's xuanhuan chain reads clean on mechanical metrics yet rewinds semantically (the same discover-the-demon-lord beat replayed 6-7 times in fresh wording — r18 had already repelled him, r19 is back to “just discovered”); 3.5's xuanhuan chain copied itself verbatim (two rounds at 100% repetition, plot frozen at the array-breach moment); 3.6's romance chain decayed monotonically into a verbatim death loop (33%→64%→73%→92%→100%); 3.5's romance chain replayed intermittently. Judge callouts matched the machine readings item by item — no hallucinated accusations.
The flash tier is a speed/price slot, not a quality slot: 3.6 runs each round in roughly half 3.5's latency (6.3s vs 12.5s) at $1.5/$7.5 per million tokens list — but dialogue rates sit far below the sources on both books (7.8% / 27.1% vs 16.1% / 48.8%), and its structural metrics trail the kimi-k3 and DeepSeek champion tiers.
Test-condition disclosure (hosted models are moving targets)
Model under test: gemini-3.6-flash (released 2026-07-21)
Evaluated: 2026-07-23 · Access channel: Eval gateway (OpenAI-compatible aggregator, not Google's official endpoint)
Review format: paired double-blind verdicts (two flipped mappings per book against position bias), not the nine-model full ranking
Matches the qwen3.8 incremental round item by item (same two books · same 55% anchor · same D2 directive · 20-round chains · temperature 0.7 · max_tokens=6000); judge gemini-3.5-flash was replaced by kimi-k2.6 since it was under test.
Honest limits
Gemini 3.6 Flash has not entered the nine-model same-protocol full-ranking review, so the board's rank column does not apply to it — this page publishes only ballot-backed paired duels and invents no rank. When it joins the full ranking depends on the next full-board rerun.
The access channel is an eval-gateway aggregator, not Google's official endpoint; 3.6-flash was a 48-hour-old hosted target, and conclusions bind to the gateway's 2026-07-23 behavior.
Opponent 3.5 Flash isn't on the main board, so this duel has no board rank to anchor against; the board's Gemini data point is 3.1-pro (a pro-tier model that can't stand in for the flash tier).
Judge-pool change: the k3/qwen3.8 rounds included gemini-3.5-flash as a judge; here it's a contestant, so kimi-k2.6 substituted — mind this when comparing tallies across rounds.
The loop ailment's scene sensitivity (3.6 clean on xuanhuan, crashed on romance) can't separate book effects from anchor luck — n=1 chain per book.
What 10,000 characters cost
$0.40 list price: in $1.5/M · out $7.5/M (models.dev snapshot 2026-07-24)
Using the app's continuation recipe: one segment ≈ 400 chars = 8k input + 550 output tokens; 10k chars ≈ 25 segments; no cache discount. For between-model comparison only.
Foreverse connects to 60+ providers with your own keys — import your book and keep writing it with Gemini 3.6 Flash.
Continue your book with itHow to cite
Foreverse Research, “How good is Gemini 3.6 Flash at writing fiction (Fiction Bench incremental duels),” 2026-07. https://foreverse.cn/research/fiction-bench/gemini-3-6-flash