Foreverse Research · Fiction Bench
How good is Kimi K3 at writing fiction?
It hasn't entered the nine-model full ranking, but its paired-duel record is clear: a 10:0 / 11:1 sweep of predecessor K2.6 — the older model's whole-passage self-copying (97.6% cross-round repetition) vanished, and stock-phrase density converged below the source baseline. Qwen3.8, released three days later, challenged it under the same protocol and lost 3:20. Known warts: ~85% of dialogue in half-width quotes, and thinking mode can't be turned off yet bills at full price.
Duel scorecard (paired double-blind, 6 judges × flipped mappings)
vs Kimi K2.6 · Same-vendor predecessor · board: #7 xuanhuan / 5-7 romance
Xuanhuan fantasy
10 : 0 (2 invalid)
Court romance
11 : 1
Tested 2026-07-18
Swept every window (10-0/10-0/10-0 xuanhuan, 11-1/11-1/11-1 romance); the single dissenting romance vote contradicted itself under mapping flip — position bias, not a stable judgment.
vs Qwen3.8-Max-Preview · Same-month challenger (K3 defending) · not on the main board
Xuanhuan fantasy
10 : 1
Court romance
10 : 2
Tested 2026-07-21
Led all six windows — the closest was still 9:2; all three of the challenger's votes contradicted themselves under mapping flip.
Where it stands against the board
The headline is generational proof: K2.6's two board-documented ailments — the field's highest simile density and whole-passage self-copying mid-chain (97.6% cross-round 12-gram repetition at r9 of the archived chain) — both vanished in K3's same-protocol retest: repetition peaked at just 6.8% (xuanhuan r10, ≈0% elsewhere), and stock-phrase density converged to 2.02‰ on xuanhuan and 0.63‰ on romance (below the 1.0‰ source baseline). Judges credited it with “perfectly cloning the source's short paragraphs, standalone onomatopoeia and '~~~' punctuation habit” — micro-texture only gpt-5.6-terra had replicated in the nine-model round.
Against the two genre champions (the DeepSeek pair) there is only a same-ruler comparison, not a shared blind duel: on xuanhuan K3's composite distance of 0.388 lands between V4 Flash (0.356, that genre's blind champion) and V4 Pro (0.454); on romance its 0.396 doesn't reach V4 Pro's 0.280. Who beats whom has no exam paper yet — that duel is on the schedule.
Its standing as the current continuation strongman has corroboration: Qwen3.8-Max-Preview, released the same month, challenged it under the identical protocol and lost 3:20 (1:10 xuanhuan / 2:10 romance). Later incremental duels use K3 as the reference strongman.
The warts have ballots too: ~85% of dialogue in half-width quotes (the same stubborn habit as glm-5.2, a constant immersion break in Chinese typography); first-person density at 16.2‰ — 1.8× the source — on first-person period prose; thinking permanently on, eating 81-86% of completion tokens at full price, with 37-54s waits per round.
Test-condition disclosure (hosted models are moving targets)
Model under test: kimi-k3 (released 2026-07-16)
Evaluated: 2026-07-18 · Access channel: Moonshot official API (thinking always on)
Review format: paired double-blind verdicts (two flipped mappings per book against position bias), not the nine-model full ranking
Protocol matches the main board item by item (same two books · same 55% anchor · same D2 directive · 20-round chains · temperature 0.7), max_tokens=2800; paired baselines = the archived K2.6 romance chain plus a freshly run K2.6-D2 xuanhuan chain to eliminate the directive variable.
Honest limits
Kimi K3 has not entered the nine-model same-protocol full-ranking review, so the board's rank column does not apply to it — this page publishes only ballot-backed paired duels and invents no rank. When it joins the full ranking depends on the next full-board rerun.
Paired blind duels ran only against K2.6 (same-vendor predecessor) and Qwen3.8 — never in the same blind room as the DeepSeek pair, so its record can't be hard-compared against the nine-model board ranks (different review formats: full ranking there, paired verdicts here).
Single book, single anchor, single chain per book (n=1), matching the main board's per-model conditions; the two invalid xuanhuan ballots were a claude judge refusing JSON output and writing its own continuation instead — excluded as recorded.
What 10,000 characters cost
$0.81 list price: in $3/M · out $15/M (models.dev snapshot 2026-07-24)
Using the app's continuation recipe: one segment ≈ 400 chars = 8k input + 550 output tokens; 10k chars ≈ 25 segments; no cache discount. For between-model comparison only.
Foreverse connects to 60+ providers with your own keys — import your book and keep writing it with Kimi K3.
Continue your book with itHow to cite
Foreverse Research, “How good is Kimi K3 at writing fiction (Fiction Bench incremental duels),” 2026-07. https://foreverse.cn/research/fiction-bench/kimi-k3