Foreverse Research · Fiction Bench

How good is Kimi K2.6 at writing fiction?

Lower-mid on both genres (#7 xuanhuan, 5-7 romance) with a clear diagnosis: the field's highest simile density (nearly one per paragraph) and whole-passage self-copying mid-chain (97.6% cross-round repetition at r9). Notably, successor K3 fixed both ailments in a same-protocol retest — see the incremental duels on the leaderboard.

Xuanhuan fantasy · Rank 7 / 9Court romance · Rank 5-7 / 9Mid-chain collapse$0.24 / 10k

Xuanhuan fantasy

7 / 9

Both packs: p303 第7 · p404 第7

Highest simile density in the field (nearly one per paragraph), whole-passage self-copying from early to mid, and setting slippage (an ancient-tree valley sprouting inside a void blood-array).

Court romance

5-7 / 9

Both packs: p505 第5 · p606 第7

One pack flagged it as “the only clear reverse-improver” — webnovel rage-cadence early, settling down by the late window.

What the 20-round chains actually showed

A unanimous #7 on xuanhuan: field-highest simile density, whole-passage self-copying from early to mid, and setting slippage (ancient trees sprouting inside a void blood-array). The self-copying has a machine number: 97.6% cross-round 12-gram repetition at r9 of the archived chain.

Romance 5-7, where one pack gave it the run's only such label: “the only clear reverse-improver” — webnovel rage-cadence early, settling by the late window. Most models decay or hold constant; this one runs backwards.

It's one of two specimens of the mid-chain-collapse failure mode (the other: Grok 4.5). Same-vendor successor K3, retested under the identical protocol, converged stock-phrase density 3.02‰→2.02‰, eliminated verbatim failure, and beat K2.6 10:0 / 11:1 in paired blind review — generational upgrades can cure, provided someone actually retests.

Long-run failure mode

Mid-chain collapse

Mid-chain collapse (whole-passage self-copying early→mid): archived chain r9 hit 97.6% cross-round 12-gram repetition. Successor K3, retested under the same protocol, showed no verbatim failure (tt peak 6.8%) — see the incremental duels section.

Structural fingerprint

Field-highest simile density (nearly one per paragraph); archived xuanhuan chain r9 measured 97.6% cross-round 12-gram repetition — machine evidence of whole-passage self-copying.

Cross-genre profile: Lower-mid on both; the simile-density ailment crosses genres.

Test-condition disclosure (hosted models are moving targets)

Model under test: kimi-k2.6 (released 2026-04-21)

Evaluated: 2026-07-16 · Access channel: Eval gateway (OpenAI-compatible pass-through)

Protocol: one 20-round continuous chain per genre · temperature 0.7 · double-blind full ranking with two shuffled mappings · structural-metric cross-checks

Directive condition: xuanhuan chains = legacy directive / romance chains = corrected D2 directive (full note in the leaderboard's method section)

What 10,000 characters cost

$0.24 list price: in $0.95/M · out $4/M (models.dev snapshot 2026-07-24)

Using the app's continuation recipe: one segment ≈ 400 chars = 8k input + 550 output tokens; 10k chars ≈ 25 segments; no cache discount. For between-model comparison only.

Foreverse connects to 60+ providers with your own keys — import your book and keep writing it with Kimi K2.6.

Continue your book with it
Kimi K2.6 for Fiction Writing — 20-Round Blind-Judged Test (Fiction Bench) · Foreverse · Xinmeng