Qwen 3.8 vs 3.7 on two novels: an 11-1 rout and a 7-5 squeak

Qwen3.8-Max-Preview shipped July 19, 2026; within 48 hours we ran it through the same protocol as our Kimi K3 test: two Chinese novels, 20 consecutive continuation rounds each, paired double-blind against its predecessor Qwen 3.7 Max. The palace novel flipped 11-1 for the new model; the fantasy epic barely moved at 7-5 with a 6-6 opening window. The twist: structural statistics say 3.7 is closer to the originals, and 3.8 writes the most uniform sentence lengths we have ever measured — yet the blind vote still went to 3.8. Honest caveats: sensory pile-ups in place of plot, a 20-round fantasy deadlock, one 'brewing the night into poison' metaphor in a period novel, and always-on reasoning eating 50-90% of output tokens. All results pinned to the 2026-07-21 hosted preview.

Two ink sketches of the same orchid side by side on aged paper: the left one wrapped in fog with only an outline visible, the right one clear down to its leaves, and an hourglass still running in the corner of the page

In our nine-model benchmark this month, Qwen 3.7 Max earned the most quotable dismissal on the board: “a consistently polished sensory stream — smells and textures the original never writes; consistently competent, consistently unlike.” It placed 4-5 on the fantasy epic and dead eighth on the palace novel. On July 19, at the World AI Conference in Shanghai, Alibaba previewed its successor: Qwen3.8-Max-Preview, 2.4 trillion parameters, with an official announcement calling it “second only to Fable 5”. No benchmark table, no model card, no activated-parameter count accompanied the claim. SiliconANGLE noted that the previous flagship, Qwen3.7-Max, had shipped in May with a full set of published results. We do not argue with slogans; we re-issue the exam. Within 48 hours, 3.8 and its predecessor each continued the same two novels for 20 consecutive rounds, judged pairwise and blind.

Half the result is a surprise and half is not. The palace novel went 11-1, a rout. The fantasy epic went 7-5, with the opening window tied 6-6. One book got a new model; the other got a nose ahead.

Same exam, one parameter moved

The protocol is copied item by item from the K3 increment test: the same two books (an 8.9-million-character Chinese fantasy epic, plus the palace classic Empresses in the Palace), the same anchor at roughly 55% depth, the same context-labeling instruction, the same 21,500-character rolling window, temperature 0.7, each round’s output appended back into context, 20 rounds per chain. The one parameter we moved was max_tokens, from 2,800 to 6,000: Qwen 3.8’s reasoning cannot be switched off, and without the higher ceiling the visible prose gets squeezed out by the thinking. Both generations ran on the same settings, so the variable is still just the model.

Why a paired increment against the predecessor rather than re-running the whole nine-model ladder? Because the ladder takes weeks and a preview does not wait. A paired vertical comparison answers the one question an upgrade actually poses (better or worse than the thing it replaces) with every variable except the model held still, and it fits inside the launch window while the hosted endpoint is still what people can subscribe to.

The judges are the same six heterogeneous models (claude-4.6-sonnet, gemini-3.1-pro-preview, gemini-3.5-flash, deepseek-v4-pro, qwen3.7-max, glm-5.2), two flipped blind mappings per book, votes counted on the overall verdict, ballots dropped when a judge’s JSON failed to parse: 12 valid votes per match. The reference excerpt in each blind pack is cut from before the anchor point, so no judge ever saw the original’s actual next passage. One detail worth keeping: qwen3.7-max sat on that panel, and on the palace novel it voted for 3.8 under both mappings; on the fantasy book it sided with its own generation.

The votes: a rout and a squeak

Match (Qwen 3.8 vs 3.7)Overall voteBy window: early / mid / late
Fantasy epic (8.9M characters)7-56-6 / 7-5 / 7-5
Palace novel (Empresses in the Palace)11-1all windows 11-1 or 11-0

The windows slice each 20-round chain into early, middle, and late thirds, voted separately, so a chain that opens strong and collapses late has nowhere to hide. The palace novel is the headline. All three review windows went the same way; no stretch of the chain wavered. Against history: Qwen 3.7 ranked eighth of nine on this book, its polished sensory stream reading “even more foreign” in period court prose, per the benchmark reviewers. Qwen 3.8 moved the genre from the bottom tier into contention in one generation. The fantasy epic is far more modest: a 6-6 tie early, 7-5 overall. The direction is improvement; the margin is a breath.

One sentence on the other duel from the same week: paired blind against Kimi K3 (released July 16), Qwen 3.8 lost 3-20 across the two books combined, and K3 remains the current continuation champion (full ballots and judge comments in the head-to-head piece). This page stays on the vertical line, 3.7 to 3.8.

The stats say 3.7, the judges say 3.8

Put the structural metrics on the table and you get a verdict that argues with the ballot box. Composite distance measures how far a chain sits from the original’s structural profile (shape, not content; lower is closer). Fantasy: 3.7 at 0.244, 3.8 at 0.442, nearly double. Palace: 3.7 slightly closer too, 0.321 against 0.385. By the numbers, the more original-like model is 3.7. The blind vote went to 3.8 on both books, 11-1 on one of them.

ChainComposite distance (lower is closer)Sentence-length CVDialogue rate
Fantasy · original baseline0.83716.1%
Fantasy · Qwen 3.80.4420.318 (late window 0.217)4.5%
Fantasy · Qwen 3.70.2440.5085.5%
Palace · original baseline0.49548.8%
Palace · Qwen 3.80.3850.39333.6%
Palace · Qwen 3.70.3210.40330.8%

The table also carries 3.8’s newest vital sign: the flattest sentence-length variance in the field, between 0.217 and 0.406 across windows, far below the originals’ 0.837 and 0.495. In plain terms, its sentences run more and more uniform. Uniform sentence length is the first AI fingerprint we ever measured, and it should have cost points; the human-feel vote went to 3.8 anyway. This is the third time our statistics and our blind judgments have pointed in opposite directions: the first was Qwen 3.7 acing nearly every metric and still placing 4-5 in the fantasy benchmark; the second was our judge-reliability experiment, where judges agreed with each other and were collectively wrong. Statistics as a regression gate, paired blind review as the verdict. That division of labor got confirmed once more. One side note: dialogue rate stays below the originals on both books; the upgrade did not move in that direction.

The repetition checkup is clean on both generations: adjacent-round verbatim overlap across all four chains measured 0.0%, none of the whole-paragraph self-copying we have caught in other model families. Qwen 3.8’s diseases live in the judges’ comments instead.

Which three degradations did the judges name?

A sourcing note first: the case-file quotes below come from the judge panel of the same chains’ other duel, the one against K3. Both matches reused the same two 3.8 chains, and that panel picked at the flaws harder. First, sensory pile-up in place of narrative progress. The claude judge on the fantasy chain, translated from the Chinese: “salt-crust shattering / sweet-metallic blood tang / withered leaves and scorched earth, stacked in place of narrative advance.” The qwen3.7-max judge’s version: “the style keeps drifting from the original, toward a fine-grained, microscopic traditional-wuxia register.” For contrast, 3.7’s cited disease was word-level looping: the same hissing onomatopoeia and sticky, teeth-aching textures recycled round after round, the same stock “lips curling faintly” micro-expression sliding into scene after scene. 3.8 fixed the word-level loop; the sensory bias itself stayed.

Second, and worse: plot deadlock. The gemini-3.5-flash judge: “severe plot rewind — after 20 rounds it is still stuck in a dead loop of fleeing and sealing the blood sigil; the narrative has stalled.” The palace book had a milder case of the same family, per the qwen3.7-max judge: “plot-pattern repetition: all three windows revolve around Wei Lin testing for poison, with no substantive progress.” This is the mid-run freeze from our three long-run failure modes: nothing repeats verbatim, the story just circles in place. Trading word-level loops for plot-level loops is the main reason the fantasy match was won by only a breath.

Third, modern literary flourishes. The palace chain produced a line that judges in the K3 duel held up for display: “brewing the night into poison” (translated), a metaphor that instantly breaks period register in a court novel set centuries ago. Curiously, two judges in this match against 3.7 quoted the same line as a point in 3.8’s favor, proof of how divisive this register is. A casual skim forgives a slip like this; a reader four hundred chapters deep does not.

Slow, pricey, and zero cache hits

Reasoning is always on: 50-90% of completion tokens are invisible thinking, and rounds took 12 to 38 seconds against 3.7’s steadier 21 to 29 — faster at its fastest, slower at its slowest. For scale, last week’s K3 test measured 37 to 54 seconds per round; waiting is simply what always-on reasoning models cost you at the button. The four chains, 80 rounds, consumed 1,227,401 input tokens and 103,229 output tokens with zero cache hits: a rolling continuation window changes its prefix every round, so this workload never touches prefix-cache discounts. Inside a reader, the asymmetry matters more than the average: a 12-second round feels fine after pressing the continue button, a 38-second one does not. We ran on a Token Plan Lite subscription (limited-time 39 CNY/month), with preview-period credits metered at 10% of standard burn; the whole experiment did not exhaust the plan’s quota.

How long does this page stay true?

A preview is a moving target. Alibaba’s own documentation states that qwen3.8-max-preview keeps iterating during the preview period and will be taken offline or replaced by a production version afterward. Every vote and metric on this page is therefore pinned to the hosted preview as of July 21, 2026 (Token Plan’s OpenAI-compatible endpoint, model IDs qwen3.8-max-preview and qwen3.7-max), and we make no promise the final release reproduces any of it. The open-weight release Alibaba has promised is the date of the second exam. When the model graduates from preview, we re-run the same 40 blind rounds, fold the result into the evergreen model guide, and pin a pointer to the new test at the top of this page.

A practical note to end on. During a preview window, resist migrating a whole book. Feed the newest chapter of whatever you are reading to both generations, three continuations each, shuffle the six passages, read them blind, then check the labels: a few cents buys an answer accountable to your book instead of ours. And the limits stay on the record (one chain per book, a single anchor point per chain), so our ballot can inform your choice, but it cannot make it.

FAQ

Is Qwen 3.8 better than Qwen 3.7 for fiction?

Depends on the genre. On the palace-intrigue novel it is a rout: same protocol, 20 consecutive continuation rounds each, six heterogeneous AI judges under two flipped blind mappings, 11-1 overall with every review window at 11-1 or 11-0, a clear break from 3.7's eighth-of-nine placement on that book. On the fantasy epic it is a squeak: 7-5 overall, with the opening window tied 6-6. The predecessor's signature polished sensory stream has receded, but judges still flagged sensory pile-ups and a 20-round plot deadlock. All results are pinned to the hosted preview as of July 21, 2026.

Is Qwen 3.8 better than Kimi K3 for fiction?

No. In paired double-blind under the same protocol, Qwen 3.8 lost to Kimi K3 3-20 across the two books combined (1-10 on the fantasy epic, 2-10 on the palace novel), and K3 remains the current continuation champion. Full ballots, judge comments, and the cost comparison live in our separate Qwen 3.8 vs Kimi K3 head-to-head; this page is about the vertical 3.7-to-3.8 comparison.

How much does Qwen 3.8 cost, and how fast is it?

We ran it on an Alibaba Token Plan subscription (verified 2026-07-21): the China-mainland personal Lite tier is 39 CNY/month at the limited-time price (the international edition lists Lite at $8, limited-time $6), and qwen3.8-max-preview credits are metered at 10% of standard burn during the preview promotion. Measured speed: 12 to 38 seconds per round, reasoning always on, thinking tokens taking 50-90% of completion output. The whole experiment (four chains, 80 rounds, 1,227,401 input tokens and 103,229 output tokens, zero cache hits) did not exhaust the Lite quota.

Will these results hold for the final Qwen 3.8 release?

No promises. Alibaba's documentation states the preview model keeps iterating and will be taken offline or replaced by a production version afterward, so every vote and metric on this page is pinned to the hosted preview as of July 21, 2026. An open-weight final release has been announced without a date; when it ships, we will re-run the same 40 blind rounds, fold the result into our evergreen model guide, and pin a pointer to the new test at the top of this page.

Is Qwen 3.8 Good at Writing Fiction? 40 Blind Rounds Against Qwen 3.7, 48 Hours After Launch · Foreverse · Xinmeng