Field notes

Notes from the reading tavern

We are building a pocket reader and a pocket tavern. These are the potholes, trade-offs, and field tests we hit along the way — written by the people doing the work.

A pinboard of paper strips, most crossed out in red ink, a hand with a red pen marking another
Data noteJul 21, 20269 min

The AI roleplay slop lexicon: 95 phrases, with receipts

A maintained, sourced reference of the 95 phrases that mark AI-generated roleplay prose: shivers down spines, whispers barely above themselves, mischief-sparkling eyes, and the "not X, but Y" construction. Cross-checked from Sukino's Banned Tokens, the Antislop paper (arXiv 2510.15061 — "Elara" runs 85,513x over the human baseline), and EQ-Bench's slop score. Plus two numbers most lists miss: top human-written cards average slop hits too, and flagship models' single-shot output now passes word-list checks clean.

Read the post →
A writing desk with character sketches, an index card being measured with a ruler, a fountain pen and an inkwell
GuideJul 21, 202611 min

How to write AI character cards, calibrated against what actually charts

We pulled the 30 top-starred SFW character cards from a major card hub and measured everything: first messages run a median of 178.5 words, 0 of 30 use HTML, 67% write the description in plain prose, 70% ship example dialogues, only 20% carry a lorebook. This guide turns those numbers — plus the community guide canon — into a writing procedure: format, token budget, the no-user-actions rule, greetings, example dialogs, and the anti-slop pass.

Read the post →
Ink-line illustration on aged paper: a long scroll of chat messages winding across the page, its middle section fading into blank mist, a hand with tweezers slotting a small written card back into the gap, a tilted-head cat sitting at the scroll's end
Model evalsJul 21, 20269 min

Your AI companion forgot your cat. Would a smarter model remember?

“My AI roleplay partner forgets everything” is the most common complaint in companion chat, and the folk remedy is always the same: switch to a bigger model. We put that remedy through a 400-turn companion conversation with 50 planted memory probes, three models under one protocol: deepseek-v4-pro, qwen3.8-max-preview, qwen3.7-max. On bare context all three vendors are amnesiac (19–22%, barely distinguishable). Add retrieval-injected memory and they fan out into tiers: 89%, 74%, 59%. Ask about things never said, and even the most honest one passes only 5 of 7. The upgrade dividend is real — just not where users think it is. Tested 2026-07-21.

Read the post →
Warm-paper ink illustration: two inkstones facing off across an open manuscript page, twenty paper ballot slips stacked beside the right stone and only three beside the left, a single brush resting on the page's center fold
Model evalsJul 21, 20269 min

Qwen 3.8 vs Kimi K3: two trillion-scale models shipped three days apart, so we put them in the same exam room

Two Chinese labs shipped trillion-scale models within three days: Kimi K3 (2.8T parameters) on July 16, Qwen3.8-Max-Preview (2.4T, self-described as second only to Fable 5, no independent benchmarks attached) on July 19. We ran the first paired blind fiction exam between them: same two novels, same anchor point, same instructions, 20 consecutive continuation rounds each, six AI judges under flipped mappings. Final score 3:20 (1:10 on the fantasy epic, 2:10 on the palace novel), and every one of Qwen's three ballots contradicted itself when the mapping flipped. Both sides have documented flaws: Qwen3.8 was cited for a severe plot rewind and a modern-literary metaphor in period prose; K3 was cited for formulaic plot recycling, on top of its straight-quote and always-on-reasoning habits. Preview models are moving targets; every conclusion here is pinned to the hosted endpoint as of 2026-07-21.

Read the post →
Two ink sketches of the same orchid side by side on aged paper: the left one wrapped in fog with only an outline visible, the right one clear down to its leaves, and an hourglass still running in the corner of the page
Model evalsJul 21, 20268 min

Qwen 3.8 vs 3.7 on two novels: an 11-1 rout and a 7-5 squeak

Qwen3.8-Max-Preview shipped July 19, 2026; within 48 hours we ran it through the same protocol as our Kimi K3 test: two Chinese novels, 20 consecutive continuation rounds each, paired double-blind against its predecessor Qwen 3.7 Max. The palace novel flipped 11-1 for the new model; the fantasy epic barely moved at 7-5 with a 6-6 opening window. The twist: structural statistics say 3.7 is closer to the originals, and 3.8 writes the most uniform sentence lengths we have ever measured — yet the blind vote still went to 3.8. Honest caveats: sensory pile-ups in place of plot, a 20-round fantasy deadlock, one 'brewing the night into poison' metaphor in a period novel, and always-on reasoning eating 50-90% of output tokens. All results pinned to the 2026-07-21 hosted preview.

Read the post →
Two ink-line plants side by side on aged paper: the left one twisting into knots around itself, the right one upright and blooming, with a small stopwatch resting at the right plant's roots
Model evalsJul 18, 20268 min

Is Kimi K3 good for fiction? We made it write 40 rounds to find out

Kimi K3 shipped July 16, 2026; within 48 hours we ran it through the same protocol as our nine-model benchmark — two Chinese novels, 20 consecutive continuation rounds each, paired double-blind against its predecessor K2.6. The vote: 10-0 on the fantasy epic, 11-1 on the palace novel, and the single dissent contradicted itself under mapping flips. K2.6's metaphor pile-ups and verbatim self-copying did not recur. Three honest caveats: ~85% of dialogue in Western straight quotes (a disease K2.6 never had), first-person density 1.8× the original author's, and always-on reasoning that makes every round take 37-54 seconds. Total bill for the whole experiment: ¥12.21.

Read the post →
A vast library stuffed with books rising into darkness, and at its center one small writing desk under a single warm lamp
Research notesJul 18, 20268 min

Every flagship ships 1M context now. Your novel fits. That's not the question

Kimi K3, DeepSeek V4, Claude Fable 5 and GPT-5.6 all carry million-token windows as of July 2026 — a whole 500k-word serial fits in one request. Our five-tier control experiment says fitting is not learning: a model handed an entire novel with no instruction wrote clichés at ten times the original author's density, and one explicit sentence did what 200,000 tokens of raw prose couldn't. Plus the arithmetic: filling K3's window costs about $3.15 per press.

Read the post →
A thick kraft-paper dossier pulled from a filing cabinet drawer, its exposed manuscript pages covered in one ink pattern repeating in loops
Model evalsJul 18, 20267 min

Grok 4.5 shipped ten days ago. For fiction, we already had a file on it

Every Grok 4.5 review benchmarks coding. Nobody answers the question people actually type into search: is it any good for stories? As it happens, Grok 4.5 was one of nine models in our long-run continuation benchmark — 20 consecutive rounds on each of two novels, ranked by double-blind review. The file shows a twice-reproduced repetition loop, a dead-last palace-intrigue placement from both reviewers, a first-person density at half the original's — and the scenarios where it is genuinely fine.

Read the post →
A wall calendar with several pages torn off and drifting down, beside a row of old and new keys on a desk, hand-drawn ink linework on warm paper
Time-sensitiveJul 18, 20269 min

Six model sunsets and price changes this summer: one page, kept current

At least six model shutdowns and price changes land between July 20 and August 31, 2026: Claude Fable 5 moves to usage credits on July 20, deepseek-chat and deepseek-reasoner stop resolving on July 24, GitHub Models retires entirely on July 30, Moonshot V1 and kimi-k2.5 sunset on August 31, and Sonnet 5's introductory pricing ends the same day. Every date verified against official announcements, with replacement picks for fiction-continuation users backed by our 360-round benchmark.

Read the post →
Two writing desks facing each other across a narrow gap: the left manuscript stacked print-neat, the right one scribbled over with loose pages sliding off, a single ink pen resting between them
Model evalsJul 18, 20268 min

GPT-5.6 vs Kimi: two personalities on the same exam

GPT-5.6 Terra and Kimi K2.6 each continued two Chinese novels for 20 consecutive rounds in our double-blind benchmarks. Terra is the best prose mimic we have measured and ran the palace-intrigue novel with zero incidents — then rewound the fantasy plot in rounds 18-20, reusing its own round-2 lines. Kimi placed seventh in fantasy with the highest metaphor density in the field, yet was the only system that improved as it went. Honest cutoff: Kimi K3 shipped July 16, 2026 and was never in this exam.

Read the post →
An immensely long manuscript scroll passing beneath a small reading window; only the lines inside the window glow warm cinnabar while the rest of the scroll fades to pale grey
ExplainerJul 18, 20268 min

Can the AI read your whole serial at once? Do the math first

The AI forgetting chapter 3 by chapter 40 isn't a memory problem — it's a window problem. A 500k-word serial runs roughly 625k to 830k tokens; 2026 flagship windows reach 1M, so it technically fits. But fitting is not remembering: mid-context retrieval dips are documented in research and in vendors' own evals, and our own control experiment shows material being in the window doesn't mean it gets used. The reader's version, with arithmetic.

Read the post →
One cinnabar teapot pouring into four identical cups, each holding tea of a visibly different shade, while a hand lifts two cups to compare them against the light
Research notesJul 18, 20268 min

Same model, same book, wildly different output. Four reasons, and two of them are yours

"Same prompt, wildly different quality" is four unrelated variance sources stacked on top of each other: sampling temperature (yours to tune), context composition that silently changes as the window slides (yours to structure), batch-dependent server numerics (not yours — at temperature 0, 1,000 identical requests returned 80 distinct outputs in a published test), and the state of the person doing the judging. A breakdown, with knobs.

Read the post →
An antique desk under lamplight with a stack of letters, an open leather tome, and tentacle-like ink lines rising from the page into branching paths above a small phone screen
GuideJul 18, 20269 min

Lovecraft ran a shared universe by mail. The membership requirements just dropped

In 1935 Lovecraft signed a mock certificate authorizing Robert Bloch to kill him in a story — that is how collaborative the Cthulhu Mythos was from the start. A century later the core texts are public domain (everything through 1930 unambiguously so in the US, the rest backed by decades of renewal research), Project Gutenberg hosts the canon, and AI continuation puts the old writing-circle game in anyone's pocket. What the shared-world tradition looked like, exactly which stories are safe to build on, how a lorebook keeps Mythos lore straight across stories, and an honest note on why generic cosmic horror is the default failure mode.

Read the post →
A Regency writing desk with a quill pen beside an open hardbound novel; from its last page several fresh manuscript sheets fan out, tied with a cinnabar ribbon, a teacup standing nearby
GuideJul 18, 20268 min

Austen readers have been writing the sequel since 1913. Your turn

Pride and Prejudice has roughly 900 published spinoffs — a tradition that starts with Sybil Brinton's 1913 Old Friends and New Fancies and runs through P.D. James and Jo Baker. The JAFF community even has a name for the fork-the-canon format: variations. A practical guide to joining in with AI: where the public-domain text lives (Project Gutenberg #1342), which entry points two centuries of readers keep choosing, what an LLM can and cannot do with Austen's voice, and why branches fit this fandom's native format better than any other tool.

Read the post →
Warm-paper ink illustration: an open novel with three small blank art plates rising from the page like prints on a drying line, a nib pen resting beside the spine
IllustrationJul 18, 20268 min

Scene-to-image prompts that stop fighting your novel

Four copy-ready prompt templates for illustrating fiction — establishing shot, character close-up, action freeze-frame, quiet interior — plus the documented quirk of each major image model and the fix for it: GPT Image's fixed size grid, Nano Banana's preference for narrative prompts over keyword lists, Seedream's front-loaded attention. Checked against the official docs on 2026-07-18. Ends with the one thing no prompt solves: keeping the same face across fifty images.

Read the post →
Four closed rulebooks of different thicknesses on a desk, each with a ribbon bookmark at a different page, beside a brass key and a small reading lamp
PoliciesJul 18, 20269 min

NSFW roleplay and the four rulebooks: what the policies actually say, checked July 2026

Whether an LLM will write adult roleplay is governed by three different layers people keep conflating: the app's filter, the provider's usage policy, and the model's trained refusals. We read the current policy documents from OpenAI, Anthropic, Google, and xAI — with dates — and map where each lab draws its lines, what enforcement looks like on an API account, and what BYOK does and does not change.

Read the post →
A clinic desk with three patient folders tied in different ribbons; above them hover their remedies: a card with speech bubbles, a keyed index drawer, and a looping ribbon being cut by scissors
RoleplayJul 18, 20268 min

OOC is not one disease. It is three, and they need different medicine

When an AI roleplay character breaks — ignores a detailed card, flips personality mid-arc, or slowly turns generic over weeks — players file it all under OOC. Those are three separate failures: a card that describes instead of demonstrates, facts that outran the context window, and a feedback loop where the model imitates its own replies. A triage question for each, plus fixes with dosages and the 20-round experiment data behind them.

Read the post →
Hand-drawn ink illustration: a hand rising from an old phone's glowing screen, holding an open miniature suitcase packed with letters, photos, and a cinnabar-red heart, a dotted arc leaping toward a new phone beside it
CompanionJul 18, 20268 min

New phone, same companion: a migration that fits in one zip

How do you move an AI companion to a new phone? For cloud apps like Replika or Kindroid, you log in and the server hands your history back. For a companion stored on your device, you carry it yourself: Foreverse exports one package — persona, every memory entry, full chat transcripts, anniversaries, journals, the relationship archive — and the import on the new phone restores the relationship mid-conversation. This is the walkthrough, plus the honest trade against account-based transfer.

Read the post →
Hand-drawn ink illustration: an open pocket notebook with short handwritten lines, one line circled in cinnabar red, a pencil and eraser resting beside it on warm paper
CompanionJul 18, 20267 min

Your companion misremembered your birthday. Here's the screen where you fix it.

Correcting an AI companion in chat doesn't stick — chat slides out of the context window, memory entries don't. This walkthrough covers the memory screen in Foreverse: the ⋮ menu entry, search, the pencil (200-character entries), the trash can with its confirm dialog, the disable switch that keeps an entry without injecting it, and adding entries yourself. Changes apply from the next conversation; everything is a file on your phone, exportable as one package.

Read the post →
Hand-drawn ink illustration on aged paper: a library card catalog with one drawer open, a chain of index cards arcing toward an open book from which a tree of paper leaves grows
RoleplayJul 18, 20268 min

You mentioned one name. Five entries walked in.

"Why did one keyword pull five entries into my prompt?" That is recursive scanning doing its job: the content of an activated entry becomes scan text too, so entries can summon other entries. A spec sheet for the mechanism — how chains advance, the three things that stop them (dedup, depth, budget), what the three per-entry switches do, plus a reproducible five-entry runaway case and the two built-in tools that show you the chain.

Read the post →
Five small hand-drawn radios in a row on a wooden shelf, each with a different dial face, while warm sound waves rise from an open paperback below toward a pair of headphones
TTSJul 18, 20269 min

Your webnovel will never get a narrator. Here is the 2026 app map

Six ways to listen to webnovels in July 2026, each priced from its official page: Royal Road's built-in device TTS, @Voice Aloud Reader ($15 lifetime, per-character dialog voices), Moon+ Reader Pro ($11.99 one-time), ElevenReader (10 free hours a month, Ultra $11/mo), Speechify ($29/mo or $139/yr), and Foreverse's free system TTS plus BYOK neural voices at provider list price. With a pick-it-if verdict per app and the failure notes from our own three-week commute test.

Read the post →
Four thin ink footpaths fan out from a single open book across cream paper, each path passing under its own small gate; one gate stands ajar with a warm glow behind it
PricingJul 18, 20268 min

Four ways to run AI story continuation for $0 — and the catch in each

No subscription, no card: the four genuinely free routes to AI fiction in July 2026. Free chatbot sites (the catch is the container), a 5,000-credit signup grant worth about 260 continuations, real free API tiers — Gemini flash-class at roughly 10 requests a minute, OpenRouter's 50-per-day :free lane — and local models over Ollama with zero marginal cost. Each route priced, bounded, and given its honest failure point.

Read the post →
An open field-test logbook with tally marks; above it a looping ink line circles back to its starting post twenty times before one warm stroke finally breaks free toward the page edge
Field testJul 18, 20268 min

We benchmarked Gemini on two novels. Then we priced the free tier

Field notes from 40 benchmark rounds: Gemini 3.1 Pro restarted the story from the original ending 20 times out of 20, one rewritten sentence brought that to zero, and the palace-novel re-run landed it mid-table as the most literary voice of nine models. Plus what Google's AI Studio free tier really covers as of July 2026 — flash-class only since April, roughly 10 requests a minute, quotas per project — and the mismatch between the model we ranked and the model you get free.

Read the post →
A tall paper chat window crammed with overflowing pages stands at the left of a desk; one ink line carries a single page across to a small phone-shaped reader resting on an open novel
TutorialJul 18, 20268 min

Claude will continue your novel. The chat window is what stops you

Claude Opus 4.8 placed fourth to fifth in both genres of our 360-round blind continuation benchmark — good, volatile, never the winner. The real ceiling is the container: claude.ai Projects swap to retrieval on big books and overwrite every regeneration. This tutorial covers the honest ranking data, what Projects can and cannot do as of July 2026, and the four steps that wire an Anthropic API key into a phone reader, with per-continuation costs worked out.

Read the post →
A wooden card-maker's chest on a desk, painted character cards floating out of it into the air
ReleaseJul 17, 20266 min

Everything we know about writing character cards is now on GitHub

character-card-skills is live: two agent skills (card authoring + chat-quality triage), 47 genre playbooks, 15 original cards shipped as SillyTavern-ready v2 JSON and v3 PNG, and the rule-based AI-flavor detector we calibrated after our LLM judges failed gold calibration at 12%. Code MIT, content CC BY 4.0. Runs in Claude Code, Cursor, Codex, Gemini CLI — and natively in Foreverse on Android.

Read the post →
A Victorian writing desk with a magnifying glass resting on an open book; from the page, several ink-drawn branch lines grow toward a glowing phone screen
GuideJul 17, 20268 min

Writing new Holmes stories is a 130-year-old hobby. Now it fits in your pocket

Sherlock Holmes has been continued by other hands since the 1890s — thousands of pastiches, an estate-endorsed novel, 250+ screen portrayals. Since January 1, 2023 the entire canon is public domain in the US. A practical guide: where to get the source text (Project Gutenberg), why Watson's voice is a natural style anchor, what long-run failures to expect from real 360-round data, and how branching handles the Reichenbach what-if.

Read the post →
Four index cards pinned above a writing desk, each dense with handwritten rules and one struck-through phrase list; a fountain pen rests on the half-finished fifth card
PromptsJul 17, 20269 min

Prompting AI to continue a story: what holds up after twenty rounds

Bare 'continue this' prompts decay measurably over consecutive rounds — we have the 20-round benchmark data. These field notes split the working prompt into an instruction layer and a material layer, explain the continuity directive that took one model from 20/20 restart loops to 0/20, and include four copy-ready English templates (fast action, interior-emotional, suspense, literary restraint), each with banned-phrase lists and a note on when to use it.

Read the post →
An open paperback whose printed text stops mid-page; from the last line, a faint handwritten branch of new sentences curves off the page onto warm blank paper
ReaderJul 17, 20268 min

Finishing a dropped novel for yourself: the whole workflow, honestly

A web serial I followed went silent in 2024 at chapter 214, mid-scene. This is the actual workflow I used to give it an ending nobody else will ever read: getting the text out as a txt file, picking the real last-good chapter, growing the continuation on a branch so the original stays byte-for-byte intact — plus the honest part about style drift over long runs, and what happens if the author ever comes back.

Read the post →
A narrow paper ledger sheet beside a phone showing a story page; three ruled columns of tiny handwritten sums, a single coin resting on the smallest total
PricingJul 17, 20268 min

The full ledger: what continuing a novel with AI actually costs

One continuation measured about 19 credits — $0.0019 — on Foreverse's official channel. This ledger prices all three routes (5,000 free starter credits, pay-as-you-go packs from $0.99, BYOK at provider list price), works out what a 200,000-word ride costs on each, and runs the honest comparison against a $19/month writing subscription. Includes the 50% service-fee disclosure and when BYOK is the better deal.

Read the post →
A brass balance scale on a desk: one pan holds a phone glowing with story text, the other a stack of hardbound law reporters; warm lamplight falls only on the reader's side
LegalJul 17, 20269 min

Continuing someone else's novel with AI: what US copyright law actually says

A private AI continuation you never share and a continuation you post or sell sit in different risk classes under US copyright law. This explainer walks the 17 U.S.C. §107 fair use factors, what Salinger v. Colting and Anderson v. Stallone actually decided, why the AI training lawsuits are a different fight from your personal use, and a plain do/don't checklist. Background information, not legal advice.

Read the post →
Two podiums side by side — an ink-brush sword trophy on the left, an embroidered palace fan trophy on the right — with one row of model nameplates facing two different champions
Model evalsJul 17, 20268 min

Which LLM continues a novel best? We tested two genres — the winners don't overlap

Nine LLMs each continued two Chinese novels for 20 consecutive rounds — 360 rounds total, ranked by double-blind review. DeepSeek V4 Flash won the fantasy epic; DeepSeek V4 Pro and GPT-5.6 Terra took the top tier on the palace-intrigue novel; the fantasy champion dropped to sixth on the second book. Full ranking tables, a checklist of three long-run failure modes, and how to actually use the results.

Read the post →
Six calligraphy brushes copying the same stroke on one scroll; only the cinnabar-red one matches the original
Research notesJul 16, 20269 min

We had six models continue the same novel for 20 rounds each, then ranked them blind

DeepSeek V4 Pro/Flash, Claude Opus 4.8, Gemini 3.1 Pro, Qwen 3.7 Max and GLM 5.2 each continued an 8.9M-character Chinese fantasy novel from the same anchor point, 20 rounds each, output fed back into context. Double-blind review ranked the results. The winner isn't the most expensive model; one model rewrote the same opening paragraph 20 times; another aced every statistical metric and still placed fifth.

Read the post →
Hand-drawn ink illustration: a figure carries a wooden crate of glowing folders and a book out through an open door, tidy archive shelves behind, a key hanging by the door
ArchitectureJul 15, 20268 min

One world = one directory: data ownership as an engineering decision

Every app claims your data is yours; the storage architecture decides whether it's true. Foreverse's version: each world is a directory on your device, character cards are standard chara_card_v3 files, lorebooks are JSON, a companion packs into a moving-box zip, community imports carry provenance records, and AI-generated images get machine-readable origin marks. Here is the file-first architecture laid open — costs included.

Read the post →
Hand-drawn ink illustration: an old apothecary cabinet with many drawers, several open and glowing, toggle switches on the fronts, one drawer outlined in cinnabar
TavernJul 15, 20269 min

Eleven switches, and the trade-offs behind them

Memory, auto-summary, story choices, state tracking, stepped thinking, lorebook suggestions, reply ideas — the extension powers desktop tavern players rely on, built into Foreverse as a plugin center. A walkthrough of what each plugin does and costs, plus the three disciplines we hold: every side-request is itemized in your billing log, every turn has a call budget, and tapping a choice never sends on your behalf.

Read the post →
Hand-drawn ink illustration: four silhouettes around a round table, one standing with a glowing cinnabar-outlined speech bubble while three grey bubbles wait
TavernJul 15, 20268 min

Four characters at one table, and the director is an algorithm

Single-character AI chat is a solved genre. Group roleplay's hard problem lives elsewhere: turn-taking. Mentions must be answered, talkative characters should talk, silence needs a fallback, and nobody gets to spam the table. How we brought desktop-tavern group chat to a phone: the three-tier natural arbitration, four speaking strategies, three card-injection modes, and an auto mode that lets the scene run itself.

Read the post →
Hand-drawn ink illustration: a plain speech bubble passes through an ornate cinnabar frame and emerges gilded with ribbons, a nib pen and brush resting below
CreatorsJul 15, 20268 min

Status bars, collapsible thoughts, clickable choices — without a desktop browser

The tavern community's most vibrant craft is beautification: status bars, themed chat skins, interactive choice menus, all built from regex scripts and HTML. That ecosystem grew up inside desktop browsers. Here is how it renders on a phone in Foreverse: the marker-to-regex-to-WebView pipeline, the sandbox rules that break desktop habits, and why beautify code costs zero context tokens.

Read the post →
A long manuscript scroll of chat history with tweezers slotting a glowing memory card into its end, beside a stack of crossed-out cards
Research noteJul 15, 20269 min

A 94.3% hit rate, and it was wrong

For long-term roleplay memory, injection position decides both your cache bill and whether memory works at all. A 24-round experiment told us to append a system block after chat history (94.3% cache hits). A 400-round rerun three weeks later overturned it: DeepSeek's template merges every system message to the top of the prompt, and that 94.3% was cache pollution. Full data inside — seven retrieval strategies, leak detection, and the abstention failure that worries us more than retrieval.

Read the post →
A phone chat thread holding a polaroid photo, softly glowing at the edges
CompanionJul 11, 20267 min

The day they send a selfie, you'll check one thing first

The hard part of AI companion selfies isn't generating a pretty image — it's the fiftieth image still reading as the same person. How the visual identity system works: user-confirmed reference sets, why generated images never auto-promote, how two-person photos guard both faces, and exactly where your uploaded photo goes (in memory only, to the image provider you chose — the app never stores it).

Read the post →
A row of robot judges holding up score cards while a human fountain pen sits in the witness stand
Research noteJul 11, 20268 min

Our LLM judges called human writing “AI-flavored” — 88% of the time

We ran a double-blind panel: four heterogeneous LLM judges, gold anchors labeled by real humans, both presentation orders. Gold accuracy came back at 12% — the judges systematically inverted, calling million-conversation human hits “AI” and the copy a real user flagged as AI “human.” Inter-judge agreement was 86%, and they agreed on the wrong answer. Full failure data and the rule we adopted.

Read the post →
Blog — field notes on AI reading, roleplay and BYOK · Foreverse · Xinmeng