Foreverse Research
Research hub:
recomputable data, defects included
We run first-hand benchmarks around AI reading, writing and roleplaying fiction: same-condition setups, double-blind review, per-call metering, archives you can recompute. This page is the front door to every research asset — if you're writing about this space and need data, take it from here.
Research assets
Research archive (citable essays)
The full experiment write-ups behind the boards, published on the blog — each with sample sizes, dates and a method section.
Annual data report
Not published yet, and we won't fake a placeholder. The raw material is accumulating automatically: the model price table archives a dated snapshot every day (since 2026-07), and every benchmark's chains and verdicts land in the research archive. Once a full year is banked, the first report appears on this page.
How we do research
- Same-condition setups: in comparative benchmarks every side runs the same model, the same key, byte-identical scripts, and every AI call passes through one metering gateway. Numbers collected under different conditions never share a table.
- Rankings come from blind review: ordering conclusions use double-blind judging with two shuffled mappings to cancel position bias; verdicts that flip between orders are discarded, not averaged.
- Judges take an exam first: AI reviewers need ≥80% on known-answer samples to earn a vote. We've published our own study of judges being unanimously wrong — so key conclusions either get human final reads or are labeled as leanings.
- Structural metrics are gates, not verdicts: sentence-length, dialogue-rate and repetition detectors catch regressions and cross-check the blind review; “does it read human” is never delegated to statistics.
- Our own defects go on the record: when a benchmark catches our app failing (the cache-hit embarrassment, the instruction-wording bug), we publish it, with the fix trail in the open.
- Recomputable: full 20-round chains, blind-pack mapping keys, judge verdicts and per-call metering logs are archived; data pages ship a downloadable JSON snapshot.
How to cite
Convention: Foreverse Research + page title + year-month + page URL. Data pages ship JSON snapshots; when you republish a number, keep the capture date we print next to it — hosted models are moving targets, and the date is part of the claim.
Foreverse Research, “Fiction Bench: novel-continuation model leaderboard,” 2026-07, https://foreverse.app/research/fiction-bench