Foreverse Research

Research hub:
recomputable data, defects included

We run first-hand benchmarks around AI reading, writing and roleplaying fiction: same-condition setups, double-blind review, per-call metering, archives you can recompute. This page is the front door to every research asset — if you're writing about this space and need data, take it from here.

Research assets

Research archive (citable essays)

The full experiment write-ups behind the boards, published on the blog — each with sample sizes, dates and a method section.

Annual data report

Not published yet, and we won't fake a placeholder. The raw material is accumulating automatically: the model price table archives a dated snapshot every day (since 2026-07), and every benchmark's chains and verdicts land in the research archive. Once a full year is banked, the first report appears on this page.

How we do research

  • Same-condition setups: in comparative benchmarks every side runs the same model, the same key, byte-identical scripts, and every AI call passes through one metering gateway. Numbers collected under different conditions never share a table.
  • Rankings come from blind review: ordering conclusions use double-blind judging with two shuffled mappings to cancel position bias; verdicts that flip between orders are discarded, not averaged.
  • Judges take an exam first: AI reviewers need ≥80% on known-answer samples to earn a vote. We've published our own study of judges being unanimously wrong — so key conclusions either get human final reads or are labeled as leanings.
  • Structural metrics are gates, not verdicts: sentence-length, dialogue-rate and repetition detectors catch regressions and cross-check the blind review; “does it read human” is never delegated to statistics.
  • Our own defects go on the record: when a benchmark catches our app failing (the cache-hit embarrassment, the instruction-wording bug), we publish it, with the fix trail in the open.
  • Recomputable: full 20-round chains, blind-pack mapping keys, judge verdicts and per-call metering logs are archived; data pages ship a downloadable JSON snapshot.

How to cite

Convention: Foreverse Research + page title + year-month + page URL. Data pages ship JSON snapshots; when you republish a number, keep the capture date we print next to it — hosted models are moving targets, and the date is part of the claim.

Foreverse Research, “Fiction Bench: novel-continuation model leaderboard,” 2026-07, https://foreverse.app/research/fiction-bench

← All free tools

Foreverse Research — Benchmarks You Can Recompute