Claims & corrections.

Every quantitative claim GhostLine publishes traces to a typed record in an internal research ledger. This page lists the live figures with their scope and provenance, and the corrections and retractions that preceded them. Corrections stay listed after they are fixed. The admission policy: only claims that have cleared confound review are published; anything still under audit is labelled where it appears.

FigureWhat it saysScope, provenance & statusRecord
91.3% Task category read from the prompt encoding alone — scoped to the five prompt-intent categories the live display renders. Qwen3-8B · 5 prompt-intent categories · 335 aggregated (non-per-head) prompt features · 900 samples over 300 prompts · prompt-grouped CV. Mean over five fold seeds: 91.3% ± 0.8 (single-seed range 90.3–92.7); macro-F1 0.911. Text baseline on identical folds: 66.4%. Sampling uncertainty at 300 prompt groups is roughly ±3 points; a cluster-bootstrap interval is queued. Distinct from the 3,791-feature figure below — the shared numeral is coincidence. CLM-2026-0731-001
98.4 / 95.4 / 94.0 / 84.5 / 78.0 Per-category recall of the headline classifier: creativity, retrieval, reasoning, uncertainty, precision (seed-mean). Supports 150 / 300 / 180 / 120 / 150, n = 900. Support-weighted these values reconcile to 91.3%, the headline. Precision is the weakest category. CLM-2026-0731-001
88.4% The seven-label variant: the five categories plus collapse and the residual "edge cases" bucket. Qwen3-8B · 1,134 samples over 378 prompts · 335 aggregated prompt features · prompt-grouped CV. Re-derived at 88.36% in a July 2026 internal audit (pipeline agreement — the audit's seed and fold assignment are not on record) and again at 88.36% from a fresh implementation on shared folds. Macro-F1 0.834. Per-label recalls 95.0 / 94.0 / 93.3 / 89.0 / 88.0 / 77.5 / 55.6, supports 180 / 150 / 180 / 300 / 150 / 120 / 54 — support-weighted they reconcile to 88.36%; excluding the residual bucket would read 90.0%. The bucket is not a claimed category and is never surfaced in the live display. CLM-2026-0729-001
85.0% The same prompt-time task, replicated on a different architecture at a different scale. Llama 3.2 3B · independent corpus · 1,380 samples over 530 prompts · prompt-grouped CV. This is a corrected figure: originally 99.6% under sample-level CV; prompt-identity leakage was removed by grouped splits in the February 2026 overfitting audit. Unlike the 8B figure, it has no 2026-07 re-run on record — its support is that audit. CLM-2026-0729-002
91.3% The 3,791-feature seven-label variant (secondary figure; shares the headline's numeral by coincidence). Reproduced at 90.74%, within its own fold variance (SD 2.17pp). Exact hyperparameters and seed were never recorded, so it reproduces within noise but not exactly. Cited only as secondary for that reason. COR-2026-0726-006
66.4% What prompt text alone achieves on the headline's own folds — the baseline any geometric result must beat. Qwen3-8B · five categories · 900 samples over 300 prompt groups · identical folds to the headline · TF-IDF + logistic regression, fitted in-fold · mean over five fold seeds. The earlier 74.6% (Feb 2026) fitted its vectorizer on the full corpus before cross-validation and ran on a 780-sample matched subset; retained as the historical configuration. An earlier "+13pp" delta was removed because it paired mismatched numbers — see corrections below. CLM-2026-0731-001
d ≈ 4.2 MLP activations become a distinct discriminating signal family at 8B, barely present at 3B. Qwen3-8B, February 2026 protocol; effect size is standardized within-model. Measured before the current confound-screening standards existed. The queued re-audit tests three named risks: generation-length confounds, scale-versus-family attribution (the 3B and 8B comparators differ in model family, not only size), and robustness to per-dimension normalization. re-audit queued
~15 ms / <2 ms / 0 Pre-generation gate latency, per-token classification latency, added forward passes. Capability measurements on a single-stream reference stack, explicitly not measured under production batching. Not classifier metrics. labelled on page
58 · 5 · 9 Signals captured per token; architecture families; models instrumented (1B–8B). Capability and breadth counts, not one uniform validation — the runs differ in corpus and protocol. Framed as a breadth inventory wherever they appear. inventory

Record IDs refer to GhostLine's internal research ledger — a typed store of claim, correction, and method records with supersession tracking. A build-time lint checks the published pages against this register; a figure that drifts from its record fails the build.

Corrections and retractions.

Claims that did not survive review — what changed, when, and why, one entry each.

2026-02

A 99.6% prompt-time accuracy on Llama 3.2 3B was corrected to 85.0%. Sample-level cross-validation on a multi-run corpus had leaked prompt identity; prompt-grouped splits are now mandatory for every published figure.

2026-07

A sealed-vault hallucination macro-F1 figure was pulled. The measurement was real; the corpus family it was measured on was later disqualified as off-distribution for the deployed setting.

2026-07

A 94-point accuracy figure was pulled after a length-only baseline explained most of its separation. Length baselines now run before any structural claim is published.

2026-07

An "85–91%" accuracy range was pulled for mixing metrics across models and mis-captioning an 8B figure. Single numbers with stated scope replaced it.

2026-07

Binary hallucination-detection F1 headlines were retired from public use in favour of three-class evaluation, which better reflects the deployed decision. The binary figures remain in the internal ledger; they are no longer quoted.

2026-07

An unsupervised-clustering ARI figure was retracted before publication.

2026-07

A "+13pp over baseline" delta was removed: the paired numbers came from different label sets and split instantiations. A matched comparison is queued.

2026-07-30

"Independently reproduced" was removed from the hero banner: the reproduction was an internal re-derivation of the pipeline, not third-party replication. Provenance for the figure now lives in this register.

2026-07-31

The text-baseline row's substrate was corrected to the five-category healthy subset of the same Qwen3-8B corpus; an intermediate revision had misattributed it to a 3B corpus.

2026-07-31

The headline was rescoped from the seven-label evaluation (88.4%) to the five prompt-intent categories the live display renders (91.3%, mean over five fold seeds). The seven-label figure remains published above. Rationale: at 8B the collapse label marks prompt intent rather than confirmed behaviour, and the residual bucket is not a category.

2026-07-31

The 74.6% text baseline was superseded by a same-fold, in-fold-fitted measurement of 66.4%; the original fitted its vectorizer before cross-validation.

How a number earns a place here.

A figure is publishable when it has a ledger record with stated scope; held-out evaluation under prompt-grouped splits at the locked generation protocol; a length-only baseline reported before any structural claim; and no open correction against its corpus or method. The site build runs a claims lint that fails on any published number missing from this register, on retracted figures, and on known-misleading phrasings. External review is invited; corrections reach the address above.