Comuvia mark — a circle in motion Comuvia ForeGlass™
ComuviaComuvia ForeGlass™

Verification Explorer

Every forecast below was sealed with a cryptographic hash before the fact it predicts was knowable, then left alone. This page shows what we predicted, when we sealed it, and — once the due dates arrive — how it scored. Hits and misses are published identically.

sealed ledger ● 0 shown
Lanei Claimi Predictedi Forecasteri Resolvesi Againsti

How to read a row: each row is one sealed forecast — append-only, hash-audited, never edited after sealing. Claim is what was committed, Against is the public source that will settle it, and rows sharing a question are competing forecasters on the same question. Hover any i for the full text; click a row for the sealed record, the seal hash and the chart against the real series.

One question, many forecasters

🛈 What this measures: how well each LLM predicts this metric using a simple, uniform approach - the neutral question plus (in grounded arms) public context, the same prompt for every model, no prompt tuning or retries. Forecast quality depends on the prompt, so a better-engineered one would likely score better; this is a like-for-like comparison of models, not the best achievable LLM forecast.

ForeGlass LLM Ensemble LLM median

Parts that must add up to the whole official published statistics

In plain terms: a forecast about a part should be consistent with the whole it belongs to. Each chain takes an officially published total, splits it into the pieces the statistics agency itself publishes, adds them back up, and shows how far off the round trip lands. A rung marked ◆ carries one of our sealed forecasts. Each chain is a group: open the one you want. The per capita switch redraws the ladders per person or per household where the chain has a stated denominator - that view is a ratio, so the shares, the running total and the residual stay on the totals.

About this tab

These are measured, published values reconciling with each other - the frame our sealed forecasts sit inside.

Where the parent total is independently compiled (the World Bank's EU aggregate), landing on it checks our member set, joins and arithmetic against an outside compiler. Where the parent is only the sum of its own published bands, exact agreement is the expected result - it verifies arithmetic, not two independent measurements.

What this tab does not show
  • Closure is not accuracy. Almost none of these rungs carry a forecast yet; coverage is stated per chain as a number.
  • Exact agreement where both sides rest on the same underlying series is expected, not corroboration.
  • A chain that cannot be built is shown as a refusal naming the blocker.
  • Per-capita values are ratios on a stated denominator - they do not add up, and where the denominator comes from a different statistical universe than the numerator the chain says so.

What this page is, and what it is not

Forecasts are extracted from committed simulation runs and external model panels, then sealed: the forecast, its due date and its resolution rule are hashed together and appended to a ledger that is never rewritten. When the due date arrives the published statistic is read from the named public source and the forecast is scored — whatever it says.

  1. Seal before witnessing. A forecast enters the ledger with its hash while the outcome is still unknown. That ordering is the whole claim; everything else on this page exists to let you check it.
  2. How a resolved lane is scored. A lane is a hit if the published value falls inside the sealed interval (or matches the sealed point rule). Alongside that we publish skill against a naive forecast: skill = 1 − |errmodel| / |errnaive| — above 0 beats the naive anchor, 0 ties it, below 0 loses to it. The anchor itself ships with every lane as naive_value, so the arithmetic is yours to redo: err = |actual − forecast| for each, then the ratio. For the first grading cohort the anchor is the year-ago committed share (2025-Q1) — the same statistic one year earlier, on the print committed at seal time. Publishing a score without the number it is measured against would be a number you could not check, which is the opposite of the point.
  3. Which print counts, and when we are late. A lane grades against the first print as fetched on the grading run. Statistical agencies revise; a later revision never reopens a resolved outcome, in either direction. If a needed release has not published by the due date the lane resolves when the data lands — never voided, never estimated from the prior quarter — and it says so itself: a late lane carries late_by_design and the observed release_observed date, so a resolved_at later than its due date is legible as a slow source rather than as us backfilling.
  4. Resolution rule fixed at seal time. Each lane names the public series and the exact computation that will grade it. The rule cannot be edited afterwards — it is inside the hash.
  5. Misses publish like hits. There is no editorial step between grading and publication. A miss is a recalibration input, not an embarrassment to bury.
  6. Competing forecasters, one question. Where several forecasters answer the same question, all of their lanes are sealed together, so the comparison cannot be assembled after the fact from the winners.
  7. Held forecasts are still committed. Forecasts we are not publishing yet appear on the Verify tab as hash-commitments only — existence and immutability now, content on release.

What this page deliberately does not show: how the simulations work inside — model internals, calibrations, run lineage. The claim on offer is the scored track record, not the recipe; those are separable, and only one of them needs to be public for you to check us.

Verify it yourself

Alongside this page sits ledger.json — the machine-readable ledger behind everything shown here.

  1. Seal algorithm:
  2. What you can check here:
  3. Commitments. ledger.json carries a commitments array: one {id, hash, sealed date} row per held forecast. When a held forecast is released, its hash on this page must equal the hash committed earlier — a forecast rewritten in between cannot match.
  4. Time-witness it. Snapshot ledger.json at an independent web archive. That timestamp is outside our control, which is what makes the ordering claim checkable by someone who does not trust us.
Grading status:

Honest limit: two of the fields inside each hash are internal run identifiers, so this ledger does not let you recompute a hash from scratch — it lets you check that a published hash never changed. Independent archive snapshots are what turn that into a time-anchored claim.