Observability Scoring Engine

This tool has been withdrawn. A review on 23 August 2026 found four separate defects. Any score, predicted accuracy, regime label or exported CSV it produced should not be used or cited.

1. The reference data came from a dataset we had already set aside

The sixteen rows the tool displayed under the heading “Comparable Empirical Domains”, and the curve fitted through them, came from a compilation that our own canonical statistics file lists among its scope exclusions as a dead artifact. The tool went on presenting it as an empirical reference set.

2. Six of those sixteen rows are unsupported, and three were fabricated

A per-domain audit we carried out on 20 May 2026 found six of the sixteen values unsupported or invalid. Three of them — a comparison of Māori maramataka knowledge across shellfish, eel and snapper prediction, shown here as 95%, 75% and 52% — were recorded in that audit as having “no derivation, no source, and no published basis”, and the comparison itself was described in our own notes as fabricated. No published source contains it.

Those values remained on this page for fifteen months after we documented that they were fabricated. The audit and the tool were never connected. We are stating this in our own words because it is our own finding about our own data, and putting it more gently now would compound the original failure. The remaining rows attribute accuracy percentages to Gunditjmara, Martu, San, Aboriginal Australian and Andean knowledge; those are drawn from cited literature, but they belong to the same withdrawn compilation and should not be relied on either.

3. The accuracy calculation was wrong

The logistic curve was implemented as b + L/(1+e) where our own reference implementation uses b + (L−b)/(1+e). Every predicted-accuracy figure this tool displayed was therefore about nineteen percentage points too low and could never exceed 72.9%, and ten of the sixteen points plotted on the chart sat above the curve labelled as their own fit.

4. Two displayed values had no source at all

The thresholds that produced the “Phase Regime” badge, and the ±15% band shown beside each accuracy figure as though it were a confidence interval, exist nowhere in any of our repositories, including a full search of pre-rewrite history. They appear to have been written directly into the page.

If you used this tool

Please discard any score or predicted accuracy taken from it, including exported CSVs. If that output went into unpublished work, the scoring needs redoing by another method. If it went into work already submitted or published, contact us and we will help you correct it.

The underlying cross-domain accuracy measure this tool scored against has itself been retired and has no replacement, so this page will not be restored in its previous form.

Other tools →   How our work is checked →