A forty-panel board driven by four underlying system states gives its owner the feeling of forty-fold coverage and the reality of four.
Nobody buys the forty panels. They accumulate — one per incident, one per migration, one per person who left. The count goes up and the number of things you can actually see does not, and there is no moment at which anyone is told the difference.
This measures that difference. Point it at a dashboard export and it reports how much independent information your telemetry carries, which panels restate one another, and which are safe to archive — with the evidence for each, and the reasons you should not.
Link three is the one everyone asserts and nobody has tested, including us. It is written down as an unscored prediction in the public repository, and it is the main thing we are looking for teams to help settle.
Runs in your browser. Your file is read in this tab and never uploaded — open the network panel and check.
A real result on the corpus that ships with the tool: a live Prometheus instance, eleven metrics, 437 samples. Small enough to check by hand, which is the point of using it as the example.
Presented as eleven independent views of the system.
# the audit figures redd run prometheus_infra.csv --basis differenced --ordered # the archive candidates, and the conflicts against your rules redd prune prometheus_infra.csv --basis differenced --ordered \ --refs ./monitoring --worksheet ws.csv
Eleven panels is a small board. The ratio is what travels: half of a production Prometheus instance was restating the other half, on a set small enough that its owners could reasonably believe they knew it.
Two of those eleven metrics are statistically interchangeable. Only one of them is safe to touch.
methodGET_status200vsrequest_rate_total
r = 0.99987
request_rate_total carries 0.0011 unique variance.
On the arithmetic alone it is a restatement of the other column and the
engine marks it ARCHIVE. Point the reference scan at your alert
rules and the worksheet leads with this instead:
** CONFLICT: engine says ARCHIVE, but this metric is behind a PAGING alert (RequestRateCollapse) **
Both facts are true, and only one of them is in the CSV. Archive it on the arithmetic alone and the next outage pages nobody. That row is sorted to the top of every worksheet, paging alerts first, because on a two-hundred-panel board the one that matters otherwise sits in the middle looking like all the others.
This is why the tool ends at a worksheet rather than a delete button. The arithmetic is the easy half.
Everything above was measured on public data. Nobody has run this on a board they would be paged for.
So there is nothing to buy here and no demo to book. There is an open call for ten teams to run the same audit on their own dashboards and say what came back — including nothing.
ADOPTION_PREREG.md — unscored, and
committed before any of you ran it. Read it first, then check
whether we were right about you.What it costs you. The example board below takes about two minutes. Your own board takes longer, and the slow part is not the audit — it is getting an export into one column per metric and one row per timestamp, which is not one click in Grafana. We predict that is where most of you will stop, and it is better said here than discovered in a reply.
What you get. Your numbers, the worksheet, and the tool — Apache 2.0, yours to keep and to fork. What we get is the first evidence that any of this survives contact with a board somebody is paged for. That is the entire trade; it is deliberately not a discount on something later.
No NDA needed to send a finding, happy to sign one if you would rather. Your team is not named in the write-up unless you ask to be.
Export any dashboard to CSV. Columns are metrics, rows are timestamps. Nothing is uploaded.
Platform and SRE teams with a board nobody will prune.
If you have two hundred tiles, four pages open during one incident, and a shared understanding that nobody has deleted a panel in two years because whoever deletes the wrong one owns the next incident review — this was built for you.
If your dashboard is ten panels and you know what every one of them does, you do not need this and it will tell you so.
The browser above proves the arithmetic. It cannot show you the part that makes archiving safe.
Apache 2.0 · numpy is the only dependency · nothing is uploaded from the CLI either, and there is a test asserting the code cannot open a socket.
Step 03 is what produced
the CONFLICT above. It answers two of the
worksheet's four attestation columns — monitors and dashboards. SLOs and
runbooks it cannot read, and it says not found in N scanned
sources rather than "unreferenced", because a monitor in a repo it
was never pointed at is invisible to it.
Every check runs on every file — and the ones that did not fire get printed too, along with the one we have not built yet. You can tell what a tool actually covers by what it admits it doesn't.
The real reason nobody prunes a dashboard is that whoever deletes the wrong panel owns the next incident review. So nothing here archives anything on its own.
One guard worth knowing about: when a total and its parts are all on the board, every one of them looks individually redundant. An unguarded tool offers you the whole family. This one protects the parts and offers only the total.
A tool only ever pointed at data whose structure nobody knows can never be caught being wrong. So the predictions go on disk before the data is pulled, and get scored afterwards — misses included. Several of the most useful results were our own bugs.
| Dashboard | Shape | Result | What it found |
|---|---|---|---|
| NYC COVID daily counts | 53 × 455 | 5.1 signals | Raw view says 1.6 — during a pandemic every metric rides the same wave |
| ACT air quality | 12 × 1,094 | 4.7 signals | NO₂ and CO at r = 0.03 but 177× their Gaussian-implied dependence |
| Prometheus infrastructure | 11 × 437 | 5.6 signals | status=200 and request_rate identical at r = 0.99987 |
| FDIC bank call reports | 39 × 1,915 | 2.5 → 14.8 | Raw totals said 2.5. Normalising per-unit revealed 14.8 — size was hiding everything |
What this cannot do, in one place, stated rather than discovered. Each of these is either a published miss or a limit we measured and could not remove.
Average column matched
at r = 1.00000000 on 541 of 541 models and nothing flagged it. Exact
subset sums are detected; means are not.The full record, including the two pre-registered predictions that failed and the cycle whose misses disproved the previous cycle's diagnosis, is in the repository — six studies, four with published misses.
One question: how many distinct things is this dashboard actually measuring? A forty-panel board driven by four underlying system states gives its owner the feeling of forty-fold coverage and the reality of four. That gap is where duplicate alerts come from, and it is what this measures.
It was built by someone who got tired of dashboards nobody trusted and panels nobody would delete, and who wanted the argument settled with evidence instead of opinion.
Before each dataset is pulled, what we expect to find gets written down
and committed. A tool only ever pointed at data whose structure nobody
knows can never be caught being wrong — so we make it possible to catch.
Every dataset in REAL_DASHBOARDS.md carries its predictions
and the score, misses included. Several of the most valuable results were
misses, because they turned out to be bugs in this tool rather than bad
guesses about the data.
You do not have to take that on trust. The source, the eight pre-registration documents — six scored, two still open — and every corpus they were scored against are public — read them on GitHub. "We wrote the predictions first" is an empty claim on a marketing page; in commit history it is checkable, which is the only reason it is worth saying.
Seven boundaries, each either a published miss or a limit that was measured and could not be removed, are set out in full under engine boundaries above — rolling averages, subset means, counter-data leniency, cross-sections, causality, and semantics.
Redd Munro is not a person. Both halves are Scots, and both describe the tool.
Redd means to clear out and set in order — you redd up a room before you can see what is in it. A Munro is a Scottish peak over 3,000 feet, and the genuinely hard part of Sir Hugh Munro's 1891 tables was never the measuring. It was deciding which summits count as separate mountains and which are just subsidiary tops of the same one. People have argued about it ever since.
That is this tool in one sentence. Not every summit is a separate mountain, and not every panel is a separate signal.
Built by Shaun Cooper, independent systems theorist. The engine grew out of a research programme that spent most of its effort trying to break its own measurements, and catalogued the distinct ways a comparison like this can produce a publishable-looking wrong answer.
Runs entirely in your browser via Pyodide. No account, no upload, no telemetry — which would be an awkward thing for this particular tool to collect.
Every "keep it just in case" argument in this industry rests on a belief that has never been tested.
The belief is that metrics which look identical come apart during incidents — that the pair sitting at r = 0.99987 all quarter diverges exactly when it matters, which is why deleting either one feels dangerous. It is the load-bearing assumption under every dashboard nobody prunes.
It has never been tested. Not here, not anywhere. It is
written down as prediction H1 in the public repository,
deliberately unscored, because settling it needs something no public
dataset contains: a window somebody labelled incident
because their service was genuinely broken.
If you have that window, you can settle a question this whole industry has been assuming the answer to.
And if you ran it and it found nothing, that is worth an email too — a tool that only hears from the cases where it worked learns nothing.