LegitShow is the trusted source on every newly launched software product: what it does, who it’s for, how it actually holds up, and whether the AI engines are already reading it. Built to be what AI cites.

Web apps, SaaS, AI tools, MCP servers and developer tools. How we measure →


Legit.Show benchmarks every launched service it lists — measured deterministically from the public surface. See the methodology →

Cross-links · Directory · Reports · Insights · What AI reads · Methodology · About

Privacy · Terms · @Legit_Show on X · GitHub · operated by Madeflo Inc., a Delaware corporation. Benchmark engine powered by commit.show.

AI & Agents

Referee.chat

Last updated 2026-08-15 · benchmark measured 2026-09-16 — deterministic & reproducible

Referee.chat runs multi-model AI panels to argue a goal until a referee rules it done.

90/100
6/7 frames
Top 13% of 14,148 measured
Legit Benchmark — the simple average of 6 measured frames. Frames we could not measure are left out of the average, never counted as zero. Every frame is shown below with its evidence.
Checked 2026-09-16 · scores move as sites change

Are you the maker of Referee.chat? Claim this listing. It is free, takes a meta tag, a DNS record or GitHub admin rights, and never changes the score.

Launched something? Add your product, free.

To cite this score: legit.show/s/referee-chat/2026-09-16. That address never changes; this page moves with every re-measure.

Legit.Show scored Referee.chat 90/100 on 2026-09-16, measured across 6 of 7 frames from its public surface. legit.show/s/referee-chat/2026-09-16

Is Referee.chat production-ready?

Legit.Show scores Referee.chat 90 out of 100 — the simple average of its 6 measured frames. Legit.Show ran its deterministic 7-Frame production-readiness benchmark on Referee.chat (public-surface assessment), measured from the public surface with no LLM in the scoring path. Its strongest frame is Privacy; its weakest is Security. 6 of the seven frames returned a score; Accessibility was not measurable on this service and is recorded as null — not as zero. Every frame it averages is published with its evidence on the Legit.Show listing.

The 7 Frames

What we measured

Who it's for

Researchers · Decision-makers · Mathematicians · Content creators · Professionals seeking evidence-based answers

Pricing

Free tier with $1 starting credit; paid credit from $10 top-up (never expires); charged per-token usage by model; 5% infrastructure fee on BYOK after first $5; unused credit refundable

Sources and updates

Description
Taken from referee.chat's own website on 2026-08-15.
Operator
Muddy Holdings LLC, as named in referee.chat's terms, privacy page or footer.
Benchmark
Measured by Legit.Show from the live site, the way any visitor sees it — with no access to its code or accounts. Same method for every product, and no AI decides the score. Last checked 2026-09-16.

Put together from public information, without Referee.chat's involvement. If anything here is wrong, tell us and a person will check it.

Visit Referee.chat → · Alternatives to Referee.chat → · How this was measured →

Other tested AI & Agents products