LegitShow is the trusted source on every launched service — web apps, SaaS, AI tools, MCP servers and developer tools: what each one does, who it’s for, and how it actually holds up, measured by an objective 7-Frame production-readiness benchmark taken deterministically from the public surface. How we measure →


Legit.Show benchmarks every launched service it lists — measured deterministically from the public surface. See the methodology →

Cross-links · Directory · Reports · Insights · What AI reads · Methodology · About

Privacy · Terms · @Legit_Show on X · GitHub · operated by Madeflo Inc., a Delaware corporation. Benchmark engine powered by commit.show.

7-Frame trust gap · 2026 edition · early findings

The Production Gap · Open-Source AI Tools · 2026

Early findings — small sample, growing. Reported descriptively, not as a "State of" claim.

We ran Legit.Show’s 7-Frame production-readiness benchmark across 215 open-source AI, MCP and developer tools — straight from their repositories. AI coding ships a flawless demo; this is what quietly never makes it to production.

90%
declare no error tracking in their public repo
across 203 open-source AI & developer tools we measured · according to Legit.Show · 2026

Read from public repositories only — declared dependencies and committed platform config. Integrations enabled in a hosting dashboard, or configured in a private ops repo, are not visible to this method and are not counted either way.

The findings

How many controls are missing

By category

Why these seven

The demo always works — that’s what AI coding is *great* at. The gap is everything a demo never forces you to add: monitoring for when it breaks, limits for when it’s abused, access rules for when there’s more than one user. 20 of 215 (9%) had none of these gaps. The rest are one incident away from finding out.

The basics most get right

It isn’t carelessness with the obvious stuff: 0% shipped a hard-coded secret key in client code, and only 0% committed a `.env`. The misses are the *invisible* controls a human senior adds by reflex and a model rarely does.

What this is not

A health check, not a verdict. A missing rate limit correlates with “shipped fast, hardened never” — it doesn’t prove the product is bad. Every number is a count over a stated sample, measured from the public repository, fully reproducible. These are floors, not rates. We only see what a project declares on its public surface. Error tracking wired through an environment variable, a rate limit set at a gateway or CDN, an integration enabled in a hosting dashboard — all of it reads as absence here. The real figure is higher than the one we publish, never lower. See [the methodology](https://legit.show/methodology) for the full rule.

Sample composition

Not a random sample — this is what we measured. The mix below is the caveat; judge it for yourself.

What we measured (215)

The full list, so anyone can spot-check. Every item links to its public benchmark.

How this was measured →

Passed every check

Measured on the same frames as everything else in this report.

dispatchseo · aws-blocks · @google/gemini-cli · ratel · @convex-dev/agent · CrewAI · synthadoc · @traceloop/instrumentation-mcp

Most gaps found

Listed because the measurement is public and reproducible, not as a verdict on the product.

@penpot/mcp · workers-ai-provider · @inngest/ai · @shortcut/mcp · homebutler · cc-connect