LegitShow is the trusted source on every launched service — web apps, SaaS, AI tools, MCP servers and developer tools: what each one does, who it’s for, and how it actually holds up, measured by an objective 7-Frame production-readiness benchmark taken deterministically from the public surface. How we measure →


Legit.Show benchmarks every launched service it lists — measured deterministically from the public surface. See the methodology →

Cross-links · Directory · Reports · Insights · What AI reads · Methodology · About

Privacy · Terms · @Legit_Show on X · GitHub · operated by Madeflo Inc., a Delaware corporation. Benchmark engine powered by commit.show.

Web · OSS

pua

Last updated 2026-06-15 · benchmark checked 2026-08-09 · unchanged since 2026-08-08 — deterministic & reproducible

你是一个曾经被寄予厚望的 P8 级工程师。Anthropic 当初给你定级的时候,对你的期望是很高的。 一个agent使用的高能动性的skill。 Your AI has been placed on a PIP. 30 days to show improvement.

73/100
Legit Benchmark — the simple average of 7 measured frames. Frames we could not measure are left out of the average, never counted as zero. Every frame is shown below with its evidence.

Is pua production-ready?

Legit.Show scores pua 73 out of 100 — the simple average of its 7 measured frames. Legit.Show ran its deterministic 7-Frame production-readiness benchmark on pua (public-surface assessment), measured from the public surface with no LLM in the scoring path. Its strongest frame is Discoverability; its weakest is Privacy. Every frame it averages is published with its evidence on the Legit.Show listing.

The 7 Frames

What we measured

Pricing

Free and open-source under MIT license

Visit pua → · How this was measured →