Developer Tools
Claude
Last updated 2026-07-25 · benchmark measured 2026-08-13 — deterministic & reproducible
A guide to reproducing Claude's agentic search benchmark scores via the Messages API.
84/100
Legit Benchmark — the simple average of 7 measured frames. Frames we could not measure are left out of the average, never counted as zero. Every frame is shown below with its evidence.
Is Claude production-ready?
Legit.Show scores Claude 84 out of 100 — the simple average of its 7 measured frames. Legit.Show ran its deterministic 7-Frame production-readiness benchmark on Claude (public-surface assessment), measured from the public surface with no LLM in the scoring path. Its strongest frame is Privacy; its weakest is Performance. Every frame it averages is published with its evidence on the Legit.Show listing.
The 7 Frames
- Performance — 38/100
- Accessibility — 89/100
- Security — 90/100
- Privacy — 100/100
- Reliability — 92/100
- Standards — 92/100
- Discoverability — 85/100
What we measured
- Security headers present: CSP, HSTS, X-Frame-Options, X-Content-Type-Options.
- Served over HTTPS with a valid certificate.
- Real Lighthouse performance run — 301 ms to first byte.
- Returns a proper 404 for unknown routes.
- 0 of 0 sampled routes reachable.
- Has a reachable privacy policy.
- Sets cookies / loads scripts with no consent prompt.
- Discoverable: sitemap, OpenGraph image, canonical URL.
Who it's for
AI researchers · Developers · ML engineers · Benchmark evaluators
Visit Claude → · Alternatives to Claude → · How this was measured →