AI & Agents
LangWatch
Last updated 2026-07-30 · benchmark measured 2026-08-09 — deterministic & reproducible
LangWatch tracks token usage, cost, and traces for Claude Code and other coding agents.
Is LangWatch production-ready?
Legit.Show scores LangWatch 67 out of 100 — the simple average of its 6 measured frames. Legit.Show ran its deterministic 7-Frame production-readiness benchmark on LangWatch (public-surface assessment), measured from the public surface with no LLM in the scoring path. Its strongest frame is Discoverability; its weakest is Privacy. 6 of the seven frames returned a score; Accessibility was not measurable on this service and is recorded as null — not as zero. Every frame it averages is published with its evidence on the Legit.Show listing.
The 7 Frames
- Performance — 80/100
- Accessibility — not measurable on this service (null — not scored as 0)
- Security — 45/100
- Privacy — 25/100
- Reliability — 75/100
- Standards — 75/100
- Discoverability — 100/100
What we measured
- Security headers present: X-Content-Type-Options, Referrer-Policy.
- No Content-Security-Policy and no HSTS.
- Served over HTTPS with a valid certificate.
- Real Lighthouse performance run — 188 ms to first byte.
- 3 of 3 sampled routes reachable.
- No privacy policy found.
- Sets cookies / loads scripts with no consent prompt.
- Discoverable: structured data, sitemap, OpenGraph image, canonical URL.
Who built it
Manouk Draisma (@alexforbesreed) · product account @LangWatchAI
Who it's for
AI engineers · development teams · agent builders · production teams
Pricing
Free tier (50k events/month), Growth plan €29/core-seat/month + €5 per 100k events, Enterprise custom pricing
Visit LangWatch → · Alternatives to LangWatch → · How this was measured →