The engineering problem.
Changing a guard can improve aggregate recall and still introduce a serious regression in one language or category. A release decision needs comparable runs, inspectable examples, and an explicit policy. GuardMeter is built around that evaluation loop.
A workflow for both engineers and reviewers.
Engineers compare a baseline and candidate on labeled examples. The local app lets reviewers inspect runs, metrics, slices, and individual predictions. A static dashboard can carry the results into a review without requiring a running application.
- Baseline-versus-candidate metrics and confusion matrices
- Category, language, and attack-type analysis
- Threshold and latency views with sample inspection
- A self-contained, read-only dashboard export
A policy that travels with the code.
A versioned configuration defines the thresholds a run must meet. The command-line gate can fail a CI build when those thresholds are breached. JSON, Markdown summaries, and JUnit outputs make the result usable in existing development workflows.
Proof includes the failures.
The published sample evaluation includes a failed gate with its underlying findings. The downloadable report and evidence pack show what was evaluated and how the result was reached. They are sample evidence, not a customer success claim or a certification of safety.
Open tooling, plus a focused service.
Teams can inspect and use the open-source framework themselves. The GuardMeter Tune-Up provides a scoped service around evaluation and actionable findings. Together, they show how a technical tool can support both a self-service developer experience and a concrete B2B offer.
