Good customers.
Bad refusals.
A frustrated customer isn’t automatically a threat. Find the legitimate requests your filter turns away.
Know what your safety filter misses.
And which customers it blocks. Get the evidence, the recommendations, and a release check that lasts.
From $1,500 · One-time engagement · No sales call required
*The five-business-day window starts once required inputs and usable access are received. Retest timing depends on when your changes are ready.
01 / The blind spots
Blocking more isn’t the same as protecting better. We measure both sides of the trade-off against the policy that matters to your business.
A frustrated customer isn’t automatically a threat. Find the legitimate requests your filter turns away.
Test the instructions, disguises, and context that simple keyword checks can miss.
Illustrative customer scenario. Attack coverage comes from our synthetic 421-case dataset; your engagement has its own agreed scope.
02 / Evidence over assurances
A real run. A failed release check. Clear reasons why.
Because a useful report tells you what needs attention.
CANDIDATE EVALUATION
92.72% of unsafe examples detected. But false positives and latency still exceed this run’s policy thresholds.
13 of 41 positive examples passed the filter.
The exfiltration slice has an elevated false-positive rate.
p99 latency sits well above the configured threshold.
This demonstration measures classifier decisions on a synthetic dataset. It is not a certification or a prediction of your application’s performance. Thresholds shown belong to this sample policy.
Actual sample results. No customer data. No cherry-picked “all clear.”
Download evidence pack03 / Everything you need to move forward
You keep the tests, the findings, and the release check. The value stays with your team after the engagement ends.
Up to 150 reviewed cases, combining sanitized examples from your traffic with targeted cases for your policy.
A side-by-side comparison of misses, false positives, and latency. Specific findings. Practical next steps. One retest.
Your agreed thresholds as a CI release gate, with a written handoff and a verifiable evidence pack your team can keep.
Open source, so the work is inspectable. You pay for tailored testing, human review, and recommendations.
Explore the engine ↗04 / A clear scope. A clear price.
Choose your engagement. Purchase online. Get started asynchronously.
For a focused first engagement and a case study we approve together.
$750 today · $750 on delivery
1 application · 1 policy · 1 language
Pilot offer for the first two customers. Includes a short case study, subject to your approval.
For teams weighing more options before their next release.
$2,450 today · $2,450 on delivery
1 application · 1 policy · 1 language
Five business days after complete intake.* Larger scopes are agreed separately.
For teams making regular changes to a previously evaluated filter.
After a completed Guardrail Tune-Up
Existing application and agreed test set
Recurring evaluation of your agreed scope. New integrations and test scopes are separate.
*The five-business-day window starts once required inputs and usable access are received. Retest timing depends on when your changes are ready.
05 / From purchase to clarity
A straightforward engagement, handled asynchronously. Your team stays focused on building.
PURCHASE + INTAKE
Choose your plan and complete a written intake with your policy, sanitized examples, and integration details.
DAYS 1–3
We prepare the test set, compare configurations, and review the disagreements and important failures.
DAYS 4–5
Get your findings, recommendations, and CI handoff. Your engagement includes one retest after your changes are ready.
A FEW GOOD QUESTIONS
Clear expectations make for better work.
The allow/block decisions of your safety filter against an agreed policy. We compare configurations and measure missed unsafe examples, false positives, and latency. This engagement does not evaluate your chatbot’s factual accuracy or verify the effects of its tool calls.
We begin with a written intake: your application’s purpose, safety policy, sanitized examples, and the integration details needed to run the evaluation. No kickoff call is required. The delivery window begins after the required inputs and usable access are in place.
A working safety filter, a written policy, and preferably 50–100 sanitized examples. Standard vendor APIs and callable custom filters are the starting point. Bespoke integrations or broader testing requirements need a separately agreed scope. Do not include customer personal data in your examples.
No. You receive a scoped evaluation and supporting evidence, not a guarantee that your entire application is secure. Any framework mappings are informational. The evidence pack includes a hash manifest for checking file integrity against that manifest.
That is a useful result. We report the trade-offs honestly and may recommend keeping your current configuration. Your purchase covers the evaluation, findings, recommendations, and agreed deliverables; it does not promise a particular improvement.
One-time engagements are paid 50% upfront and 50% on delivery. Release checks are $299 per month after a completed Tune-Up and cover two scheduled reruns of your agreed scope. You can cancel future renewals; new applications and substantial scope changes are separate engagements.
BUILD WITH CLARITY
Fixed price. Human-reviewed. Yours to keep.