Your Detector's Threshold is a Benign-Only Quantity
An in-depth look at mathematical proofs and 726 benchmark samples showing why guardrail thresholds belong to benign traffic, not attacks.

Stock photo for illustration only, not from the actual event
- Guardrail thresholds are not model parameters; they are traffic properties.
- Calibrating thresholds using real attack data is a fundamental category error.
- Benign score distributions exclusively dictate optimal thresholds and operating points.
- Shifting production traffic can render quarterly threshold calibrations off by 5x.
Software developers often treat a guardrail's threshold as a model parameter, but a simple mathematical proof reveals it is actually a property of your traffic. Consequently, the datasets most engineers use to calibrate these thresholds are often completely misguided.
Measurements on a public benchmark containing 629 real prompt-injection attacks—embedded inside ordinary tool outputs like bills, emails, and web pages—alongside 97 benign tool outputs across nine open-source detectors, highlight a massive scale discrepancy spanning five orders of magnitude.
For instance, Prompt Guard 2 scored attacks around 0.009 and benign traffic near 0.0008, while its decision cutoff sat at 0.5—roughly 50 times above the model's entire operational range. This configuration caught a mere 6 out of 629 attacks (1.0%) and never triggered on benign data. Conversely, other detectors scored benign traffic near ~0.999, placing a 0.5 cutoff entirely below their functional range.
This wide span of score scales illustrates that a default threshold is merely an assumption about a scale your model may not share. Relying on default settings without accounting for actual traffic distributions inevitably leads to extremes where a system either catches nothing or triggers false alarms constantly.
Assuming a target false-alarm budget f, achievable operating points are set strictly by the benign score distribution. The maximum-TPR threshold corresponds to the k-th highest benign score, where k = floor(f · n_benign). While attack labels help decide how much budget you are willing to pay for, benign traffic alone dictates where the threshold sits.

Stock photo for illustration only, not from the actual event
Sweeping an attack-labelled threshold to maximize TPR subject to each detector's benign-derived FPR reproduces the benign-derived TPR exactly, down to the decimal, across all nine detectors. Treating attack traffic as the primary tuning knob is a category error; attacks measure the payoff, but they do not move the threshold.
The underlying trap is revealed when splitting the benchmark into four domain suites, calibrating the threshold on three at a 2% false-alarm budget, and applying it to the fourth. The threshold breaks the budget on 11 of 36 folds, yielding a held-out false-alarm rate of 4.9%—2.5 times higher than promised—because the traffic changed while the detector remained static.
Source: Dev.to
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment