Reflex-Guard cuts prompt-safety filtering to 37.6 ms locally with dense semantic embeddings, reaching 95.9% harmful-prompt recall on 30,568 samples
Synopsis
The work introduces Reflex-Guard, a lightweight locally running guardrail for LLM prompt safety that combines jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers, achieving 95.9% recall on harmful prompts at 37.6 ms end-to-end latency on a balanced dataset of 30,568 samples drawn from five complementary sources, faster than Llama Guard 2 (255 ms) and SafeDecoding (723 ms), detecting 100% of GCG suffix attacks and Base64-encoded prompts at the default threshold, while DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection.
Fig. 1: Overall Reflex-Guard pipeline architecture.
arXivInterpretation
Reflex-Guard reaches 95.9% recall on harmful prompts at 37.6 ms end-to-end latency on a balanced dataset of 30,568 samples. Existing guardrails such as LLM-as-a-judge and cloud-based safety APIs typically add about 250-900 ms per request, while real-time applications often need to respond in under 100 ms; this work pushes latency well below that bar. Systematic evaluation on a strategically balanced dataset drawn from five complementary sources, with direct comparison to Llama Guard 2 (255 ms) and SafeDecoding (723 ms).
It detects 100% of GCG suffix attacks and Base64-encoded prompts at the default threshold. This indicates that jailbreak-aware preprocessing plus compact sentence-transformer embeddings give stable coverage of these two attack forms. Detection rates for these two attack types are reported as 100% on the evaluation dataset.
DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection because they produced a distinct probability distribution. This reveals that different attack types occupy distinct regions in the embedding probability space, so a single default threshold may not fit every attack form. Observed through threshold-adjustment experiments showing DrAttack's probability distribution differs from other attack types.
Reflex-Guard achieves Reflex Efficiency Score (RES) scores up to 16.79, above Llama Guard 2's 11.90 and SafeDecoding's 9.80. This composite efficiency metric compares accuracy and latency together, showing the local lightweight approach leads clearly on efficiency. RES values reported against the two baselines under the same evaluation setup.
Perspective
The result targets real-time LLM applications that need to respond within 100 ms and settings where routing user prompts to external moderation endpoints is undesirable; the combination of local execution, compact embeddings, and seven binary classifiers makes this guardrail deployable locally. For GCG suffix attacks and Base64-encoded prompts, the default threshold suffices; for DrAttack structured prompts, the threshold must be lowered to 0.03. The deployment advice in the text centers on how different attack types are distributed in the embedding probability space.
Different attack types occupying distinct regions in the embedding probability space means threshold choice is tied to attack form, with DrAttack needing the lower 0.03 threshold as an example; how this holds across a broader, continuously evolving attack surface remains an open question. In addition, the current text is abstract-level information without figures or per-category detailed results, so the specific distributions and false-positive behavior of each attack type cannot be further confirmed from the available material.
