Reflex-Guard achieves 95.9% harmful-prompt recall at 37.6 ms locally, faster than Llama Guard 2 and SafeDecoding
Synopsis
The work introduces Reflex-Guard, a lightweight locally running guardrail for LLM prompt safety that combines jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers, achieving 95.9% recall on harmful prompts at 37.6 ms end-to-end latency on a strategically balanced dataset of 30,568 samples drawn from five complementary sources, faster than Llama Guard 2 (255 ms) and SafeDecoding (723 ms), detecting 100% of GCG suffix attacks and Base64-encoded prompts at the default threshold, while DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection.
Fig. 1: Overall Reflex-Guard pipeline architecture.
arXivInterpretation
Reflex-Guard reaches 95.9% recall on harmful prompts at 37.6 ms end-to-end latency on a balanced dataset of 30,568 samples. Existing guardrail methods such as LLM-as-a-judge and cloud-based safety APIs typically add about 250-900 ms per request, while real-time applications usually need to respond in under 100 ms; this work cuts latency to 37.6 ms while keeping high recall. Evaluation on a strategically balanced dataset of 30,568 samples drawn from five complementary sources, reporting both recall and end-to-end latency.
Reflex-Guard is faster than Llama Guard 2 (255 ms) and SafeDecoding (723 ms), and reaches a Reflex Efficiency Score (RES) of up to 16.79 versus 11.90 for Llama Guard 2 and 9.80 for SafeDecoding. The work introduces RES as a metric combining efficiency and effectiveness and reports direct comparison values against two baselines. Systematic evaluation reporting latency and RES comparisons against two named baselines.
At the default threshold Reflex-Guard detects 100% of GCG suffix attacks and Base64-encoded prompts, but DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection. The work shows that different attack types occupy distinct regions in the embedding probability space and derives deployment advice for adjusting thresholds by attack type. Detection performance is reported separately for GCG suffix attacks, Base64-encoded prompts, and DrAttack structured prompts under default and adjusted thresholds.
Reflex-Guard runs locally and is composed of jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers. Compared with routing user prompts through external moderation endpoints, local execution avoids the associated data privacy concerns. The component composition is explicitly described and serves as the basis for low-latency, high-accuracy filtering.
Perspective
The result targets real-time LLM application deployments that need to respond within 100 ms and prefer not to send user prompts to external moderation endpoints; the lightweight local design lets such systems perform prompt safety filtering in their own environment. The threshold advice by attack type (for example, a 0.03 threshold for DrAttack structured prompts) can be applied directly to deployment configuration.
Different attack types occupy distinct regions in the embedding probability space, implying a single default threshold may not be optimal for every attack form, as the 0.03 threshold for DrAttack illustrates; readers may watch how this threshold sensitivity behaves across more attack types and real traffic distributions. In addition, the visible text is the abstract and browse context without figures or full per-category results, so the detailed numeric distributions and false-positive behavior per attack type remain open questions.
