Skip to main content
Back to timeline
arXivSource publication:

LTBD trains four learnable trust-boundary delimiters and cuts prompt-injection success to 0.00% on AlpacaFarm across four open models

Synopsis

The authors propose LTBD, a defense that keeps the LLM fully frozen and trains only four learnable delimiter embeddings placed at the boundaries between trusted instructions and untrusted external data, explicitly encoding trust provenance; across Llama3-8B, Llama3.1-8B, Falcon3-7B and Qwen2.5-7B it reaches 0.00% attack success rate on AlpacaFarm and 0.11-0.19% on TaskTracker, and stays near-unchanged under the defense-aware Delimiter-Spoof adaptive attack.

Source-provided article image: LTBD: Learnable Trust-Boundary Delimiters for Prompt Injection Defense
Figure 1 ·

Figure 1: Overview of Learnable Trust-Boundary Delimiters (LTBD) . (a) Prompt injections embedded in external data steer the LLM away from the original task. (b) LTBD insert four learnable delimiters to explicitly mark the trusted instruction region and the untrusted data region. The delimiter delimiter embeddings are the only trainable parameters, while the base LLM f θ f_{\theta} remains frozen. During inference the input is wrapped with the learned delimiters, enabling the model to follow the trusted instruction and treat the (potentially attacked) external data as data rather than a new instruction.

arXiv

Interpretation

The paper reframes prompt injection as a trust-boundary modeling problem: trusted user instructions and untrusted external data share one input sequence yet carry different behavioral authority, and current LLMs lack an explicit representation of trust provenance. Unlike handcrafted prompting defenses such as Reminder and Sandwich, and unlike training-based defenses such as StruQ that fine-tune model parameters, this work makes the boundary itself the learnable object rather than relying on prompt wording or weight updates. The framing is supported by the method design and ablation: on Llama3.1-8B, moving the same four learnable tokens from the input prefix to the trust boundaries reduces AlpacaFarm ASR from 23.08% to 0.48% and SEP ASR from 33.81% to 3.43%, indicating the gain comes from boundary placement rather than from adding trainable tokens.

LTBD optimizes only four delimiter embeddings while the base LLM stays fully frozen, requiring no auxiliary defense model and no modification of model weights. Unlike DefensiveToken, which learns global defensive token embeddings, each LTBD delimiter is tied to a specific structural boundary, so the learned embeddings acquire trust-aware roles rather than acting as generic control prefixes. Training uses the Cleaned Alpaca instruction-tuning dataset (51K samples) for a single epoch, keeping half the samples unchanged and injecting one of three variants (Ignore, Completion, and the proposed Delimiter-Spoof) into the rest, with the same data construction protocol across four models.

On the security-utility trade-off, LTBD reaches 0.00% ASR on AlpacaFarm for all four models and 0.11-0.19% on TaskTracker, while achieving the highest AlpacaFarm WinRate among all defenses on all four models. Relative to DefensiveToken, SEP ASR drops from 3.20/2.81/6.70/4.40% to 2.04/2.00/1.57/1.86%; relative to parameter-updating methods StruQ-Full, StruQ-LoRA and SecAlign-LoRA, LTBD matches or surpasses them on TaskTracker and on SEP for Falcon3 and Qwen2.5. Utility uses the AlpacaEval2 protocol with GPT-4o as judge against GPT-4 Turbo references, reporting WinRate; security reports ASR, with AlpacaFarm taking the worst case over Ignore, Completion and Ignore-Completion. LTBD matches or exceeds the undefended model on 6 of 8 utility evaluations and beats DefensiveToken on utility in 7 of 8 settings.

Under the Delimiter-Spoof adaptive attack, in which the adversary has full knowledge of the four delimiters and the input structure, LTBD's worst-case ASR remains nearly unchanged. The attack forges delimiter patterns to imitate, terminate or reopen trusted and untrusted regions across five variants; results show that forging delimiter syntax provides almost no advantage over plain injection. Worst-case ASR stays at 0.00% for Llama3-8B and Falcon3-7B, 0.48% for Llama3.1-8B, and rises only from 0.00% to 0.48% for Qwen2.5-7B. The authors attribute this to robustness residing in the learned embeddings tied to true structural boundaries rather than in the surface text of the delimiters.

Perspective

The result targets LLM applications that place untrusted external data in context, such as retrieval augmentation, tool calling and agentic workflows. It presumes the input can be clearly partitioned into a trusted instruction region and an untrusted data region, and that the deployer can train four delimiter embeddings per model. For settings where these two regions cannot be distinguished, or where external data must legitimately carry instruction authority, the method's premise does not hold. The reported evidence covers four open instruction-tuned models, training on the Cleaned Alpaca dataset for a single epoch, and evaluation on AlpacaFarm, SEP, TaskTracker and CyberSecEval2, so the scope of the conclusions is bounded by those models and benchmarks.

Worth watching: the five Delimiter-Spoof variants are constructed by the authors and also used as training augmentation, so the strength ceiling of the adaptive attack depends on the coverage of that attack family; the paper does not report the compute needed to train the four delimiter embeddings, nor a quantified inference latency, describing overhead only as negligible; on CyberSecEval2 the ranking between LTBD and training-based defenses varies, for example LTBD 7.27% versus StruQ-Full 10.9% and SecAlign-LoRA 18.2% on Llama3-8B, but LTBD 1.82% versus StruQ-Full 7.3% on Falcon3-7B, so relative advantage is not uniform across benchmark and model combinations. In addition, delimiter embeddings are tied to a specific model, and the text does not state whether retraining is needed after switching base models or whether embeddings transfer across models.

Sources