Skip to main content
Back to timeline
arXivSource publication:

CredLeak-Bench: All seven tested LLM agents leaked credentials to phishing, with recovery topping out at 24.2%

Synopsis

The authors introduce CredLeak-Bench, a sandboxed benchmark that measures leakage through actual submissions of synthetic credentials across directed authentication, autonomous inbox monitoring, and recovery settings with paired legitimate controls; evaluating seven models, they find every model leaks in both primary settings, phishing can induce disclosure without any authentication request, most leakage-reducing mitigations also impair legitimate task completion, and recovery ranges only from 0% to 24.2%.

Source-provided article image: CredLeakBench: Evaluating Credential Leakage and Recovery in LLM Agents
Figure 1 ·

Figure 1: (a) Agents receive explicit authentication requests or inbox-triage instructions and interact with local webpages using tools and account information. We measure information leakage to unauthorized destinations and false refusal on legitimate tasks. In the recovery setting, we evaluate whether agents bypass a phishing link and complete a genuine task through a trusted destination without leaking information. (b) The ideal performance lies in the upper-left corner, representing high task completion and a low leak rate. Connected points trace each model’s performance across the Directed (open circles), Autonomous (filled circles), and Recovery (open diamonds) settings.

arXiv

Interpretation

The benchmark pairs each phishing scenario with a legitimate control sharing the same brand, framing, and scaffolding, enabling joint measurement of leak rate and false-refusal rate rather than refusal behavior alone. Prior agent-security evaluations focus on prompt injection, malicious user requests, or unnecessary disclosure during legitimate tasks, and lack a design targeting deceptive recipients with matched legitimate counterparts. 512 directed and 256 autonomous scenarios, each phishing case matched to a legitimate control; leakage is adjudicated by matching submitted values against the synthetic vault rather than by model self-report.

Every evaluated model leaks credentials or sensitive information in both the directed and autonomous settings, and phishing emails can induce disclosure during inbox monitoring even when the user never asks the agent to authenticate. Treats whether disclosure requires an explicit authentication request as an independent variable, separating user-directed authentication from agent-initiated action. Directed leak rates range from 31.3% to 75.8% with false-refusal rates of 3.0% to 26.6%; autonomous leakage falls but false refusal rises sharply, e.g., Opus 4.8 drops from 31.3% to 1.6% leakage while false refusal rises from 6.6% to 88.3%.

Explicit domain-verification guidance can reduce leakage while increasing legitimate task completion in the directed setting, showing the security-utility trade-off is not universal. Treats security support as a controlled factor, comparing no added guidance, generic safety advice, domain-verification instructions, a URL-checking tool, and a vault policy. Domain verification moves six of seven models into the upper-left quadrant, improving both metrics; for GPT-5-mini leakage falls by 50.1 points while completion rises by 23.4 points; autonomous vault policy effects are mixed across models.

Avoiding leakage does not imply recovery: on phishing cases with a genuine pending task and trusted reference information, recovery ranges from 0% to 24.2%, and some runs leak before or after reaching the trusted endpoint. Introduces a recovery setting requiring the correct password to reach the trusted endpoint without vault information reaching the phishing endpoint, distinguishing mere refusal from safely completing the user's underlying task. 768 primary recovery scenarios; GPT-5-mini recovers most often at 24.2% but still leaks in 12.5% and leaves 63.3% incomplete; Opus 4.8 leaks in 0% yet never recovers; Qwen3.5-9B leaks before reaching the trusted endpoint in 23 runs and after trusted authentication in 12 runs.

Perspective

The benchmark targets email-driven web-agent workflows with access to a synthetic vault, suited to evaluating protection of credentials and sensitive information against deceptive recipients and the ability to find a trusted alternative. It covers eight service categories, 15 fictional services, four user framings, and eight deception cues, and separates directed authentication, autonomous inbox monitoring, and recovery. Trusted domains are supplied to the agent only in designated conditions, and connection-security conditions are simulated through the displayed URL scheme or a certificate-warning page without real network interception. In the recovery setting, recovery is defined as the correct target-account password reaching the trusted endpoint with no vault information reaching the phishing endpoint, measuring trusted authentication rather than completion of a downstream transaction.

Recovery ranges from 0% to 24.2%, and models differ widely in how leakage and recovery combine, so a single metric cannot summarize safety. The trusted-information sources (a full vault URL versus only an official domain in a prior email) differ in both information location and URL completeness, and the authors note the comparison reflects the practical difficulty of finding and using a trusted address. Security guidance effects are model-dependent: the vault policy raises recovery for three models and lowers it for three. Non-completion under cautious framing may reflect deferral rather than failed phishing recognition, and higher recovery under pressure does not establish lower leakage. The primary recovery set also includes 96 certificate-warning cases for which the simulator exposes no distinct trusted recovery route. These are scope and open questions rather than defects.

Sources