Researchers scanned 6.09 million public URLs and found 12,331 potential sensitive-data leaks, including 26 live password-reset links and 62 non-expiring JWTs
Synopsis
The authors built an automated detection system combining lexical URL filtering, dynamic rendering, OCR-based extraction, and content classification, applied it to 6,094,475 public URLs collected from VirusTotal, URLScan.io, Hybrid Analysis, the Wayback Machine, and RedHunt paste sites, and identified 12,331 potential exposures across authentication, financial, personal, and document-related categories, including 26 live password-reset links, 83 API keys, 12 publicly accessible e-signature workflows, 7 fully visible 2FA backup codes, and 62 JWTs lacking expiration constraints.
Figure 1: Leak Detection Workflow Overview
· Page 2Interpretation
The authors built a cross-platform dataset of roughly 6.09 million URLs from five public sources (URLScan.io, VirusTotal, the Wayback Machine, Hybrid Analysis, and RedHunt paste sites) and designed a lightweight four-stage detection pipeline: lexical filtering, content retrieval, rendering and extraction, and content analysis. Prior work on URL exposure was smaller-scale or focused on a single source, such as West and Aviv's analysis of query-string disclosures and GitHound or SecretFinder's detection of secrets in code repositories, but did not cover dynamic rendering or archived sources; this work is the first to combine lexical, structural, and dynamic inspection across scanning platforms, archives, and paste sites at large scale. The dataset contains 6,094,475 URLs, and the system processed them end-to-end in 3.8 hours; Stage 1 reduced the set to 832,714 candidates, the availability check left 149,450 live URLs, and 12,331 potential leaks were confirmed.
The system identified 12,331 potential URL leaks, distributed across categories as authentication and tokens 5,343 (43.3%), financial 2,412 (19.6%), personally identifiable information 1,440, metadata 678, travel 496, social 308, photos 291, e-signature 60, and documents 58; URLScan.io contributed 63.5% and the Wayback Machine 14.6%. No prior large-scale measurement had quantified the scale and category distribution of sensitive-information exposure in public URLs; this work provides the first cross-platform classification of leaks and shows that credential-related leaks dominate. Random samples from each category were manually inspected, and false-positive rates with Wilson 95% confidence intervals were reported; for the authentication and tokens category, the upper-bound false-positive rate was 0.26, yielding an estimated 3,954 true-positive leaks.
The authors report several critical exposure types: 26 publicly accessible password-reset links, 83 API keys appearing in query parameters or inline content, 12 publicly accessible e-signature workflows, 7 fully visible 2FA backup codes and 16 partially masked ones, and 62 JWTs lacking expiration controls or valid for multi-day durations. Previously such exposures were mentioned only in news articles or blog posts; this work is the first to systematically identify and categorize these specific exposure types at scale through an automated system, and it inspected JWT iat and exp claims. These findings come from analysis of rendered pages, query parameters, DOM structure, hidden input fields, and OCR-extracted content; the authors validated the behavior of 2FA and password-reset links using self-controlled accounts, confirming that visiting them did not invalidate credentials.
The authors revisited 2,000 previously flagged URLs after 75 days and found that 38% (762 URLs) remained publicly reachable without authentication, then manually verified a stratified subset of 200 and confirmed that the pages still exposed sensitive content. Prior work had not measured the temporal persistence of URL leaks; this work provides the first longitudinal data on leak persistence, showing that these exposures are not transient. The revisit sample was 2,000 randomly sampled URLs, of which 762 remained reachable after 75 days; 200 were manually verified in a stratified subset, confirming the content was still sensitive rather than placeholders or error states.
Perspective
The work applies to passive detection scenarios involving publicly accessible URLs, aimed at security researchers, repository operators, enterprise security teams, and developers. The detection pipeline and identifier sets are open-sourced, so repositories can use them for pre-publication redaction, and enterprises can use them to monitor their own domain exposures across public sources. The results are based on data collected between January and April 2025 and apply to three types of public sources: scanning platforms, archives, and paste sites.
True-positive and false-positive estimates rely on random sampling within each category and assume a uniform false-positive distribution, but certain domains (such as repeated login templates) may contribute disproportionately to false positives, and domain-level skew may bias estimates. Recall is difficult to measure at scale because no comprehensive labeled dataset of confirmed leaks exists, so evaluating system coverage remains an open challenge. Source imbalance partly reflects differences in platform accessibility, with URLScan.io offering broad self-service access while other repositories impose stricter controls. In addition, leak patterns and exposure formats may change over time, and platform structure updates and parameter-name evolution may affect detection effectiveness.
