GateScope black-box audits 10 commercial LLM API gateways: some identify gpt-5 as the claimed model only 13.09% of the time, and o*ey bills 62.8% above expected
Synopsis
The authors introduce GateScope, a lightweight black-box auditing framework that uses only public APIs to evaluate LLM API gateways along response content, multi-turn conversation consistency, billing accuracy, and latency characteristics; controlled validation on official endpoints yields an average F1 of 0.968±0.085 across 24 models, and auditing 10 commercial gateways reveals identification rates as low as 13.09% for gpt-5, degraded multi-turn memory checkpoints, a 62.8% billing gap for o*ey on gpt-4o, and markedly higher latency variation for b*ie.
Figure 1: Overview of LLM API gateway architecture and the
· Page 2Interpretation
GateScope uses structured probes and behavioral signature vectors (answer quality, reasoning structure, scale, style, parsing quality) to train one-vs-rest XGBoost classifiers per model, deciding whether a response came from the claimed model without seeing gateway internals. Prior LLM fingerprinting work such as LLMmap, instructional fingerprints, UTF, AuditLLM, and FDLLM targets models or downstream artifacts; this work moves the audit target to third-party commercial gateways and enforces JSON output to reduce free-text noise. On official vendor endpoints, the average F1 across 24 models is 0.968±0.085, with most models above 0.95; five unseen models (Claude sonnet 4.5, Claude haiku 4.5, Qwen-plus, Qwen-turbo, DeepSeek-V3.2-Exp) are never misclassified as a known model.
In response-content audits of 10 commercial gateways, several platforms show identification rates far below the official baseline for specific models, suggesting possible model substitution or downgrading. The paper turns the question of whether the claimed model actually served a request into a quantifiable fraction, with each percentage computed over 275 responses per model. Official baselines identify gpt-4o, gpt-5, Gemini-2.5-pro, and Claude Sonnet 4.0 at 98.18%, 97.09%, 96.00%, and 97.09% respectively; b*ie identifies gpt-5 only 13.09% of the time, a*yi 48.00%, and several platforms fall below 60%.
A 25-turn conversation template with memory checkpoints, system_fingerprint counts, and cache-hit rates exposes differences in long-context retention and infrastructure behavior across gateways. Multi-turn consistency is decomposed into three observable indicators, and system_fingerprint changes plus cache-rate drops are used as supporting signals of model switching. The gpt-4o official baseline passes T10/T24/T25 with 1 fingerprint and a 48.7% cache rate; a*yi on gpt-4o passes T24 only once, shows 4 fingerprints, and a 4.2% cache rate; o*ey on gpt-4o-mini passes T24/T25 above baseline while reporting a 0.0% cache rate.
Billing audits compute expected cost from published rates and gateway-reported token usage and compare it with actual charges, finding notable deviations on a few platforms. The paper turns billing trust into a recomputable gap percentage and handles cached-token pricing separately. For gpt-4o most gateways show a 0.0% gap, while a*ix shows +7.6% and o*ey +62.8%; extended results show gpt-5 at +14.8% on a*ix and +15.0% on o*ey, and gemini-2.0-flash-lite at +19.3% on o*ub.
Perspective
GateScope targets ordinary clients and detects observable consistency gaps and supporting signals through public APIs rather than inferring gateway internals. It suits developers and enterprise buyers who use OpenAI-compatible interfaces and need to verify the claimed model, long-context memory, and billing trustworthiness; the paper explicitly frames results as a point-in-time snapshot rather than a long-term ranking of individual services and anonymizes the platforms.
Latency is framed by the authors as a supporting signal rather than direct evidence, since high variation may also come from gateway load or network conditions; all latency measurements come from a single fixed vantage point and should be read as point-in-time observations. Multi-turn and billing results rest on five repetitions per model per gateway and are snapshot in nature. The explanations offered for anomalies such as a*yi and o*ey are several "could suggest" mechanisms rather than attribution to a specific internal behavior; future work includes multiple vantage points, long-term continuous monitoring, and broader measurement dimensions for better attribution.
