TP-CRIV lets a third party with no white-box or API access verify an AI model's identity using only its ordinary black-box inference interface
Synopsis
The work proposes Third-Party Challenge-Response Identity Verification (TP-CRIV) for AI models, targeting a setting where the verifier has neither white-box nor API access to the claimant's model, can interact with the suspicious deployed service only through its ordinary black-box inference interface, and does not require protocol-specific cooperation from the service provider; under fresh, previously undisclosed requirements and network isolation it obtains empirical evidence as to whether the claimant locally possesses a model satisfying a predeclared identity, and the authors instantiate it for CNN image classifiers using probability-control-based witness generation, with experiments on ten ImageNet-pretrained TorchVision models showing clear same/cross-model separation and finite-chal
Fig. 1: Comparison of model-to-model verification approaches from the perspective of the evidence available to an independent third party and the resulting inference about the relationship between a claimant’s model and a suspicious deployed model.
arXivInterpretation
It proposes TP-CRIV, a third-party challenge-response identity verification framework that explicitly frames the verification goal as whether the claimant currently possesses and can utilize model-dependent information relevant to the claimed model identity. Existing watermarking, fingerprinting, and model similarity analysis mainly rely on predefined evidence or direct behavioral comparison and do not explicitly evaluate whether the claimant currently possesses and can utilize such model-dependent information; TP-CRIV makes this capability question the object of verification. At the abstract level the paper states the framework and its verification setting; this is a conceptual and methodological contribution whose formal details require the full text.
The framework operates under strict constraints: the verifier has neither white-box nor API access to the claimant's model, can interact with the suspicious deployed service only through its ordinary black-box inference interface, and does not require protocol-specific cooperation from the service provider. Compared with verification schemes that typically need model internals or provider cooperation, TP-CRIV relaxes the setting to a purely black-box, non-cooperative third-party scenario. The abstract explicitly states these three constraints, so this is a setting-level claim.
Verification is conducted under fresh, previously undisclosed requirements and network isolation so that the demonstrated capability cannot rely on online external assistance after challenge disclosure, and the resulting evidence is interpreted relative to independently specified and calibrated matching and non-matching operating situations and is statistical rather than cryptographic. Fresh challenges plus network isolation rule out post-disclosure online help, and independently calibrated matching/non-matching situations are used to interpret evidence rather than offering cryptographic certainty. The abstract states this design point and the nature of the evidence, without giving specific statistics or calibration procedures.
The authors instantiate TP-CRIV for CNN image classifiers using probability-control-based witness generation, and experiments on ten ImageNet-pretrained TorchVision models show clear same/cross-model separation and finite-challenge verification using independently calibrated thresholds. It grounds the framework in a concrete model family and dataset and reports experimental signals of same/cross-model separability and finite-challenge verifiability. The abstract reports the number of models (ten), their source (ImageNet-pretrained TorchVision), and the experimental conclusions (clear separation, finite-challenge verification), but gives no specific values, thresholds, or challenge counts.
Perspective
The framework targets a third-party verification setting: the verifier has no white-box or API access, can interact with the suspicious deployed service only through its ordinary black-box inference interface, and does not require protocol-specific cooperation from the service provider; verification runs under fresh, previously undisclosed requirements and network isolation, and evidence is interpreted relative to independently specified and calibrated matching and non-matching situations, being statistical rather than cryptographic. The abstract's instantiation targets CNN image classifiers using probability-control-based witness generation and is evaluated on ten ImageNet-pretrained TorchVision models.
The abstract gives no specific statistics, threshold values, challenge counts, error rates, or calibration procedure details, and does not state whether the instantiation extends to other model types or tasks; these are open questions for readers of the full text. The available text here is the abstract and page navigation information, without body figures or experimental details, so descriptions of experimental scale and statistical properties are bounded by what the abstract states.
