PromptGate lifts query purity in open-set federated active learning from about 60% to above 95% using federated learnable prompts
Synopsis
PromptGate introduces a client-adaptive vision-language gating module for open-set federated active learning (OS-FAL): it learns class-specific context (CSC) prompts on a frozen BiomedCLIP backbone, splitting them into global tokens aggregated via FedAvg and client-local tokens, uses VLM pseudo-labels to filter the unlabeled pool into a high-purity ID candidate pool before querying, and then hands that pool to any downstream active learning strategy; on the FedISIC and FedEMBED federated medical imaging benchmarks, static VLM prompting degrades to roughly 50% ID purity, whereas PromptGate maintains above 95% purity with 98% OOD recall.
Interpretation
It proposes PromptGate, a learnable-prompt VLM gating module for open-set federated active learning, decomposing prompts into global tokens aggregated by FedAvg and client-private local tokens. Prior work such as OpenPath warms up open-set active learning with VLM text prompts in a centralized setting and treats OOD as a fixed global notion; PromptGate instead uses the VLM as a client-adaptive filtering front-end rather than defining an acquisition rule, adapting to heterogeneous OOD behavior across institutions. The method is fully specified: frozen BiomedCLIP image and text encoders, class-specific context prompts, a mixed configuration of 8 global plus 8 local learnable vectors per class (16 total), SGD (lr 0.002, momentum 0.9, weight decay 5e-4) for 15 epochs on a 128-shot subset per client, with only global tokens sent to the server.
A VLM-gated pseudo-labeling mechanism partitions the unlabeled pool into an ID candidate pool and a discarded pool before querying, acting as a plug-in pre-selection gate in front of any downstream acquisition strategy. The gate is strategy-agnostic and can precede Random, Entropy, FEAL, PAL, or LfOSA, whereas existing OS-FAL approaches rely on task-specific representations and treat OOD geometrically in feature space. On FedISIC (R=5), PromptGate variants average 96.5%-96.8% purity versus 60.7% for Coldstart and 74.3% for the static Baseline; on FedEMBED (R=10), variants average 90.3%-90.7% purity versus 88.0% for the Baseline; the paper reports average gains of +22% on FedISIC and +2-3% on FedEMBED, with BMA improvements of +1-3%, all averaged over three seeds.
Prompts are updated each round as new annotations arrive, progressively sharpening the ID/OOD boundary and turning the VLM into a dynamic gatekeeper. Static VLM prompting is fixed at R=0, whereas PromptGate updates both global and local tokens with CoOp-style prompt learning after every active learning round. On FedISIC, VLM query precision starts near 60% at R=1, dips temporarily at R=2 as the first CoOp update shifts the boundary, and rises above 95% from R=3 onward; Table 2 reports final-round OOD recall of 98.2% (Mixed), 98.5% (Global), and 98.8% (Local) versus 0.5% for the Baseline.
Ablation shows local token adaptation gives the strongest gating purity, while the mixed configuration achieves the best BMA when paired with a strong active learning strategy. The paper compares Mixed (8G-L8), Global (16G), and Local (16L) prompt configurations, indicating that semantic alignment tailored to client-specific artifacts outperforms generalized global prompts. On FedISIC, Local achieves the highest average purity (96.8%) and BMA (60.3%); on FedEMBED, Local keeps peak purity (90.7%) with only a -0.2% BMA trade-off against Mixed; on FedISIC, Mixed with Entropy reaches 66.1% BMA versus 64.4% for Random.
Perspective
The work targets federated medical imaging settings where data cannot be centralized and unlabeled pools contain OOD noise, applied to tasks such as dermoscopy and breast density classification; the method is designed as a plug-and-play module that can precede any OS-FAL acquisition strategy, adding only 16 prompt vectors (about 12K parameters) per client with a frozen backbone, making it suitable for resource-constrained hospital networks that avoid centralized data curation. The paper notes that dedicating a small annotation budget to prompt initialization is more effective than expanding static OOD prompts, pointing toward future deployment of the gate across more sites and modalities.
The paper itself notes that PromptGate's fine-grained ID classification remains limited (about 42% BMA on FedISIC), confirming the VLM should act as an OOD gatekeeper rather than a standalone classifier; early-round prompt adaptation can temporarily reduce query precision, and the system inherits the VLM backbone's domain biases. The authors propose future directions including enforcing cross-modal agreement between the VLM and task model to stabilize early rounds and selectively querying from the exploration pool. Readers should also note that FedISIC OOD is injected at a 50% ratio as a simulation, whereas FedEMBED OOD consists of naturally present clinical artifacts, and the differing OOD prevalence across the two settings leads to different purity gains; moreover, VLM-gated methods initially show slightly lower BMA than Coldstart on FedEMBED, so the effect of a smaller and less diverse filtered pool deserves continued observation under other data distributions.
