Skip to main content
Back to timeline
arXivSource publication:

DisParQ learns discrete part concepts from a frozen DINOv2 backbone, reaching 83.2% top-1 on ImageNet linear probing

Synopsis

DisParQ learns spatially grounded, discrete concept representations from a powerful frozen vision-only self-supervised backbone without class labels or language supervision: each image patch is assigned to exactly one concept from a learnable prototype dictionary, only a sparse subset of concepts may activate per image, and continuous residuals quantized into discrete attributes capture how each concept varies; a spatial decoder reconstructs the backbone's representation from concepts and attributes alone. Across seven datasets it reaches 83.2% top-1 on ImageNet linear probing, closely matching its frozen DINOv2 teacher, achieves higher concept consistency than language-aligned models, remains competitive on fine-grained recognition, and enables cross-category part-based retrieval.

AI-generated editorial illustration: DisParQ: Self-Supervised Part Concepts for Interpretable Vision Foundation Models

Interpretation

DisParQ learns spatially grounded, discrete concept representations from a frozen vision-only self-supervised backbone, assigning each image patch to exactly one concept from a learnable prototype dictionary while allowing only a sparse subset of concepts to activate per image. Prior concept-based models are often limited to fixed categories or depend on language to define their concepts; DisParQ requires no class labels and no language supervision. The method description specifies the patch-to-concept assignment rule and the sparse activation constraint, and states that concepts come from a learnable prototype dictionary.

To capture how each concept varies across images (for example, the type of a "wheel"), DisParQ learns continuous residuals alongside the concepts and then quantizes them into discrete attributes. Separating the concept itself from its variable attributes lets the discrete representation keep interpretable part concepts while expressing how the same part differs across instances. The text explicitly describes learning continuous residuals and quantizing them, and gives the type of a "wheel" as an example of concept variation.

A spatial decoder reconstructs the backbone's representation from the concepts and attributes alone, so successful reconstruction means the discrete representation preserves the backbone's information. Using reconstruction as a check on information preservation turns whether discrete concepts sufficiently express backbone features into a verifiable objective. The text directly links successful reconstruction to information preservation as the methodological basis for the representation's validity.

Across seven datasets (ImageNet, PartImageNet, Places, CUB, Cars, Dogs, Flowers), DisParQ reaches 83.2% top-1 on ImageNet linear probing, closely matching its frozen DINOv2 teacher, achieves higher concept consistency than language-aligned models, remains competitive on fine-grained recognition, and enables cross-category part-based retrieval. It covers both general recognition and fine-grained benchmarks, surpasses language-aligned models on concept consistency, and demonstrates cross-category part-based retrieval. Evaluation spans seven datasets including general recognition and fine-grained benchmarks; ImageNet linear probing reports a specific 83.2% top-1 and compares against the frozen DINOv2 teacher.

Perspective

This work targets vision researchers and practitioners who want inspectable concept representations without class labels or language supervision, in settings built on a frozen vision-only self-supervised backbone that require spatially grounded part concepts and cross-category part-based retrieval. Its evaluation covers general recognition (ImageNet, PartImageNet, Places) and fine-grained benchmarks (CUB, Cars, Dogs, Flowers), so the conclusions mainly apply to the recognition and retrieval settings these datasets represent.

The current text is summary-level information and does not give the concept dictionary size, sparsity settings, attribute quantization scheme, decoder architecture, or full metrics on each dataset, so the influence of these design choices on results cannot be judged. The exact protocol and evaluation metrics for cross-category part-based retrieval are also not stated. Readers may watch how the number of discrete concepts and the sparsity constraint affect the balance between concept consistency and fine-grained recognition, and how far the reconstruction objective guarantees that the concept representation preserves backbone information.

Sources