Skip to main content
Back to timeline
arXivSource publication:

Hyvärinen argues unsupervised learning has no single goal, proposing four goals that are orthogonal or even conflicting

Synopsis

The paper argues that unsupervised learning, including self-supervised learning, has no single goal or single maximized objective function, and proposes four distinct goals that are partly orthogonal or even conflicting: estimating the distribution, generating new data points, extracting features for downstream tasks, and understanding the data; for each goal it gives a problem statement, representative methods, and validation approaches, noting that understanding depends on identifiability while validation is often the hardest part.

AI-generated editorial illustration: What is the goal of unsupervised machine learning?

Interpretation

The paper argues that unsupervised learning is a heterogeneous field in which no single goal can be identified even conceptually, let alone a single maximized objective function, in contrast to supervised learning maximizing prediction accuracy and reinforcement learning maximizing total reward. Rather than treating unsupervised learning as one problem, the author makes the non-uniqueness of its goals the central claim and notes that the goals may be orthogonal or even contradictory. This is a conceptual argument based on classifying and reasoning about existing methods (kernel density estimation, score matching, noise-contrastive estimation, normalizing flows, GANs, diffusion models, autoencoders, contrastive learning, ICA, causal discovery), not on new experiments or benchmark results.

The author proposes four goals: estimating the distribution, generating new data points, extracting features for downstream tasks, and understanding the data; estimation and generation need a nonlinear function approximator, but that approximator can be a total black box, making these goals conceptually orthogonal to feature extraction and understanding. Beyond the related evaluation perspective of Theis et al. (2016), the author adds the goal of understanding the data and emphasizes that pursuing one goal may hinder another. The argument rests on structural analysis of method families, for example that GANs learn to generate directly without density estimation, and that latent variable models such as VAEs can generate but may not be competitive with methods optimized for generation.

The paper discusses validation separately for each goal: objective function values in density estimation usually lack clear interpretation and only allow comparison between architectures; validation of generated data lacks a general measure, inception scores are restricted to domains with labelled benchmarks, and discriminative measures are already optimized in GAN training and depend on the network architecture and training procedure. It ties the question of how to validate to which goal is being pursued, showing that evaluation criteria are not interchangeable across goals. Based on a review of existing validation practice and logical analysis; no new quantitative experiments are reported.

The paper stresses that identifiability is key to the goal of understanding the data: many widely used feature extraction methods determine features only up to a linear transformation, and nonlinear models such as VAEs define features only up to an orthogonal transformation or with even stronger indeterminacies, making interpretation questionable, whereas linear ICA achieves identifiability through non-Gaussianity, and related methods through temporal dependencies or nonnegativity. It uses identifiability as a criterion distinguishing feature extraction from understanding the data, noting that the former may require identifiability far less. Draws on existing results in linear ICA, nonlinear ICA, and causal discovery; it is a literature synthesis and conceptual argument.

Perspective

The framework addresses the basic case of real-valued vector data under a probabilistic view with an i.i.d. sample, an assumption the author notes can be relaxed. It is meant for researchers and practitioners who need to choose unsupervised methods for a specific task and design matching validation: first decide which of the four goals is being pursued, then choose methods accordingly (for example, identifiable linear or nonlinear component models when interpretation is needed, sampling methods when generation is needed) and select evaluation criteria that fit that goal. The author also states the list is not exhaustive, since causal discovery, data compression, and data restoration or cleaning could each be additional goals, so the framework is better treated as a starting point for discussion and classification than as a closed taxonomy.

Readers should note that the author explicitly calls the list of four goals tentative and possibly incomplete, so different researchers may draw different divisions; the judgments of orthogonality and contradiction among goals are conceptual arguments whose boundaries depend on the specific method and data; validation is described as difficult for both data generation and understanding the data, and no general solution is offered; and the author himself raises the debatable point of whether density estimation should count as a separate goal.

Sources