A survey unifies joint, cross-modal and joint-editing video-audio generation on one distribution and gives the first systematic taxonomy of joint audio-visual editing: nine categories, 28 edit types
Synopsis
Observing that video and audio are perceived together yet most generative models treat them in isolation, this survey casts joint generation, cross-modal generation (one modality from the other) and joint editing as three problems defined on a single distribution over audio-visual pairs, organizes the field around one question—how the output is kept coherent across modalities in time and semantics—proposes a taxonomy along five design axes (generation strategy, audio representation, video representation, alignment enforcement, pretraining reuse), and, to the authors' knowledge, provides the first systematic taxonomy of joint audio-visual editing, mapped as nine edit categories spanning 28 edit types, together with methods, datasets and metrics for each setting and a closing list of open pr
Interpretation
The paper offers a unified formulation that treats joint generation, cross-modal generation and joint editing as three problems on one distribution over audio-visual pairs, with audio-visual correspondence—semantic agreement plus temporal localization—as the alignment score that separates the joint and cross-modal setting from two independent unimodal problems. Earlier overviews treat video generation, audio generation or audio-visual understanding in isolation; this work covers generation and editing of the two streams as a coupled pair, using cross-modal coherence as the organizing principle. The contribution is conceptual and formal: definitions of audio-visual correspondence, latent representations, temporal synchronization tolerance and datasets, plus tables instantiating the three problems task by task.
The paper proposes a five-axis design taxonomy: generation strategy (single-tower, dual-tower, cascaded, unified-token, guidance-based), audio representation (continuous latent, discrete tokens, mel-spectrogram, waveform), video representation (3D-VAE, 2D-VAE plus temporal module, discrete tokens, pixel), alignment enforcement (cross-attention, shared positional encoding, discriminator, classifier guidance, explicit prior) and pretraining reuse (from-scratch, single pretrained, dual pretrained). The taxonomy attributes method variation to a small number of axes that constrain different components of the generative tuple, notes that the axes are complementary rather than orthogonal, and marks the unoccupied cells. A table annotates dozens of methods axis by axis, and the text reports distributions such as continuous-latent audio representation used by twenty-eight of the twenty-nine methods in the table, 3D-VAE by nineteen methods, and cross-attention by nineteen methods.
The paper states that, to its knowledge, it is the first overview to systematically taxonomize joint audio-visual editing, mapping the space as nine edit categories spanning 28 edit types with representative operations and use cases. No prior work systematically classified editing of video and audio as a coupled pair in which an edit specified in one modality must propagate to the other. A dedicated table lists category, edit type, modality, representative edits and example use cases, and the text discusses synchronization, joint content, cross-modal transfer, identity and performance, scene and environment, narrative, generative, quality and restoration, and cross-cutting dimensions section by section.
The paper collects methods, datasets and metrics for each setting and identifies structural gaps: no joint method covered generates either stream as discrete tokens, dedicated audio-to-video generation has only isolated attempts, and joint audio-visual editing has no shared benchmark. These gaps are presented as concrete openings exposed by the taxonomy rather than as shortcomings of existing work. Dataset, benchmark and metric tables list representative resources, and the text notes that each editing method evaluates on data it assembled or repurposed itself, which makes editing results mutually incomparable today.
Perspective
The survey is aimed at researchers and practitioners who want to understand joint generation, cross-modal generation and joint editing within one framework, and at settings where one must judge how a method trades off representation, alignment and pretraining reuse. Its scope requires that at least one of video or audio be an output and the other appear in the pipeline; single-modality generation and audio-visual understanding without a generative or editing component are excluded, and fully closed commercial systems are excluded from the taxonomy because they disclose neither representations nor synchronization mechanisms, appearing only as user-facing capabilities in an appendix. The taxonomy is positioned as a map of the documented literature, reflecting coverage through mid-2026.
The authors note that coverage reflects the literature through mid-2026, that several covered systems are described only in preprints or technical reports whose details may change, and that empty regions of the taxonomy may fill rapidly; the five axes are complementary rather than mutually exclusive, so assigning a method occasionally requires judgment where papers are ambiguous, and the per-method tables record the authors' reading while the cited sources remain authoritative. In addition, joint audio-visual editing has no shared benchmark, so each method evaluates on data it assembled or repurposed itself and editing results are hard to compare directly; long-horizon coherence is evaluated on horizons far shorter than the minutes-long content the applications assume. The loaded text is the full paper, but tables appear as text and some cells are incomplete, which may limit precise restatement of individual methods' axis values.
