PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 5, 20260 citationsOpen Access

Investigating Permutation-Invariant Discrete Representation Learning for Spatially Aligned Images

JSJamie S. J. StirlingNANoura Al-MoubayedHSHubert P. H. Shum

Key Points

  • This research explores whether positional information is necessary for discrete representations of spatially aligned images.
  • Proposed permutation-invariant vector-quantized autoencoder (PI-VQ) to eliminate positional information.
  • Introduced matching quantization algorithm to enhance effective bottleneck capacity.
  • Evaluated performance on image datasets CelebA, CelebA-HQ, and FFHQ.
  • PI-VQ captures global semantic features effectively without a learned prior.
  • Matching quantization increases effective capacity by 3.5 times compared to nearest-neighbor quantization.
  • Achieved competitive metrics for precision, density, and coverage in image synthesis.

Abstract

Vector quantization approaches (VQ-VAE, VQ-GAN) learn discrete neural representations of images, but these representations are inherently position-dependent: codes are spatially arranged and contextually entangled, requiring autoregressive or diffusion-based priors to model their dependencies at sample time. In this work, we ask whether positional information is necessary for discrete representations of spatially aligned data. We propose the permutation-invariant vector-quantized autoencoder (PI-VQ), in which latent codes are constrained to carry no positional information. We find that this constraint encourages codes to capture global, semantic features, and enables direct interpolation between images without a learned prior. To address the reduced information capacity of permutation-invariant representations, we introduce matching quantization, a vector quantization algorithm based on optimal bipartite matching that increases effective bottleneck capacity by 3. 5 relative to naive nearest-neighbour quantization. The compositional structure of the learned codes further enables interpolation-based sampling, allowing synthesis of novel images in a single forward pass. We evaluate PI-VQ on CelebA, CelebA-HQ and FFHQ, obtaining competitive precision, density and coverage metrics for images synthesised with our approach. We discuss the trade-offs inherent to position-free representations, including separability and interpretability of the latent codes, pointing to numerous directions for future work.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Stirling et al. (2026) studied this question.

synapsesocial.com/papers/69d1fcd4a79560c99a0a2904https://doi.org/10.48550/arxiv.2604.01843
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Neural Discrete Representation Learning2017 · 1,923 citations
  2. 2HyperVQ: MLR-based Vector Quantization in Hyperbolic Space2024
  3. 3Purrception: Variational Flow Matching for Vector-Quantized Image Generation2025
  4. 4Activation Map-based Vector Quantization for 360-degree Image Semantic Communication2024
  5. 5A lightweight perceptual-guided VQVAE for high-fidelity image compression2026