Canonical Correlation Analysis (CCA) is a fundamental method for multiview shared-space learning. However, its strict reliance on paired data poses a significant limitation, as such data is often difficult to obtain or entirely unavailable.
In this paper, we present Unpaired CCA (UCCA), a novel method that learns linear projections to maximize the correlation of the true underlying pairing without access to any paired samples. We first establish theoretical results connecting the Quadratic Assignment Problem (QAP) to CCA. Leveraging these theoretical insights, we derive a practical method to maximize correlation exclusively from unpaired data.
To the best of our knowledge, UCCA is the first approach to learn maximally correlated projections in a strictly unpaired setting. We validate UCCA on real-world multi-modal datasets, demonstrating that it significantly outperforms recent unpaired alignment baselines in recovering the underlying true correlation.
Our approach is grounded in a theoretical investigation of the unpaired CCA problem. The key insight is that finding the optimal pairing for CCA is equivalent to solving a Quadratic Assignment Problem with linear kernels.
First we define a proxy pairing for Unpaired CCA scenaiors, taking into account the orthogonal projections in the CCA objective (see Fig. 1).
Directly optimizing over permutations and orthogonal matrices simultaneously is computationally prohibitive. We relax the orthogonal group O(d) to the Frobenius sphere SF, yielding MCPSF, which we prove is tractable:
A critical question is whether the relaxation alters the optimal solution. Under a mild majorization condition on the singular values of the cross-correlation matrix - theoretically justified and empirically verified across all benchmarks - the two pairings are identical:
The key insight of our work is the definition of the MCPO(d) proxy pairing. Standard Maximum Correlation Pairing (MCP) is highly dependent on input rotation and fails to retrieve a valid pairing when data views are unaligned. MCPO(d) is entirely invariant to orthogonal projections, enabling it to successfully retrieve the true pairing. Its relaxation, MCPSF, inherits this invariance and likewise succeeds in recovering the correct correspondence.
Derived from our theoretical insights, UCCA operates in four steps. Given two unpaired datasets X and Y, the algorithm extracts representative anchor points from each modality, matches them using a QAP solver, and then applies standard CCA to the matched pseudo-pairs.
We evaluate UCCA on six multi-modal benchmark configurations spanning single-cell multi-omics (SNARE), image-text pairs (Flickr, COCO), and handwritten digit views (PC, KL, PA). The training data is strictly unpaired: each view is drawn from a disjoint subset of the original samples.
Using the learned embeddings from unpaired samples, we train a kNN classifier on one view and test it on the other. Higher is better. UCCA significantly outperforms all baselines across four datasets while maintaining comparable performance on the fifth.
| Method | PC-KL | PC-PA | KL-PA | SNARE | COCO |
|---|---|---|---|---|---|
| J-MDS-CCA | 0.083 | 0.087 | 0.211 | 0.247 | 0.237 |
| SCOTv1-CCA | 0.074 | 0.121 | 0.102 | 0.650 | 0.141 |
| SCOTv2-CCA | 0.148 | 0.090 | 0.113 | 0.646 | 0.164 |
| SCA | 0.097 | 0.083 | 0.100 | 0.312 | 0.161 |
| UCA | 0.088 | 0.100 | 0.159 | 0.378 | 0.132 |
| UCCA (ours) | 0.292 | 0.315 | 0.564 | 0.849 | 0.232 |
| Paired CCA † | 0.562 | 0.516 | 0.629 | 0.887 | 0.232 |
† Paired CCA uses the ground-truth pairing.
If you find this work useful, please cite: