Alligat0R
NeurIPS 2025 — Spotlight

Alligat0R
Pre-Training through Covisibility Segmentation for Relative Camera Pose Regression

Thibaut Loiseau1  ·  Guillaume Bourmaud2  ·  Vincent Lepetit1

1LIGM, Ecole des Ponts, Univ. Gustave Eiffel, CNRS    2IMS, Univ. de Bordeaux, CNRS

Abstract

Pre-training techniques have greatly advanced computer vision, with CroCo's cross-view completion approach yielding impressive results in tasks like 3D reconstruction and pose regression. However, cross-view completion is ill-posed in non-covisible regions, limiting its effectiveness. We introduce Alligat0R, a novel pre-training approach that replaces cross-view learning with a covisibility segmentation task. Our method predicts whether each pixel in one image is covisible in the second image, occluded, or outside the field of view, making the pre-training effective in both covisible and non-covisible regions, and provides interpretable predictions. To support this, we present Cub3, a large-scale dataset with 5M image pairs and dense covisibility annotations derived from the nuScenes and ScanNet datasets. The experiments show that Alligat0R significantly outperforms CroCo in relative pose regression.

Alligat0R teaser: covisibility segmentation vs cross-view completion

Alligat0R vs. CroCo. CroCo attempts to reconstruct potentially non-covisible regions (an ill-posed task), resulting in blurred outputs. Alligat0R instead explicitly segments pixels as covisible, occluded, or outside FOV, overcoming this fundamental limitation.

At a Glance

5M
Image Pairs (Cub3)
60.3%
Success on RUBIK (SoTA)
57ms
Inference Time

Contributions

Method

Alligat0R uses the same encoder-decoder architecture as CroCo (ViT encoder + cross-attention decoder), but replaces the reconstruction objective with covisibility segmentation. During pre-training, both images are processed symmetrically (no masking), and the model learns to classify every pixel into three classes. For downstream pose regression, a regression head is added on top, and fine-tuning proceeds in two phases: first with a frozen backbone, then jointly with the segmentation head.

Alligat0R architecture overview

Architecture overview. (a) During pre-training, each pixel is classified as covisible, occluded, or outside FOV. (b) For fine-tuning on pose regression, features from both views are pooled, processed through a shared MLP, and fed to separate translation and rotation heads.

Main Results

We evaluate Alligat0R against CroCo pre-training on the RUBIK benchmark (autonomous driving) and ScanNet1500 (indoor), using the same architecture and fine-tuning protocol for fair comparison.

Pre-training Data Backbone RUBIK ScanNet1500
5°/2m10°/5m 10°/0.25m10°/1m
Frozen backbone
CroCoCub3-50Frozen21.848.148.369.3
CroCoCub3-allFrozen8.925.28.724.9
Alligat0RCub3-50Frozen32.458.447.563.4
Alligat0RCub3-allFrozen55.382.378.292.5
Unfrozen backbone
CroCoCub3-50Unfrozen38.366.775.791.5
CroCoCub3-allUnfrozen38.666.264.185.0
Alligat0RCub3-50Unfrozen53.378.182.893.9
Alligat0RCub3-allUnfrozen60.381.985.595.1
Baseline
From scratchCub3-allUnfrozen29.152.337.057.2
Performance comparison across geometric challenges

Performance across geometric challenges. Success rate at 5°/2m on RUBIK, broken down by overlap, scale ratio, and viewpoint angle. Alligat0R trained on Cub3-all (red) consistently outperforms other configurations, especially in challenging low-overlap and extreme viewpoint scenarios.

Qualitative Results

Comparisons of CroCo's reconstructions and Alligat0R's covisibility segmentations on image pairs from nuScenes (outdoor) and ScanNet (indoor). While CroCo produces blurred results in non-covisible regions, Alligat0R correctly identifies the geometric relationships.

Covisible Occluded Outside FOV
Qualitative result on nuScenes Qualitative result on ScanNet

Each row shows: Reference image, Target image, Masked input for CroCo, CroCo reconstruction, and Alligat0R segmentation outputs for both views.

Citation

@article{loiseau2026alligat0r,
  title={Alligat0r: Pre-training through covisibility segmentation
         for relative camera pose regression},
  author={Loiseau, Thibaut and Bourmaud, Guillaume and Lepetit, Vincent},
  journal={Advances in Neural Information Processing Systems},
  volume={38},
  pages={13762--13789},
  year={2026}
}

Acknowledgments

This project has received funding from the Bosch Research Foundation (Bosch Forschungsstiftung), the European Union (ERC Advanced Grant Explorer, Funding ID #101097259), and was granted access to the HPC resources of IDRIS under the allocation 2024-AD011014905R1 made by GENCI.