One model for all 25 sensor pairs
One set of weights matches RGB, event, thermal, depth, and LiDAR in every direction.
Towards Universal Dense Correspondence
Across Modalities
Robotics Institute, Carnegie Mellon University
Robotic platforms combine complementary sensors to maintain perception across diverse operating conditions, but dense correspondence methods are typically designed for RGB or for a small set of modality pairs. Relying on pair-specific models scales poorly with the number of sensors and forgoes potential positive transfer across modality pairs. We introduce TartanMatch, a unified model for dense correspondence across all 25 ordered pairs among RGB, event, thermal, depth, and projected LiDAR. TartanMatch adapts an RGB-pretrained correspondence model to modality-specific representation and learns all sensor pairs within one shared network using a mixture of real and synthetic data. On zero-shot real-world benchmarks, TartanMatch reduces endpoint error over the strongest per-pair baseline by 61.9% (cross-modal) and 49.6% (same-modal). We further show that joint training improves average matching accuracy over pair-specific training, especially on challenging cross-modal pairs. These results demonstrate that broad multimodal coverage and strong matching performance can coexist in a single model.
Explore dense correspondence across the full modality space.
Choose a demo
SELECT A MODALITY PAIR
DSERT-RoLL real sensor data. All methods share the same inputs. Compare each warped source with the target.
One set of weights matches RGB, event, thermal, depth, and LiDAR in every direction.
61.9% lower cross-modal error and 49.6% lower same-modal error on unseen real-world datasets.
See the accuracy resultsJoint training improves 11 of 15 pairs, reducing average cross-modal error by 34.1%.
See all 15 pairsTartanAir V2 validation
Preserve each sensor’s representation and learn dense correspondence together.
RGB enters directly; thermal intensity is replicated across three channels. Events use 15 temporal bins. Depth and projected LiDAR use log-depth + validity, preserving missing measurements explicitly.
Learned 1 × 1 projections map event and geometric inputs to three channels. A shared DINOv2 encoder, information-sharing transformer, and DPT heads predict dense flow and covisibility.
First, freeze the encoder while the input projections and matcher adapt. Then, fine-tune the full network to absorb the remaining modality shift.
No single dataset covers every sensor pair. We combine 16 sources across indoor, outdoor, aerial, and driving scenes, using native flow labels, geometry-derived correspondence, and synthesized thermal imagery.
Evaluated on real sensors from datasets held out of training.
Dense matching and relative pose on unseen real-world datasets.
One shared model versus 15 pair-specific models on TartanAir V2 validation. Lower error is better.
11 of 15 pairs improve. Average error drops by 34.1% across modalities and 17.8% within the same modality.
Milliseconds per image pair on H100. Shorter bars are faster.
27.8 ms per pair, about the same as UFM and 7.4× faster than MatchAnything.
@article{kwon2026tartanmatch,
title = {TartanMatch: Towards Universal Dense
Correspondence Across Modalities},
author = {Kwon, Hyeokjoon and Cai, Jiting and
Li, Ruogu and Bhandari, Kritan and
Hemkumar, Geethika and Maheshwari, Parv and
Qiu, Yuheng and Zhang, Yuchen and
Scherer, Sebastian and Wang, Wenshan},
journal = {arXiv preprint},
year = {2026}
}