TartanMatch

Towards Universal Dense Correspondence
Across Modalities

Hyeokjoon Kwon*, Jiting Cai*†, Ruogu Li*, Kritan Bhandari, Geethika Hemkumar,
Parv Maheshwari, Yuheng Qiu, Yuchen Zhang, Sebastian Scherer, Wenshan Wang

* Equal contribution † Corresponding author

Robotics Institute, Carnegie Mellon University

Robotic platforms combine complementary sensors to maintain perception across diverse operating conditions, but dense correspondence methods are typically designed for RGB or for a small set of modality pairs. Relying on pair-specific models scales poorly with the number of sensors and forgoes potential positive transfer across modality pairs. We introduce TartanMatch, a unified model for dense correspondence across all 25 ordered pairs among RGB, event, thermal, depth, and projected LiDAR. TartanMatch adapts an RGB-pretrained correspondence model to modality-specific representation and learns all sensor pairs within one shared network using a mixture of real and synthetic data. On zero-shot real-world benchmarks, TartanMatch reduces endpoint error over the strongest per-pair baseline by 61.9% (cross-modal) and 49.6% (same-modal). We further show that joint training improves average matching accuracy over pair-specific training, especially on challenging cross-modal pairs. These results demonstrate that broad multimodal coverage and strong matching performance can coexist in a single model.

One model. Any sensor. Any pair.

Explore dense correspondence across the full modality space.

Choose a demo

SELECT A MODALITY PAIR

Source ↓Target →

RGB to Event

Source RGBfixed viewTarget Eventmoving viewSource moved into target viewshould line up with the target
Synthetic sequences from TartanAir V2. Use the grid or arrow keys to choose a pair.

One model for all 25 sensor pairs

One set of weights matches RGB, event, thermal, depth, and LiDAR in every direction.

Much lower error than existing matchers

61.9% lower cross-modal error and 49.6% lower same-modal error on unseen real-world datasets.

See the accuracy results
Average endpoint error (px) ↓
TartanMatchBest other method

Training on all pairs together beats one model per pair

Joint training improves 11 of 15 pairs, reducing average cross-modal error by 34.1%.

See all 15 pairs
Average endpoint error (px) ↓
Joint modelOne model per pair

TartanAir V2 validation

Five inputs with one shared backbone.

Preserve each sensor’s representation and learn dense correspondence together.

01

Keep the sensor representation

RGB enters directly; thermal intensity is replicated across three channels. Events use 15 temporal bins. Depth and projected LiDAR use log-depth + validity, preserving missing measurements explicitly.

02

Share the matching model

Learned 1 × 1 projections map event and geometric inputs to three channels. A shared DINOv2 encoder, information-sharing transformer, and DPT heads predict dense flow and covisibility.

03

Adapt in two stages

First, freeze the encoder while the input projections and matcher adapt. Then, fine-tune the full network to absorb the remaining modality shift.

Freeze 10 epochs→Fine-tune 35 epochs
TRAINING DATA

Real + synthetic data

No single dataset covers every sensor pair. We combine 16 sources across indoor, outdoor, aerial, and driving scenes, using native flow labels, geometry-derived correspondence, and synthesized thermal imagery.

16training sources650kpairs / epoch5zero-shot benchmarks
TartanAirV2 36.3%MegaDepth 20.2%DSEC 10.9%TartanAirV2 WB 9.7%12 other sources 22.9%

Universal coverage with strong matching performance

Evaluated on real sensors from datasets held out of training.

Accuracy

Dense matching and relative pose on unseen real-world datasets.

Cross-modal correspondence

EPE (px) ↓
TartanMatchStrongest baseline per evaluation
View exact values and baselines +

Joint training

One shared model versus 15 pair-specific models on TartanAir V2 validation. Lower error is better.

One model per pairJoint model (TartanMatch)

11 of 15 pairs improve. Average error drops by 34.1% across modalities and 17.8% within the same modality.

Speed

Milliseconds per image pair on H100. Shorter bars are faster.

27.8 ms per pair, about the same as UFM and 7.4× faster than MatchAnything.

Build on TartanMatch.

Paper (PDF) ↗
BIBTEX
@article{kwon2026tartanmatch,
  title = {TartanMatch: Towards Universal Dense
           Correspondence Across Modalities},
  author = {Kwon, Hyeokjoon and Cai, Jiting and
            Li, Ruogu and Bhandari, Kritan and
            Hemkumar, Geethika and Maheshwari, Parv and
            Qiu, Yuheng and Zhang, Yuchen and
            Scherer, Sebastian and Wang, Wenshan},
  journal = {arXiv preprint},
  year = {2026}
}
Enlarged TartanMatch architecture