LoFTR: Detector-Free Local Feature Matching with Transformers

Jiaming SunZehong ShenYuang WangHujun BaoXiaowei Zhou

article2021CVPR2,061 citations

Presents a detector-free local feature matching framework that uses Transformer attention to establish dense coarse-to-fine correspondences, successfully matching low-texture regions where traditional keypoint detectors fail.

Listen

Local image feature matching is foundational for core 3D computer vision tasks such as visual localization, mapping, and camera pose estimation. Traditional workflows rely on detecting distinct interest points (such as sharp corners) before describing and matching them. However, this detector-based framework frequently breaks down in realistic, challenging environments—such as indoor scenes with blank walls, motion blur, repetitive textures, or drastic viewpoint and lighting changes—because feature detectors cannot reliably identify repeatable points in indistinct regions.

The article introduces and evaluates Local Feature Transformer (LoFTR), a detector-free feature matching method designed to establish accurate, dense correspondences across images, including within low-texture and repetitive areas. The core objective was to demonstrate that eliminating the initial feature detection step and leveraging global attention mechanisms yields superior matching performance compared to established detector-based and detector-free techniques.

The approach operates in a coarse-to-fine sequence. First, a standard convolutional network extracts feature maps at both coarse (1/8 resolution) and fine (1/2 resolution) scales. The coarse features are enriched with positional encodings and processed through interleaved self-attention and cross-attention Transformer layers, which allow the network to establish global context across both images. Coarse matches are established using differentiable matching (such as optimal transport or dual-softmax) and then refined to sub-pixel accuracy within cropped local windows on the fine feature maps. To maintain computational feasibility, the architecture incorporates linear attention, reducing computational complexity from quadratic to linear relative to feature length.

Key evaluations across public indoor and outdoor benchmarks demonstrate substantial performance advantages. On the HPatches dataset, LoFTR achieved a homography estimation accuracy (AUC at 3 pixels) of 65.9%, markedly outperforming the top detector-based baseline SuperPoint paired with SuperGlue (53.9%) and previous detector-free approaches like DRC-Net (50.6%). In relative camera pose estimation on the indoor ScanNet benchmark, LoFTR achieved an AUC of 40.8% at a 10-degree error threshold, outperforming SuperGlue (33.8%) and DRC-Net (17.9%). On the outdoor MegaDepth dataset, LoFTR outperformed DRC-Net by 61% and SuperGlue by 13% at the 10-degree threshold. Furthermore, LoFTR achieved state-of-the-art visual localization rankings on the indoor InLoc benchmark and the outdoor Aachen Day-Night dataset.

These findings indicate that removing the interest point detector eliminates a major operational failure point in visual navigation and mapping systems. By integrating global context and position-aware descriptors, computer vision systems can reliably navigate and localize within previously intractable environments like featureless corridors and varying day-night cycles. The system processes a 640x480 image pair in roughly 116 to 130 milliseconds, making it practical for near-real-time deployment while significantly mitigating tracking failures and operational risk in robotic and autonomous platforms.

Teams developing visual localization, robotics, or augmented reality systems should consider transitioning from sparse detector-based pipelines to coarse-to-fine detector-free architectures where low-texture scenes cause reliability issues. For practical implementation, dual-softmax matching is recommended when lowest inference latency is needed, whereas optimal transport provides robust performance in complex indoor settings. Future work should evaluate the architecture across more severe environmental shifts, such as multi-season appearance changes, and explore further latency optimizations for resource-constrained hardware.

  • Paper: VGGT: Visual Geometry Grounded Transformer, Jianyuan Wang et al. (2025). VGGT extends transformer-based cross-image geometric attention mechanisms beyond pair-wise local feature matching to full feed-forward 3D reconstruction and camera pose estimation.
Cover for LoFTR: Detector-Free Local Feature Matching with Transformers

Abstract

We present a novel method for local image feature matching. Instead of performing image feature detection, description, and matching sequentially, we propose to first establish pixel-wise dense matches at a coarse level and later refine the good matches at a fine level. In contrast to dense methods that use a cost volume to search correspondences, we use self and cross attention layers in Transformer to obtain feature descriptors that are conditioned on both images. The global receptive field provided by Transformer enables our method to produce dense matches in low-texture areas, where feature detectors usually struggle to produce repeatable interest points. The experiments on indoor and outdoor datasets show that LoFTR outperforms state-of-the-art methods by a large margin. LoFTR also ranks first on two public benchmarks of visual localization among the published methods.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methods
  • 3.1 Local Feature Extraction
  • 3.2 Local Feature Transformer (LoFTR) Module
  • 3.3 Establishing Coarse-level Matches
  • 3.4 Coarse-to-Fine Module
  • 3.5 Supervision
  • 3.6 Implementation Details
  • 4 Experiments
  • 4.1 Homography Estimation
  • 4.2 Relative Pose Estimation
  • 4.3 Visual Localization
  • 4.4 Understanding LoFTR
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — LoFTR Detector-Free Feature Matching Architecture

    model/method

    LoFTR (Local Feature TRansformer) is a detector-free feature matching architecture that establishes dense correspondences across an image pair (IA,IB)(I^A, I^B) in a coarse-to-fine manner without sparse interest point detection:

    1. Local Feature Extraction: A standard convolutional backbone with Feature Pyramid Network (FPN) extracts multi-scale feature representations from both input images: coarse-level feature maps F~A,F~B∈RH8×W8×C\tilde{F}^A, \tilde{F}^B \in \mathbb{R}^{\frac{H}{8} \times \frac{W}{8} \times C} at 1/81/8 of the original image dimension (H,W)(H, W), and fine-level feature maps F^A,F^B∈RH2×W2×Cf\hat{F}^A, \hat{F}^B \in \mathbb{R}^{\frac{H}{2} \times \frac{W}{2} \times C_f} at 1/21/2 dimension.

    2. Coarse-Level Feature Transformation: The flattened coarse features F~A\tilde{F}^A and F~B\tilde{F}^B are augmented with 2D sinusoidal positional encodings (added once at the input) and passed through a Transformer module containing NcN_c interleaved self-attention and cross-attention layers, yielding context- and position-dependent transformed features F~trA\tilde{F}_{tr}^A and F~trB\tilde{F}_{tr}^B.

    3. Coarse-Level Matching: A differentiable matching module (using either dual-softmax or optimal transport) computes a coarse matching confidence matrix PcP_c. Matches satisfying confidence threshold θc\theta_c and mutual nearest neighbor (MNN) criteria form candidate coarse matches Mc\mathcal{M}_c.

    4. Coarse-to-Fine Match Refinement: For each coarse match (i~,j~)∈Mc(\tilde{i}, \tilde{j}) \in \mathcal{M}_c, local windows of size w×ww \times w centered at mapped fine positions (i^,j^)(\hat{i}, \hat{j}) are cropped from fine features F^A\hat{F}^A and F^B\hat{F}^B (which are concatenated with upsampled coarse transformed features F~tr\tilde{F}_{tr}). A fine-level LoFTR module with NfN_f attention layers transforms the window features to produce F^trA(i^)\hat{F}_{tr}^A(\hat{i}) and F^trB(j^)\hat{F}_{tr}^B(\hat{j}). The center feature of F^trA(i^)\hat{F}_{tr}^A(\hat{i}) is correlated with all vectors in F^trB(j^)\hat{F}_{tr}^B(\hat{j}) to produce a 2D matching probability heatmap. The final sub-pixel match location j^′\hat{j}' on image IBI^B is computed as the expectation over this probability distribution, producing fine matches Mf={(i^,j^′)}\mathcal{M}_f = \{(\hat{i}, \hat{j}')\}.

  2. Knowl 2 — Linear Attention in the LoFTR Module

    model/method

    Standard dot-product self- and cross-attention scales quadratically (O(N2)\mathcal{O}(N^2)) with sequence length N=H8×W8N = \frac{H}{8} \times \frac{W}{8}, computing:

    Attention⁡(Q,K,V)=softmax⁡(QKT)V\operatorname{Attention}(Q, K, V) = \operatorname{softmax}(QK^T)V

    where Q,K∈RN×DkQ, K \in \mathbb{R}^{N \times D_k} and V∈RN×DvV \in \mathbb{R}^{N \times D_v} are query, key, and value matrices.

    To enable global receptive fields over long 2D feature sequences with manageable compute cost, LoFTR utilizes linear attention by replacing the exponential kernel with an explicit feature map kernel:

    sim⁡(Q,K)=ϕ(Q)⋅ϕ(K)T\operatorname{sim}(Q, K) = \phi(Q) \cdot \phi(K)^T

    where ϕ(x)=elu⁡(x)+1\phi(x) = \operatorname{elu}(x) + 1.

    By the associativity property of matrix multiplication, the attention computation is evaluated as:

    LinearAttention⁡(Q,K,V)=ϕ(Q)(ϕ(K)TV)\operatorname{LinearAttention}(Q, K, V) = \phi(Q) \left( \phi(K)^T V \right)

    Since ϕ(K)TV∈RDk×Dv\phi(K)^T V \in \mathbb{R}^{D_k \times D_v} is computed first and feature dimension Dk,Dv≪ND_k, D_v \ll N, computational and memory complexity is reduced from O(N2)\mathcal{O}(N^2) to O(N)\mathcal{O}(N).

  3. Knowl 3 — Coarse-Level Differentiable Matching and Selection Formulation

    equation

    Given transformed coarse feature representations F~trA∈RN~×D\tilde{F}_{tr}^A \in \mathbb{R}^{\tilde{N} \times D} and F~trB∈RM~×D\tilde{F}_{tr}^B \in \mathbb{R}^{\tilde{M} \times D} at 1/81/8 image resolution, a pairwise score matrix S∈RN~×M~S \in \mathbb{R}^{\tilde{N} \times \tilde{M}} is computed as:

    S(i,j)=1τ⟨F~trA(i),F~trB(j)⟩S(i, j) = \frac{1}{\tau} \langle \tilde{F}_{tr}^A(i), \tilde{F}_{tr}^B(j) \rangle

    where τ\tau is a temperature hyperparameter and ⟨⋅,⋅⟩\langle \cdot, \cdot \rangle denotes the inner product.

    Under the dual-softmax matching formulation, the soft mutual nearest neighbor matching probability matrix Pc∈[0,1]N~×M~P_c \in [0, 1]^{\tilde{N} \times \tilde{M}} is computed by applying softmax normalization over rows and columns:

    Pc(i,j)=softmax⁡(S(i,⋅))j⋅softmax⁡(S(⋅,j))iP_c(i, j) = \operatorname{softmax}(S(i, \cdot))_j \cdot \operatorname{softmax}(S(\cdot, j))_i

    Alternatively, an optimal transport (OT) formulation treats −S-S as the cost matrix of a partial assignment problem solved via Sinkhorn iterations.

    The coarse-level match set Mc\mathcal{M}_c is selected by enforcing a confidence threshold θc\theta_c and mutual nearest neighbor (MNN) constraints:

    Mc={(i~,j~)  |  (i~,j~)∈MNN⁡(Pc),  Pc(i~,j~)≥θc}\mathcal{M}_c = \left\{ (\tilde{i}, \tilde{j}) \;\middle|\; (\tilde{i}, \tilde{j}) \in \operatorname{MNN}(P_c), \; P_c(\tilde{i}, \tilde{j}) \ge \theta_c \right\}

  4. Knowl 4 — Fine-Level Sub-Pixel Refinement via Correlation Heatmap and Spatial Expectation

    model/method

    To refine each coarse match (i~,j~)∈Mc(\tilde{i}, \tilde{j}) \in \mathcal{M}_c to continuous sub-pixel coordinates on image IBI^B:

    1. Coarse grid centers (i~,j~)(\tilde{i}, \tilde{j}) are mapped to fine-level feature maps F^A,F^B\hat{F}^A, \hat{F}^B (1/21/2 image resolution) at locations (i^,j^)(\hat{i}, \hat{j}).

    2. Local windows of spatial size w×ww \times w centered at i^\hat{i} on F^A\hat{F}^A and j^\hat{j} on F^B\hat{F}^B (concatenated with bilinear-upsampled coarse transformed features) are cropped and processed by NfN_f LoFTR transformer layers, producing transformed local patch features F^trA(i^)∈Rw×w×Cf\hat{F}_{tr}^A(\hat{i}) \in \mathbb{R}^{w \times w \times C_f} and F^trB(j^)∈Rw×w×Cf\hat{F}_{tr}^B(\hat{j}) \in \mathbb{R}^{w \times w \times C_f}.

    3. The center feature vector F^trA(i^)(i^)\hat{F}_{tr}^A(\hat{i})(\hat{i}) is correlated with every vector kk in window F^trB(j^)\hat{F}_{tr}^B(\hat{j}). A spatial softmax is applied over the correlation map to obtain a discrete 2D probability distribution P(k)P(k) across pixel locations k∈N(j^)k \in \mathcal{N}(\hat{j}) in the w×ww \times w window.

    4. The refined sub-pixel target location j^′\hat{j}' on image IBI^B is computed as the expectation over the spatial probability distribution:

    j^′=∑k∈N(j^)k⋅P(k)\hat{j}' = \sum_{k \in \mathcal{N}(\hat{j})} k \cdot P(k)

    Gathering pairs yields the final refined matches Mf={(i^,j^′)}\mathcal{M}_f = \{(\hat{i}, \hat{j}')\}.

  5. Knowl 5 — LoFTR Joint Dual-Level Training Loss Formulation

    equation

    The end-to-end training objective for LoFTR is defined as L=Lc+Lf\mathcal{L} = \mathcal{L}_c + \mathcal{L}_f, consisting of a coarse-level negative log-likelihood loss Lc\mathcal{L}_c and an uncertainty-weighted fine-level ℓ2\ell_2 regression loss Lf\mathcal{L}_f:

    Lc=−1∣Mcgt∣∑(i~,j~)∈Mcgtlog⁡Pc(i~,j~)\mathcal{L}_c = -\frac{1}{|\mathcal{M}_c^{gt}|} \sum_{(\tilde{i}, \tilde{j}) \in \mathcal{M}_c^{gt}} \log P_c(\tilde{i}, \tilde{j})

    where Mcgt\mathcal{M}_c^{gt} denotes the ground-truth coarse matches defined as mutual nearest neighbors between 1/81/8-resolution grid centers based on reprojection distance computed from ground-truth depth and camera poses.

    Lf=1∣Mf∣∑(i^,j^′)∈Mf1σ2(i^)∥j^′−j^gt′∥22\mathcal{L}_f = \frac{1}{|\mathcal{M}_f|} \sum_{(\hat{i}, \hat{j}') \in \mathcal{M}_f} \frac{1}{\sigma^2(\hat{i})} \|\hat{j}' - \hat{j}'_{gt}\|_2^2

    where j^gt′\hat{j}'_{gt} is the ground-truth warped location of query position i^\hat{i} onto image IBI^B, and σ2(i^)\sigma^2(\hat{i}) is the total spatial variance of the correlation heatmap representing prediction uncertainty. Gradients are not backpropagated through σ2(i^)\sigma^2(\hat{i}), and candidate matches whose warped location j^gt′\hat{j}'_{gt} falls outside the w×ww \times w local window are omitted from Lf\mathcal{L}_f.

  6. Knowl 6 — Indoor Relative Camera Pose Estimation on ScanNet

    data/table

    Evaluation of indoor relative camera pose estimation on 1,500 test image pairs from ScanNet with wide baselines and extensive texture-less regions. Essential matrices are estimated from predicted matches using RANSAC. Pose error is defined as the maximum angular error of rotation and translation. Performance is reported as Area Under the Cumulative Curve (AUC) of pose error at thresholds of 5°, 10°, and 20°:

    Category Method Pose estimation AUC
    @5° @10° @20°
    Detector-based ORB + GMS 5.21 13.65 25.36
    D2-Net + NN 5.25 14.53 27.96
    ContextDesc + Ratio Test 6.64 15.01 25.75
    SP + NN 9.43 21.53 36.40
    SP + PointCN 11.40 25.47 41.41
    SP + OANet 11.76 26.90 43.85
    SP + SuperGlue 16.16 33.81 51.84
    Detector-free DRC-Net†^\dagger 7.69 17.93 30.49
    LoFTR-OT†^\dagger 16.88 33.62 50.62
    LoFTR-OT 21.51 40.39 57.96
    LoFTR-DS 22.06 40.80 57.62

    †^\dagger indicates models trained on MegaDepth to assess cross-dataset generalizability. OT and DS denote differentiable matching via Optimal Transport and Dual-Softmax, respectively. LoFTR variants outperform both detector-based methods (SuperPoint + SuperGlue) and detector-free methods (DRC-Net) across all error thresholds.

  7. Knowl 7 — Outdoor Relative Camera Pose Estimation on MegaDepth

    data/table

    Evaluation of outdoor relative camera pose estimation on 1,500 test pairs sampled from MegaDepth validation scenes (Sacre Coeur and St. Peter's Square) characterized by large viewpoint variations and repetitive architectural patterns. Camera poses are recovered by solving the essential matrix using RANSAC from predicted matches, evaluated by pose estimation AUC at 5°, 10°, and 20° thresholds:

    Category Method Pose estimation AUC
    @5° @10° @20°
    Detector-based SuperPoint + SuperGlue 42.18 61.16 75.96
    Detector-free DRC-Net 27.01 42.96 58.31
    LoFTR-OT 50.31 67.14 79.93
    LoFTR-DS 52.80 69.19 81.18

    LoFTR with dual-softmax matching (LoFTR-DS) achieves the highest AUC across all thresholds, outperforming SuperPoint + SuperGlue by 10.62 percentage points at @5° and 8.03 percentage points at @10°, and detector-free DRC-Net by 25.79 percentage points at @5°.

  8. Knowl 8 — Planar Homography Estimation Evaluation on HPatches

    data/table

    Evaluation of homography estimation on the HPatches benchmark, comprising 52 sequences under illumination variations and 56 sequences under viewpoint variations. Images are resized with shorter dimension 480. Homographies are estimated via OpenCV RANSAC on up to 1K matches. Corner error AUC is reported up to error thresholds of 3px, 5px, and 10px:

    Category Method Homography est. AUC #matches
    @3px @5px @10px
    Detector-based D2Net + NN 23.2 35.9 53.6 0.2K
    R2D2 + NN 50.6 63.9 76.8 0.5K
    DISK + NN 52.3 64.9 78.9 1.1K
    SP + SuperGlue 53.9 68.3 81.7 0.6K
    Detector-free Sparse-NCNet 48.9 54.2 67.1 1.0K
    DRC-Net 50.6 56.2 68.3 1.0K
    LoFTR-DS 65.9 75.6 84.6 1.0K

    LoFTR-DS outperforms both detector-based and detector-free baselines across all thresholds, with the relative margin growing largest at the strict 3px error threshold (12.0 percentage points higher than SuperPoint + SuperGlue).

  9. Knowl 9 — Visual Localization on InLoc and Aachen Day-Night Benchmarks

    data/table

    Evaluation of 6-DoF visual localization on the InLoc indoor benchmark (handheld device track using the HLoc pipeline, evaluated at (0.25 m,10∘)/(0.5 m,10∘)/(1.0 m,10∘)(0.25\text{ m}, 10^\circ) / (0.5\text{ m}, 10^\circ) / (1.0\text{ m}, 10^\circ) on DUC1 and DUC2 queries) and Aachen Day-Night v1.1 outdoor benchmark (evaluated on night queries at (0.25 m,2∘)/(0.5 m,5∘)/(1.0 m,10∘)(0.25\text{ m}, 2^\circ) / (0.5\text{ m}, 5^\circ) / (1.0\text{ m}, 10^\circ)):

    InLoc Indoor Visual Localization
    Method DUC1 DUC2
    ISRF 39.4 / 58.1 / 70.2 41.2 / 61.1 / 69.5
    KAPTURE + R2D2 41.4 / 60.1 / 73.7 47.3 / 67.2 / 73.3
    HLoc + SP + SuperGlue 49.0 / 68.7 / 80.8 53.4 / 77.1 / 82.4
    HLoc + LoFTR-OT 47.5 / 72.2 / 84.8 54.2 / 74.8 / 85.5
    Aachen Day-Night v1.1 Night Queries
    Method Benchmark Track Night Accuracy
    R2D2 + NN Local Feature Eval 71.2 / 86.9 / 98.9
    LISRD + SP + AdaLam Local Feature Eval 73.3 / 86.9 / 97.9
    ISRF + NN Local Feature Eval 69.1 / 87.4 / 98.4
    SP + SuperGlue Local Feature Eval 73.3 / 88.0 / 98.4
    LoFTR-DS Local Feature Eval 72.8 / 88.5 / 99.0
    SP + SuperGlue Full HLoc Pipeline 77.0 / 90.6 / 100.0
    LoFTR-OT Full HLoc Pipeline 78.5 / 90.6 / 99.0

    On InLoc, LoFTR-OT sets state-of-the-art performance across the higher accuracy thresholds on both DUC1 and DUC2 indoor scenes. On Aachen Day-Night v1.1, LoFTR-OT achieves 78.5% localization rate at the strict 0.25 m / 2° threshold on night queries within HLoc.

  10. Knowl 10 — Ablation Study on LoFTR Architectural Components

    data/table

    Ablation study on the ScanNet dataset evaluating architectural choices of LoFTR using Optimal Transport matching, reporting pose estimation AUC at 5°, 10°, and 20° angular error thresholds:

    Ablation Variant Pose estimation AUC
    @5° @10° @20°
    1) Replace LoFTR module with convolution 14.98 32.04 49.92
    2) 1/16 coarse-resolution + 1/4 fine-resolution 16.75 34.82 54.00
    3) Positional encoding per layer (DETR-style) 18.02 35.64 52.77
    4) Larger model (Nc=8,Nf=2N_c = 8, N_f = 2) 20.87 40.23 57.56
    Full model (Nc=4,Nf=1N_c = 4, N_f = 1) 20.06 40.80 57.62

    Key takeaways:

    • Replacing the Transformer with standard convolution layers with comparable parameters causes a sharp drop in AUC@5° (from 20.06 to 14.98), demonstrating the importance of the global receptive field.
    • Coarser resolution feature maps (1/161/16 and 1/41/4) reduce AUC@5° from 20.06 to 16.75 while only slightly reducing inference runtime (104 ms vs. 116 ms).
    • Adding positional encodings at each Transformer layer (DETR-style) degrades performance compared to adding positional encodings once at the backbone output.
    • Doubling attention depth (Nc=8,Nf=2N_c=8, N_f=2) yields negligible performance gains over the default configuration (Nc=4,Nf=1N_c=4, N_f=1).

Coverage note — None was omitted; all key architectural components, mathematical formulations, training objectives, and benchmark experiments were captured.

References

  1. 1.Vassileios Balntas, Karel Lenc, Andrea Vedaldi, and Krystian Mikolajczyk. HPatches: A benchmark and evaluation of handcrafted and learned local descriptors. In CVPR, 2017.
  2. 2.JiaWang Bian, Wen-Yan Lin, Yasuyuki Matsushita, Sai-Kit Yeung, Tan-Dat Nguyen, and Ming-Ming Cheng. GMS: Grid-based motion statistics for fast, ultra-robust feature correspondence. In CVPR, 2017.
  3. 3.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, 2020.
  4. 4.Luca Cavalli, Viktor Larsson, Martin Ralf Oswald, Torsten Sattler, and Marc Pollefeys. Handcrafted Outlier Detection Revisited. In ECCV, 2020.
  5. 5.Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al. Rethinking attention with performers. ICLR, 2021.
  6. 6.Christopher B Choy, JunYoung Gwak, Silvio Savarese, and Manmohan Chandraker. Universal correspondence network. NeurIPS, 2016.
  7. 7.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. ScanNet: Richly-annotated 3d reconstructions of indoor scenes. In CVPR, 2017.
  8. 8.Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. Toward geometric deep slam. arXiv:1707.07410.
  9. 9.Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. SuperPoint: Self-supervised interest point detection and description. In CVPRW, 2018.
  10. 10.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. ICLR, 2021.
  11. 11.Mihai Dusmanu, Ignacio Rocco, Tomas Pajdla, Marc Pollefeys, Josef Sivic, Akihiko Torii, and Torsten Sattler. D2-Net: A trainable cnn for joint detection and description of local features. CVPR, 2019.
  12. 12.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016.
  13. 13.Jared Heinly, Enrique Dunn, and Jan-Michael Frahm. Comparative evaluation of binary features. In ECCV, 2012.
  14. 14.Martin Humenberger, Yohann Cabon, Nicolas Guerin, Julien Morat, Jer´ ome Revaud, Philippe Rerole, No ˆ e Pion, Cesar ´ de Souza, Vincent Leroy, and Gabriela Csurka. Robust Image Retrieval-based Visual Localization using Kapture. arXiv:2007.13867.
  15. 15.Wei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi, and Kwang Moo Yi. COTR: Correspondence Transformer for Matching Across Images, 2021.
  16. 16.Chaitanya Joshi. Transformers are Graph Neural Networks. https://thegradient.pub/transformers-are-graph-neural-networks/, 2020.
  17. 17.Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and Franc¸ois Fleuret. Transformers are RNNs: Fast autoregressive transformers with linear attention. In ICML, 2020.
  18. 18.Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. Reformer: The efficient transformer. ICLR, 2020.
  19. 19.Xinghui Li, Kai Han, Shuda Li, and Victor Prisacariu. Dual-resolution correspondence networks. NeurIPS, 2020.
  20. 20.Zhaoshuo Li, Xingtong Liu, Francis X Creighton, Russell H Taylor, and Mathias Unberath. Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with Transformers. arXiv:2011.02910.
  21. 21.Zhengqi Li and Noah Snavely. Megadepth: Learning single-view depth prediction from internet photos. In CVPR, 2018.
  22. 22.Tsung-Yi Lin, Piotr Dollar, Ross B. Girshick, Kaiming He, ´ Bharath Hariharan, and Serge J. Belongie. Feature Pyramid Networks for Object Detection. CVPR, 2017.
  23. 23.Ce Liu, Jenny Yuen, and Antonio Torralba. SIFT Flow: Dense correspondence across scenes and its applications. T-PAMI, 2010.
  24. 24.X. Liu, Y. Zheng, B. Killeen, M. Ishii, G. D. Hager, R. H. Taylor, and M. Unberath. Extremely Dense Point Correspondences Using a Learned Feature Descriptor. In CVPR, 2020.
  25. 25.Yuan Liu, Zehong Shen, Zhixuan Lin, Sida Peng, Hujun Bao, and Xiaowei Zhou. GIFT: Learning transformation-invariant dense visual descriptors via group cnns. NeurIPS, 2019.
  26. 26.David G Lowe. Distinctive image features from scale-invariant keypoints. IJCV, 2004.
  27. 27.Zixin Luo, Tianwei Shen, Lei Zhou, Jiahui Zhang, Yao Yao, Shiwei Li, Tian Fang, and Long Quan. ContextDesc: Local Descriptor Augmentation with Cross-Modality Context. CVPR, 2019.
  28. 28.Zixin Luo, Lei Zhou, Xuyang Bai, Hongkai Chen, Jiahui Zhang, Yao Yao, Shiwei Li, Tian Fang, and Long Quan. ASLFeat: Learning local features of accurate shape and localization. In CVPR, 2020.
  29. 29.Iaroslav Melekhov, Gabriel J Brostow, Juho Kannala, and Daniyar Turmukhambetov. Image Stylization for Robust Features. arXiv:2008.06959.
  30. 30.Krystian Mikolajczyk and Cordelia Schmid. A performance evaluation of local descriptors. T-PAMI, 2005.
  31. 31.Remi Pautrat, Viktor Larsson, Martin R Oswald, and Marc ´ Pollefeys. Online Invariance Selection for Local Feature Descriptors. In ECCV, 2020.
  32. 32.Jerome Revaud, Philippe Weinzaepfel, Cesar De Souza, Noe ´ Pion, Gabriela Csurka, Yohann Cabon, and Martin Humenberger. R2D2: repeatable and reliable detector and descriptor. NeurIPS, 2019.
  33. 33.Ignacio Rocco, Relja Arandjelovic, and Josef Sivic. Efficient ´ neighbourhood consensus networks via submanifold sparse convolutions. In ECCV, 2020.
  34. 34.Ignacio Rocco, Mircea Cimpoi, Relja Arandjelovic, Akihiko ´ Torii, Tomas Pajdla, and Josef Sivic. Neighbourhood consensus networks. NeurIPS, 2018.
  35. 35.Ethan Rublee, Vincent Rabaud, Kurt Konolige, and Gary Bradski. ORB: An efficient alternative to SIFT or SURF. In ICCV, 2011.
  36. 36.Paul-Edouard Sarlin, Cesar Cadena, Roland Siegwart, and Marcin Dymczyk. From coarse to fine: Robust hierarchical localization at large scale. In CVPR, 2019.
  37. 37.Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. SuperGlue: Learning feature matching with graph neural networks. In CVPR, 2020.
  38. 38.Torsten Sattler, Tobias Weyand, Bastian Leibe, and Leif Kobbelt. Image Retrieval for Image-Based Localization Revisited. In BMVC, 2012.
  39. 39.Tanner Schmidt, Richard Newcombe, and Dieter Fox. Self-supervised visual descriptor learning for dense correspondence. RAL, 2016.
  40. 40.Johannes L Schonberger and Jan-Michael Frahm. Structure-from-Motion revisited. In CVPR, 2016.
  41. 41.Hajime Taira, Masatoshi Okutomi, Torsten Sattler, Mircea Cimpoi, Marc Pollefeys, Josef Sivic, Tomas Pajdla, and Akihiko Torii. InLoc: Indoor visual localization with dense matching and view synthesis. In CVPR, 2018.
  42. 42.Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. Efficient transformers: A survey. arXiv:2009.06732.
  43. 43.Carl Toft, Will Maddern, Akihiko Torii, Lars Hammarstrand, Erik Stenborg, Daniel Safari, Masatoshi Okutomi, Marc Pollefeys, Josef Sivic, Tomas Pajdla, et al. Long-Term Visual Localization Revisited. T-PAMI, 2020.
  44. 44.Prune Truong, Martin Danelljan, L. Gool, and R. Timofte. Learning Accurate Dense Correspondences and When to Trust Them. ArXiv, abs/2101.01710, 2021.
  45. 45.Prune Truong, Martin Danelljan, Luc Van Gool, and Radu Timofte. GOCor: Bringing Globally Optimized Correspondence Volumes into Your Neural Network. In NeurIPS, 2020.
  46. 46.Prune Truong, Martin Danelljan, and Radu Timofte. GLU-Net: Global-Local Universal Network for dense flow and correspondences. In CVPR, 2020.
  47. 47.Michał Tyszkiewicz, Pascal Fua, and Eduard Trulls. DISK: Learning local features with policy gradient. NeurIPS, 2020.
  48. 48.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. NeurIPS, 2017.
  49. 49.Huiyu Wang, Yukun Zhu, Bradley Green, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. Axial-Deeplab: Stand-alone axial-attention for panoptic segmentation. In ECCV, 2020.
  50. 50.Qianqian Wang, Xiaowei Zhou, Bharath Hariharan, and Noah Snavely. Learning feature descriptors using camera pose supervision. In ECCV, 2020.
  51. 51.Kwang Moo Yi, Eduard Trulls, Vincent Lepetit, and Pascal Fua. LIFT: Learned invariant feature transform. In ECCV, 2016.
  52. 52.Kwang Moo Yi, Eduard Trulls, Yuki Ono, Vincent Lepetit, Mathieu Salzmann, and Pascal Fua. Learning to find good correspondences. In CVPR, 2018.
  53. 53.Jiahui Zhang, Dawei Sun, Zixin Luo, Anbang Yao, Lei Zhou, Tianwei Shen, Yurong Chen, Long Quan, and Hongen Liao. Learning Two-View Correspondences and Geometry Using Order-Aware Network. ICCV, 2019.
  54. 54.Zichao Zhang, Torsten Sattler, and Davide Scaramuzza. Reference Pose Generation for Long-term Visual Localization via Learned Features and View Synthesis. IJCV, 2020.

Citation

MLA
Sun, J., et al. “LoFTR: Detector-Free Local Feature Matching with Transformers”. arXiv, 2021, http://arxiv.org/abs/2104.00680v1.
APA
Sun, J., Shen, Z., Wang, Y., Bao, H., & Zhou, X. (2021). LoFTR: Detector-Free Local Feature Matching with Transformers. arXiv. http://arxiv.org/abs/2104.00680v1
Chicago
Sun, J., Z. Shen, Y. Wang, H. Bao, and X. Zhou. 2021. “LoFTR: Detector-Free Local Feature Matching with Transformers”. arXiv. http://arxiv.org/abs/2104.00680v1.
Harvard
Sun, J. et al. (2021) “LoFTR: Detector-Free Local Feature Matching with Transformers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2104.00680v1.
Vancouver
1. Sun J, Shen Z, Wang Y, Bao H, Zhou X (2021) LoFTR: Detector-Free Local Feature Matching with Transformers. arXiv

BibTeX

@article{sun2021loftr,
  title = {LoFTR: Detector-Free Local Feature Matching with Transformers},
  author = {Sun, Jiaming and Shen, Zehong and Wang, Yuang and Bao, Hujun and Zhou, Xiaowei},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2104.00680v1},
  eprint = {2104.00680}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/