Feature association is an essential step for vision-based localization methods. These methods rely on feature matching to estimate relative motion between consecutive frames using projective geometry. Regardless of the advances in feature association, most existing methods still rely on pairwise feature matching approaches and ignore the rich temporal context in image sequences. In this paper, we revisit the well-known Kanade-Lucas-Tomasi (KLT) feature tracker algorithm and propose a differentiable tracker model. The Proposed method is a fully convolutional neural network that learns spatio-temporal features to track keypoints across videos. Experimental results show that the proposed method outperforms the KLT method, especially in challenging environments.
Dias et al. (Tue,) studied this question.