Lip reading via event cameras could capture subtle and similar lip motions with large dynamic range and microsecond latency, offering high temporal resolution over traditional frame-based cameras. However, existing approaches often overlook the exploitation of distinctive lip motion patterns, instead opting to adapt existing video recognition architectures. In this paper, we propose an event-specified framework named the Triplane Fusion Network (TF-Net) to harness lip movements through analysis from three distinct but complementary views. Specifically, while adhering to the standard XYT view, we further incorporate two additional perspectives: XT and YT, aiming to capitalize on the unique streaming nature of events over time. Given the three views, TF-Net contains multiple expert blocks for every view and mutual information exchange blocks that facilitate the multi-directional exchange of motion information among the different views. We observe that these three views complement each other, further enhancing the learning of the event-specified distribution. Extensive experiments validate the effectiveness of the proposed method, surpassing +1.6% and +2.3% accuracy over other competitive methods on the real-world dataset DVS-Lip and the synthetic dataset Modality, respectively.
Ju et al. (2026) studied this question.