Introduction: One of the first steps in analyzing viral nucleotide/protein sequences is multiple sequence alignment (MSA). Due to the global effort of rapid diagnosis and advancement of sequencing technologies, more than three million SARS-CoV-2 genomes have been sequenced. This has given us an unprecedented opportunity to examine the capabilities of the MSA tools in handling datasets of various sizes. Materials and methods: In this study, we evaluated the speed, capacity, and user-friendliness of four frequently used MSA tools MAFFT (Multiple Alignment using Fast Fourier Transform), Clustal Omega, MUSCLE (Multiple Sequence Comparison by Log-Expectation), and T-Coffee on three laptops (ProArt Studiobook 16 OLED, ASUS Vivobook 17X, and ASUS Vivobook 14) using the Windows and Linux operating systems to align SARS-CoV-2 genomes and S protein sequences, which involved up to 2000 and 1,280,000 sequences, respectively. Results: Generally, runtime performance was similar across the laptops; however, only ProArt Studiobook 16 OLED could handle larger datasets and run Clustal Omega on Linux. Through using MSA tools to analyze S protein sequences with downloaded versions for the small dataset (≤2000), no significant difference in runtime was found for Clustal Omega, MUSCLE-super5 and MAFFT on Windows, whereas on Linux, MAFFT was the fastest. For the medium dataset (4000–64,000), MUSCLE-super5 (Windows) had the shortest runtime, while Clustal Omega took longer on both operating systems. For the large dataset (≥128,000), MUSCLE-super5 (Windows) was the fastest. As for SARS-CoV-2 genome sequences, T-Coffee failed to process them, whereas MAFFT consistently had the shortest runtime, irrespective of dataset size. Conclusions: Overall, the downloaded versions of MAFFT and MUSCLE were the most efficient for analyzing SARS-CoV-2 genome and S protein sequences, respectively.
Huang et al. (Fri,) studied this question.