Deoxyribonucleic acid (DNA) data storage has attracted significant attention as an alternative to conventional digital storage media due to its high information density and long-term durability.However, biochemical processes involved in DNA synthesis, polymerase chain reaction (PCR) amplification, and sequencing inevitably introduce errors, which lead to an important challenge to reliable data recovery.In this study, an error-resilient DNA encoding scheme based on the classical Hamming coding system is described.Specifically, the 8-bit American standard code for information interchange data representing a single character is expanded into a 12-bit codeword by inserting four parity bits using a Hamming code, which enables the detection and correction of single-bit errors.Subsequently, the resulting 12-bit binary codeword is converted into a DNA sequence using a base-4 encoding scheme, in which the quaternary symbols 0, 1, 2, and 3 are mapped to the nucleotides adenine, cytosine, guanine, and thymine, respectively.During decoding, parity checks inherent to the Hamming code are employed to identify and correct erroneous bits, thereby allowing accurate reconstruction of the original data.Furthermore, PCR-induced DNA sequence errors were experimentally introduced and analyzed, demonstrating that the embedded Hamming code effectively detects and corrects errors under practical biochemical conditions.This work shows that fundamental digital error-correcting codes can be effectively integrated into DNA-based storage systems, providing a modular encoding framework that can be extended to other coding schemes for reliable DNA data storage.
Lee et al. (Thu,) studied this question.