The Voynich Manuscript (Beinecke MS 408, c. 1404-1438 CE) resists classification because standard statistical metrics are reproduced by both natural language and human-produced gibberish (Gaskell & Bowern 2022). We introduce Byte-Pair Encoding mean vocabulary morpheme length (BPE VMML) as a writing system classifier and apply it to 21 corpora spanning gibberish, 15 alphabetic languages, two cipher controls, and two logographic references. Three clusters emerge: gibberish (range 2.0-4.1), alphabetic natural and constructed language (4.06-5.25), and high sub-word regularity (5.62-5.92, containing Chinese Pinyin and Voynich). The Voynich VMML of 5.918 (95% CI 5.77-6.05) lies entirely above the alphabetic cluster maximum (4.51), with no overlap. Five mechanistic hypotheses are tested. The central finding is a systematic suffix inventory analysis: 37 distinct 3-character endings account for 80% of the entire Voynich corpus (vs. 172 for Latin, 109 for German). BPE applied to the last four characters of Voynich tokens alone yields VMML=5.868, matching the full-corpus value and exceeding Latin and German full-corpus values. The suffix inventory is stable across scribal sections (Currier A and B share a core of 9 dominant endings). The rank ordering Gibberish = Pinyin is preserved across all nine BPE parameter configurations tested. All code and data are publicly available.
Felipe J. F. da Silva (Sat,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: