PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 24, 20260 citations

Advancing explanatory and tonal dialectometry

View Full Paper
HSH.W.M. Sung

Key Points

  • This research aims to enhance dialectometry by integrating tonal analysis into the study of Yue dialects.
  • Utilized Levenshtein distance to calculate phonetic distances between Yue-Pinghua dialects.
  • Employed multidimensional scaling and cluster analysis for segmental classification.
  • Implemented multiple sequence alignment to enhance qualitative analysis of segmental features.
  • Applied modified Onset-Contour-Offset for measuring tonal distances and variations.
  • Identified key segmental dialect groups that fall within a continuum, ranging from 2 to 5 groups.
  • Established that tonal variation displays distinct categorical patterns compared to segmental variation.
  • Uncovered that not all dialect areas align in segmental and tonal patterns, indicating areas for future research.

Abstract

Dialectometry is a quantitative branch of dialectology, which makes use of computational and statistical methods on dialect data in order to understand language variation in space. The current dissertation presents how dialectometry can deepen our understanding of the variation of Yue dialects spoken in Southern China, as well as how Yue can help us broaden the scope of computational methods used in dialectometry, in order to account for tonal languages which are common in the world, but not so common as a subject of study within dialectometry. A number of research questions are addressed in this dissertation, and they fall under the following themes: 1) segmental classification of Yue dialects, 2) identification of characteristic features of Yue dialects, 3) tonal classification of Yue dialects and 4) comparison between segmentaland tonal variation of Yue dialects. The dataset used in the current dissertation consists of the IPA transcriptionof around 130 words in 113 Yue-Pinghua dialects. For the segmentalclassification, Levenshtein distance was used to calculate the phoneticdistances, followed by multidimensional scaling as a dimensionalityreduction technique and cluster analysis to explore the internal structureof the Yue-Pinghua dialect landscape. Traditional Northern Pinghua hasbeen found to be outside the Yue continuum, but that does not applyto traditional Southern Pinghua. A deeper analysis was then proceededwith the remaining 104 dialects (with Northern Pinghua removed as outliers).Yue dialects are found to lie in a big continuum on the segmentallevel, and the dialects can be divided into 2 to 5 big groups, dependingon the level of detail one seeks for. A dialectometric classification often receives criticisms for the lack of details or explanations of the identified dialect groups. This is becauseclassifications were based on distances, and it is difficult to retrieve thequalitative information (dialect features) after the quantification intodistances. For this reason, multiple sequence alignment (MSA hereafter)was employed before calculating the dialect distances for the segmentalclassification of Yue. MSA breaks the phonetic transcriptions downto historically related segments, and this was done to all the dialects simultaneously.The transformation of the transcription data makes soundsegments 1) historically more accurately aligned and 2) more suitable tobe analysed with post-hoc analyses, such as automatic feature extraction.Multi-aligned data can be easily integrated to the usual dialectometricworkflow: dialect distances can be calculated with the MSA data,followed by analyses like cluster analysis and multidimensional scaling.Using normalised Pointwise Mutual Information, an association measurecommonly used in natural language processing, characteristic featuresclosely associated to each dialect group (identified with the cluster analysis)can be identified. This technique has increased the explanatorycomponent of dialectometry, which goes beyond a mere dialect classification. On the other hand, tone languages, despite being quite commonwithin the world’s languages, are not studied a lot using dialectometrictechniques. One problem is that it is not immediately clear howtone distances can be measured. There have been several approaches,from simply binary similarities to sophisticated perception-based distancecalculation methods. The current dissertation applies Levenshteindistance on a representation of tones, modified Onset-Contour-Offset(mOCO hereafter), to obtain dialect distances and to explore how tonesvary between different dialects in space. mOCO is a more suitable tonerepresentation for dialectometry comparing to others since it can differentiate72 of the 73 tones attested in the dataset, with gradual distancesfrom one tone to another, which matches human perception. When applyingmOCO to a dialectometric analysis of Yue dialects on the tonallevel, the pattern of variation differs from segments. While segmentalvariation shows a continuum pattern, tonal variation shows a more categorical,dialect area pattern. A further analysis has shown that not allthe dialect areas show the same pattern on the segmental and tonal levels.The discrepancies shown in other dialect groups are intriguing andpose further research questions which await for further research.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

H.W.M. Sung (2026) studied this question.

synapsesocial.com/papers/699d401ade8e28729cf65291https://doi.org/10.48273/lot0709
Ask AI
Helpful
Bookmark
Share
View Full Paper