1. Introduction
The same acoustic results can often be achieved through different articulatory gestures. Studies of articulatory trading relations show that listeners often categorize two articulatorily-different variants as the same sound, such as tensing vs. nasalization of /æ/ (De Decker & Nycz, 2012), laminal vs. apical English /s/ (Bladon & Nolan, 1977), interdental vs. dental English /θ/ (McGuire & Babel, 2012), and bunched vs. retroflex /ɹ/ (Delattre & Freeman, 1968). While these different articulations may produce the same acoustics, they can result in very different gestures that are more or less visible to an interlocutor. The co-occurence of acoustic and visual properties can further contribute to the strength or weakness of a particular contrast, which may have implications for sound change. For example, English [θ] is articulatorily more variable than [f] (sometimes produced as an interdental and other times as a dental), while the labiodental contact of [f] is more consistent between speakers (McGuire & Babel, 2012). This makes the visual cues for [θ] less reliable than its typological target of neutralization, i.e., [f], which may explain why /θ/ so often fronts, or neutralizes, to /f/ diachronically in English varieties, Germanic, Athabascan, Oceanic and other language families (Dubois & Horvath, 1998; Jones, 2002; Kjellmer, 1995; Smith, 2009).
In the current study, we examine the co-occurence of acoustic and visual properties of Mandarin alveolar and retroflex sibilants. Through an audiovisual production experiment with 30 native Mandarin speakers, we report an additional lip rounding gesture for retroflexes that is not seen in alveolars. We measure horizontal and vertical lip apertures, construct roundness scores, and then compare roundness with the acoustic properties of the sibilants. We find more variation in both the acoustic and visual properties of retroflexes, which may contribute to their merger with alveolars in particular varieties of Mandarin.
In the following subsections, we review how phonological contrasts are realized phonetically and how this leads to different cue relationships. We then turn to audiovisual integration in speech and, finally, introduce the Mandarin sibilant contrast and merger. We outline our methods for the audiovisual experiment in Section 2 and report results in Section 3. In Section 4, we discuss the Mandarin sibilant contrast in terms of both audio and visual properties, the stability of these properties, and the implications for sound change. We conclude with future avenues for perception work.
1.1. Phonetic cues and cue weighting
Phonological contrasts are known to be multi-dimensional, as distinctions are frequently achieved along various articulatory and acoustic parameters (Ladefoged, 1980). For example, the obstruent voicing contrast in English [±voicing] can be found in up to 16 different acoustic properties, such as voice onset time (VOT), closure duration, preceding vowel duration, and formant and f0 transitions, among other cues (Lisker, 1986). One hallmark of phonetic cues is that they have different “weights,” in that different cues contribute to the phonological contrast to various degrees (e.g., Repp, 1982). For example, in the English word-final voicing contrast, the duration of the preceding vowel is found to be more widely used than glottal pulsing or the fading-out of formants (Greenlee, 1980; Krause, 1982; Wardrip-Fruin & Peach, 1984; Nittrouer, 2005).
Cues are also perceptually integrated in the processing of a phonological contrast and can “trade” with each other (Fitch et al., 1980; Repp, 1982; Kingston et al., 2008). Primary and non-primary cues can be traded between each other to strengthen a contrast, either in an enhancing relationship (strong primary cue + strong secondary cue) or a compensatory relationship (weaker primary cue + stronger secondary cue). This occurs perceptually without changing the label of the category (Fitch et al., 1980; Francis et al., 2000; Iverson et al., 2005; Shultz et al., 2012).
Cue weights in perception and production are highly variable. They can differ between speakers of the same language (Guion, 1998; Jiang et al., 2020; Pfiffner, 2021; Hauser, 2023) and further within different speaker populations (Coetzee et al., 2018; Pfiffner, 2021, 2025). For example, Coetzee et al. (2018) show that the word-initial plosive voicing contrast in Afrikaans is produced with different weights given to VOT and the following vowel’s f0, which varies by speaker. Some speakers produce the contrast through VOT alone, and others devoice all plosives and rely on differences in the following vowel’s f0. Pfiffner (2021, 2025) shows that these differences in Afrikaans are further influenced by particular social factors of the speaker and listener; women, both younger and older, are found to devoice more often than men, for example. This also translates to perception, where in the absence of any prevoicing, the following vowel’s f0 at onset is a significant factor in the perception of phonological voicing, where plosives followed by low f0s are perceived as phonologically voiced, and plosives followed by high f0s are perceived as phonologically voiceless. However, this effect is greatest when listeners are hearing female talkers in comparison to hearing male talkers (Pfiffner, 2025).
One of the many unanswered questions in studies of phonetic cues is whether these trading and weighting relationships can be extended to other modalities, such as between acoustic and visual dimensions. If visual information is among the total cues employed by the speaker in production and available to listeners to recover intended speech gestures and phonological categories, then visual cues should also be perceptually integrated, weighted, and traded with each other and acoustic cues. This may then have important implications for how studies analyze sound change. In this study, we begin to fill this gap by examining the co-occurrence of visual and acoustic properties in the production of the sibilant contrast in Mandarin. This particular contrast is known to be neutralizing in particular regional varieties of Mandarin (e.g., Taiwan), so by examining which visual and acoustic properties are readily available to listeners and how much they vary in a particular contrast, we can better understand how mergers progress in both the acoustic and visual dimensions. Further, this provides background information that can then be used to test multimodal perceptual cue weighting and integration in future studies. In the following sections, we discuss audiovisual integration in speech and the Mandarin sibilant contrast/merger. Then, we introduce the present study, its results, and implications for sound change.
1.2. Audiovisual cue integration in speech
Acoustic accounts of sound change are widely accepted, but it has long been known that speech perception is multimodal, so a pure auditory perceptual model may not suffice in accounting for all sound changes. Listeners are also found to be sensitive to cues in the visual dimension (Sumby & Pollack, 1954; McGurk & MacDonald, 1976; Mayer et al., 2015; Ménard et al., 2016) or coupling sensorimotor information with the speech signal, such as tactile movement of the jaw and facial muscles (Gick & Derrick, 2009). Cues from all modalities are argued to be integrated without conscious awareness (Marks, 2004) to improve detection of speech and intelligibility. For example, Sumby and Pollack’s (1954) seminal study demonstrated that visual exposure to speakers’ facial movements, in addition to auditory cues, significantly improves speech intelligibility under noise when the amount of information is held consistent. When there is less noise, the contribution of visual cues is reduced and non-significant. Similar results are also reported by Grant and Seitz (2000), where visually matched and unmatched videos were played simultaneously with speech under noise. Compared to the audio-only condition, the detection of speech was improved to around 1.6 dB in the audio-visual (matched) condition. Since there was no difference between audio-visual (unmatched) and audio-only condition, it was concluded that the visual dynamic movements of articulators enhanced the accuracy of speech perception.
Introducing conflicting cues between audio and visual dimensions further supports the role of visual input in speech perception, as illustrated by the McGurk Effect (McGurk & MacDonald, 1976; Sekiyama & Tohkura, 1991), where audio-only inputs of [ba] or [ɡa] are accurately perceived, but when audio utterances are dubbed with incongruent videos, most listeners report hearing a fusion syllable of perceptual cues in both dimensions. For example, [da] is reported when the lip movement of [ba] is overlaid onto the auditory signal of [ɡa]. Among the three critical cues (i.e., place of articulation, manner of articulation, and voicing) that constitute [ba], [da], and [ɡa], the place and manner of articulation are visible, and voicing is audible. When visual information of [ba] (i.e., presence of lip closure) replaces that of [ɡa] and is combined with manner and voicing, listeners perceive [da].
Finally, brain imaging studies have shown that areas once believed to be sensitive only to auditory cues also respond to visual stimulation in a similar way (Calvert et al., 1997, 2001). Such activities also happen at the early stage of auditory processing, suggesting a more plastic, multimodal speech processing in humans (Musacchia et al., 2006). For example, lipreading with no auditory input stimulates activities in the auditory cortex measured by functional magnetic resonance imaging (Calvert et al., 1997), as well as meaningless pseudo-speech expressions, such as lip or jaw movement. However, non-linguistic facial expressions do not activate any response in the area. Musacchia et al. (2006) found that audio-visual stimulation of congruent or conflicting cues may influence auditory processing as early as 11 ms and last as long as 30 ms after exposure. Taken together, these studies show that subcortical auditory processing is also sensitive to visual input, implying a multi-modal nature of speech perception in underlying neurological and physiological mechanisms.
Despite the numerous studies of audiovisual cues in speech perception and their implications for cue relationships between modalities, there is a lack of research examining the co-occurrence of these cues in production and their subsequent cue weighting in perception. The present study will address the co-occurrence of visual and audio properties in the production of the Mandarin sibilant contrast, described in the following section.
1.3. The Mandarin sibilant contrast and merger
In Mandarin, there is traditionally a three-way contrast of sibilants between alveolars (/s/, /ts/, /tsʰ/), retroflexes (/ʂ/, /ʈʂ/, /ʈʂʰ/), and palatoalveolars (/ɕ/, /tɕ/, /tɕʰ/). These sets of sounds include fricatives (/s/, /ʂ/, /ɕ/), affricates (/ts/, /ʈʂ/, /tɕ/), and aspirated affricates (/tsʰ/, /ʈʂʰ/, /tɕʰ/), all of which include a period of frication noise. Affricates have an additional transient burst preceding the frication noise. Aspirated affricates in Mandarin do not reliably have a low-frequency, low-amplitude noise following the frication period (Tseng et al., 2011, p. 791–793); instead, they are frequently articulated as one combined gesture (see discussion in Section 2.4). Alveolars are typically characterized by a high amplitude, high frequency frication noise, in contrast with retroflexes, where the energy is more concentrated in a lower frequency range. Although there is a disagreement among existing studies for the place of articulation (POA) of the two sets of sounds in Mandarin, ranging from the dental region to alveolar ridge for the alveolars, and retroflex or post-alveolar for the retroflexes (Hauser, 2023, and see a discussion in Chang & Shih, 2015), the fact that there is a difference in POA is widely accepted. Palatoalveolars, on the other hand, can be seen as a sequence of a sibilant and the high front vowel [i] in complementary distribution with the other two sets of sibilants.
The traditional three-way contrast between retroflexes (/ʂ/, /ʈʂ/, /ʈʂʰ/), alveolars (/s/, /ts/, /tsʰ/), and palatoalveolars (/ɕ/, /tɕ/, /tɕʰ/), is notable because it is being reduced in some regional varieties of Mandarin, especially in Southern regions and Taiwan (Cheng, 1985; Lee-Kim & Chou, 2022). Specifically, /ʂ/, /ʈʂ/, and /ʈʂʰ/ have been observed to be both acoustically and perceptually more similar to /s/, /ts/, /tsʰ/ respectively, due to a change in spectral properties. This includes an increase in the center of gravity (CoG) in the frication portion of the retroflexes for speakers that participate in the merger in comparison to those that do not (Chang, 2012).1
One dimension that has not yet been considered in this contrast and merger are potential differences in visual properties. Traditionally, retroflexes are described as curling the tongue tip backwards to lengthen the front resonating tube, thus lowering the filtered frequency (Fant, 1960, 1971) in the frication proportion. However, it is also possible for rounding gestures to achieve similar acoustic properties by lengthening the front resonating tube, and previous studies have described Mandarin retroflexes as rounded (e.g., Chang, 2012). However, to our knowledge, these visual cues have never been tested and quantified. Previous studies of this contrast further focused solely on their acoustic properties (e.g., Hauser, 2023; Lee-Kim, 2022; Chiu et al., 2020), potentially leaving the rounding properties unaccounted for.
Existing experiments have suggested the possibility of an additional gesture; for example, an ultrasound study by Chiu et al. (2020) found variable degrees of tongue retroflexion in Taiwan Mandarin sibilants, ranging from complete articulatory distinction between alveolars and retroflexes to complete articulatory overlap. However, in the same study, such low-frequency CoG was given a lesser weighting by native speakers in perception. The authors reported that compared to [i] and [o], the following vowel [a] introduced more overlap in CoG between retroflexes and alveolars among four of the seven speakers they tested. It is possible that the following vowel had an anticipatory effect of tongue lowering, which causes less tongue retroflexion and more fronted contact closer to the alveolar ridge. However, this might not be the only answer, as speakers exhibiting context-specific merger patterns had a mismatch between acoustic results and reconstructed tongue position using ultrasound data. Chiu et al. further examined formant frequencies and found that these speakers had a lowered F2 in vowel onsets, which pointed to the possibility that there are additional articulatory gestures of lip rounding that may enhance or maintain this contrast.
1.4 The present study
The current study aims to fill the gap by documenting the visual properties and variability of lip rounding in the production of the Mandarin sibilant contrast. We further investigate the relationship between acoustic and visual properties co-occurring in speech production.
In an audiovisual production experiment, we employ the non-intrusive, non-obstructive, easily adaptable, and openly available computer vision tool, OpenFace 2.0, an independent video processing toolkit using Python to analyze facial movements and behavior (Baltrušaitis et al., 2018). The methods have been adopted and expanded in a laboratory phonology context (Krause et al., 2020, 2024). For example, Broś and Krause (2024) examine lip apertures in stop lenition in Canary Islands Spanish, finding that aperture, alongside acoustic measurements, can be used to probe the degree of lenition. Furthermore, they found that lip aperture captures focus-related, sentence-level prosodic differences in lenition more clearly than acoustic measures, while acoustics are more sensitive to stress-based differences. This lends further support for studying phonological contrasts in the audiovisual domain, as visual cues can contribute information independent of acoustic cues. Also, these studies in general demonstrate the feasibility and reliability of using motion capture and computer vision tools to study articulation and visual properties of phonological change. Our study therefore employs a similar audiovisual production paradigm to answer our research questions about the visual properties in this contrast, their variability, and their relationship to acoustic cues. Our research questions are as follows:
What are the visual properties of alveolar and retroflex sibilants? What patterns can we see between the two sets of sounds among different visual dimensions?
What is the relationship between acoustic properties and (visual) lip rounding? Is there a cue enhancing relationship in articulation whereby the visual parameters strongly predict acoustic properties?
2. Methods
We conducted an audiovisual production experiment in which participants were recorded while reading stimuli from a teleprompter set in front of them. This allowed us to obtain simultaneous high-fidelity acoustic recordings and videos of participants’ faces. The following sections report the stimuli, procedure, participants, and the data processing pipeline.
2.1. Stimuli
All stimuli in the experiment were monosyllabic CV words in Mandarin, balanced for tone (Tone 1, 2, 3, and 4). Due to lexical constraints, only 34 tokens were included (Appendix A), as well as 10 additional fillers. All items were repeated three times. Participants were presented with the Chinese characters, and for characters with multiple readings, we also presented Pinyin,2 a romanized representation of Chinese phonology, above the character. Word frequency was controlled to ensure minimal difficulty recognizing the character in production. Because the stimuli were single Chinese characters, frequency was extracted using character-frequency values from the SUBTLEX-CH-CHR corpus (Cai & Brysbaert, 2010). One target character (玼) was not listed in the corpus and assigned a value of 0, but it was still selected due to lexical constraints.3 Across all 44 items, mean raw character frequency was 41,980 (SD = 88,486), corresponding to 896 characters per million (SD = 1,889). Alveolar (n = 10) and retroflex items (n = 12) were matched in frequency: Frequency statistics for each member of the contrast (/s, ʂ, ts, ʈʂ, tsʰ, ʈʂʰ/) are given in Appendix B. Frequency did not differ significantly between alveolar and retroflex items in a Welch t-test on log-transformed frequency (i.e., log(raw frequency+1), t(12.29) = –0.38 , p = .711) or in a Wilcoxon rank-sum test on raw frequency (W = 65 , p = .767); the same pattern held for characters per million. All tokens were sibilant-initial and followed by homorganic vocalic nuclei (represented by /i/ in Appendix A).4 Two additional pairs between alveolars and palatoalveolars were formed, excluding retroflexes due to lexical constraints. Filler words had non-sibilants onsets, including stops, fricatives, and approximants, and only non-homorganic vocalic nuclei were used.
2.2. Procedure
Recordings were made in a sound-attenuated booth. Participants were seated about 40 inches away from a teleprompter, which consisted of an 11-inch screen and a piece of glass to reflect lights emitted by the screen, and a Razer Kyio Pro webcam that was mounted behind the glass. The whole teleprompter system was mounted on a tripod that was adjusted to the height of the participant’s eye level so that their face stayed in focus at the center of the camera’s view. An additional AKG C535 EB directional microphone was placed approximately 20 inches away to record high-quality audio for acoustic analysis. It was placed below the teleprompter and outside the view of the webcam so that the tracking of facial landmarks would not be affected. All participants were told to keep their upper body still during the experiment, although the facial feature extraction algorithm that we used is capable of accurate measurements despite slight head movements or variations in the distance between the speaker and the camera (Krause et al., 2020).
Audio was recorded using Praat (Boersma & Weenink, 2023) and a Steinberg UR22 audio interface at a sampling rate of 44.1kHz and a depth of 16 bit. Before video recording, the camera was calibrated following the methods of Krause et al. (2020). The calibration (or camera resectioning) is deemed necessary because it helps recover the exact pose and location of the camera in relation to the captured objects in three-dimensional space (Quan & Lan, 1999; Fraser, 2001; Lu, 2018). To calibrate, we took 30 pictures of an angled chessboard and fed the pictures into a Python script to estimate four parameters,5 which were later specified in OpenFace (Baltrušaitis et al., 2018) to accurately track the video. An extra copy of index audio was recorded using the webcam through Open Broadcaster Software 29.0.2 at 1080P and 60 frames per second, encoded in H.264 compression standard and .mp4 container format. Audio and video were aligned using Adobe Premiere Pro with reference to the index audio so that the start time and duration of the high-quality audio recording was synchronized with the video recording.
Each session consisted of one familiarization block with four trials and 13 test blocks. Stimuli were pseudo-randomized within each block. The familiarization trials were not included in the analysis. In each trial, participants were presented with a countdown timer from three seconds to zero, followed by a metronomic beep and a two-second presentation of stimuli. The metronomic beep was designed to be the starting point of visual parameter tracking. Participants were instructed to utter the word as soon as they saw the character but to keep their mouth closed before and after stimuli presentation. There was a short break at the end of each block. The experiment took an average of 15 minutes to complete and produced 132 audio and video tokens (i.e., 34 target items plus 10 fillers, with 3 repetitions each) per participant.
2.3. Participants
Thirty native Mandarin speakers participated in the study (19 female and 11 male speakers), all undergraduate students from China who were studying abroad in the United States. All were born and raised in China until they were 18 years old and continued to use Mandarin daily. Participants were also fluent in English or had at least had 10 years of exposure to formal instruction in English. Participants also reported use of local Chinese varieties, though their dominant language for communicating with friends and family was Mandarin, and they classified themselves as native Mandarin speakers. Finally, no one reported any cognitive, hearing, or speech difficulties, and all had normal or corrected-to-normal vision. Participants were paid 10 US dollars for their participation.
2.4. Data Processing
Audio recordings were annotated, aligned, and segmented by the first author using Praat (Boersma & Weenink, 2023). For each trial, the metronomic beep and the utterance were marked. The utterance was further segmented to isolate the frication portion of the sibilant. For both aspirated (n = 715) and unaspirated (n = 1239) retroflexes and alveolars, the frication portion began at the end of the transient burst visible in the waveform and ended at the onset of voicing and visible formant structure. We did not distinguish any subsegmental boundaries between the burst noise and aspiration for the aspirated affricates (ʈʂʰ and tsʰ) because no boundary was often visible. This is not unusual for Mandarin, where the internal boundary between frication and aspiration is not clearly marked (see Tseng et al., 2011, p. 791–793, who reported that approximately 10% of Mandarin tokens surveyed had a clear boundary). Further, cross-linguistically, it is not uncommon for aspiration noise to be partially or completely blended or subsumed by high-amplitude higher-frequency frication noise in both stops and affricates, which suggests a substantial gestural blending between abducted glottis and the stricture at the oral cavity (Stevens, 2000). We observed these patterns in our data (Figure 1) and thus analyzed the entire frication portion of the aspirated affricates.
Following previous research on sibilants (e.g. Lee-Kim & Chou, 2022; Rao & Shaw, 2021), tokens were high-pass filtered at 300 Hz, and four spectral moments were calculated from the middle 50% of frication. The four moments calculated for analysis include spectral mean (center of gravity, or CoG), spectral dispersion (SD), skewness, and kurtosis. It should be noted that CoG is prone to changes in sampling rate, or parameter settings for CoG calculation. Alternative approaches include examining the whole spectral slice (e.g., Lee-Kim et al., 2022), or Multitaper Spectral Analysis (e.g., Shadle, 2023). However, CoG has been widely used in previous studies (e.g. Chang, 2012; Chiu et al., 2020; Rao & Shaw, 2021), and it provides an easier way to perform inferential statistics and estimate multiple effects, such as visual cues.
We further segmented audiovisual-aligned videos into chunks that corresponded to the exact timestamp of each trial using the MoviePy library in Python and then imported videos to OpenFace (Baltrušaitis et al., 2018) to track and extract facial landmarks, head-turns, and head poses. A scheme of facial landmarks from the OpenFace documentation (Baltrušaitis, 2019) is shown below in Figure 2a, followed by the scheme overlaid on a real person producing round (left) and unround (right) sibilants in Figure 2b.
Figure 2a: The scheme of critical landmarks, adapted from Baltrušaitis (2019), where red dots are the focus of the current study.
With a reference point of the position of the camera, coordinates inside the three-dimensional space were recorded for the top and bottom points of the upper and lower lips, and the outer and inner sides of the two lip corners. We then found the midpoint between lip edges and obtained a new set of horizontal and vertical apertures, using point numbers 51, 62, 57, 66, 48, 60, and 54 and 64, respectively (Figure 2a). Excluded were data entries that had an average inferential confidence reported by OpenFace below 0.95 or head-turns that exceeded two standard deviations away from the position of the camera. Corresponding to TextGrid annotation, video frames were marked as either pre-speech, consonant, or vowel phase. We used the frames from speech onset to offset (i.e., excluding pre-speech) for the main analyses. The visualization of the rounding trajectory included all three phases. This leaves 1,954 alveolar and retroflex sibilant tokens overall, 56,755 video frames for the main analyses, and 144,725 video frames for the trajectory analysis. Palatoalveolar tokens served as additional fillers and are not analyzed here.
For variables of visual properties, we first computed Euclidean distances (Formula 1) in both directions using the midpoint between the inner and outer edges, i.e., the midpoints between 51 and 62, 57 and 66, 48 and 60, and 54 and 64, respectively. To account for the thickness of the lips, the distances were measured between midpoints and not landmarks at the edges. The distances form the vertical and horizontal apertures of the mouth. It should be noted that although the reported distance uses millimeters as the unit of measurement, it is only meaningful across participants in this study and should not be entirely equated to millimeters in the real world (Krause et al., 2020).
Formula 1: Euclidean Distance formula used to find the distance between a(Xa, Ya, Za) and b(Xa, Ya, Za).
We then quantified the roundness. We assume that the shape of the mouth in its resting position can be approximated to an ellipse (e.g., the right panel of Figure 2b), and when the lips are rounded, their shape approximates a circle (the left panel of Figure 2b). Under this assumption, to derive roundness, we calculated the eccentricity (Formula 2) of this ellipse, which is equal to the square root of 1 minus the square of the semi-minor axis (i.e., the shortest radius of the ellipse) divided by the square of the semi-major axis (i.e., the longest radius of the ellipse). In the context of the current study, the semi-major and semi-minor axes correspond to half of the horizontal and the vertical distances that were mentioned in the preceding paragraph, respectively. The eccentricity of ellipses, by definition, ranges from 0 to 1. For a perfect circle, where the major and the minor axes are equal, the eccentricity is 0. The value is larger than 0 for any ellipse. Therefore, the less rounded, the higher the eccentricity value. This could be counterintuitive if we adopt eccentricity as the direct representation of roundness, as more rounded tokens would have smaller values. Therefore, we decided that the roundness in this study will be defined as 1 minus the eccentricity (Formula 3). It then still ranges from 0 to 1, but more rounding will lead to larger roundness values. There are other potential ways to derive roundness (i.e., circularity, see Montero & Bribiesca, 2009, for a review), such as the variance of distances from the centroid to each critical landmark on the lips, which will be smaller in value for more rounded tokens. The eccentricity method was chosen for its simple application to the form of data available to us, i.e., coordinates in the three-dimensional space. As a result, we include the variables of horizontal aperture, vertical aperture, and roundness score (1 — eccentricity) in our analysis.
Formula 2: The eccentricity of an ellipse, where ɑ (i.e., the semi-major axis) is half of the horizontal distance, and β (i.e., the semi-minor axis) is half of the vertical distance.
Formula 3: Roundness of sibilants, where e is the eccentricity in Formula 2.
2.5. Analysis
Data analysis was conducted in R (R Core Team, 2024).6 Our first research question asked about the visual properties and variabilities of the retroflex-alveolar contrast. To answer this, we first report the summary statistics of the horizontal and vertical apertures, as well as derived roundness scores. Then we examine individual variability by plotting speakers on a spectrum of differences in roundness scores, operationalized with Cohen’s d (Cohen, 1988; see e.g. Brunelle et al., 2020) between the two categories, i.e., retroflex vs. alveolar. Male and female speakers are analyzed separately for their potentially different baseline in overall lip size. To test if the two categories can be differentiated by roundness score, we fit a linear mixed-effects regression model to predict the visual property with category (i.e., retroflex vs. alveolar), manner of articulation (i.e., affricates vs. fricatives) and the interaction between them (see the results section, as well as the supplementary materials for model details). To account for differential baselines of subjects and item-specific effects, by-participant and by-item random intercepts, as well as by-participant random slope, were added to the model.
For our second research question, the relationship between acoustic and visual properties in this contrast, we first establish whether acoustic differences are realized among the participants we sampled, given that there is an ongoing merger in particular varieties of Mandarin (e.g., Taiwan Mandarin). Although all of our participants are from mainland China, where the merger is more limited and stigmatized (Chang et al., 2013; Duanmu, 2007; Lin, 2007), we consider the step necessary as a sanity check for the status of the primary acoustic correlate, i.e., whether CoG has undergone merger among our participants in recent development of the sound change. We divided our participants by gender and examined individual productions with respect to the spectral distance between the two categories using Cohen’s d. Gender groups were also analyzed separately because there are known gender differences in sibilant production, where male speakers tend to have lower CoG measurements in comparison to female speakers. Further, it is common to see gender variation in a sound change (Trudgill, 1974; Labov, 1990, 2001), so it is helpful to analyze by gender in case any of our speakers are participating in the merger. Separate linear mixed-effects regressions with random intercepts for participants and items were fitted to predict the four spectral moments by category.
We then paired visual data with acoustic data and performed Pearson correlation tests to find which acoustic properties best correlate with visual properties. As mentioned above, previous studies focused on CoG as the primary acoustic cue, so we expect that CoG would best correlate with any visual properties. We then present a scatterplot of changes in CoG against changes in the roundness score with fitted regression lines separated by gender, which is predicted to show a positive correlation between changes in both measurements, as more changes in the length of the front resonating tube should result in more changes (i.e., lowering) of frequency in the frication (see the background section for details). We also fitted a linear mixed-effects regression model predicting CoG by measurements in the visual domain. For a cue enhancing relationship among audiovisual cues, we would expect changes in CoG to be significantly predicted by roundness.
To show that acoustics and visual properties may together support the distinction, we also fitted a conditional random forest model for all speakers to categorize property (i.e., alveolars vs. retroflexes) by both CoG and roundness score. While the method of using the generalized linear mixed-effects model for predicting categories with multiple cues and assessing cue weights has been widely used in previous studies (see, e.g., Jiang et al., 2020; Shultz et al., 2012) to assess weightings between different acoustic cues in production and perception (e.g., Berg, 1989; Doherty & Turner, 1996; Yang & Sundara, 2019), the analysis in the current study is limited by the fact that the acoustic and visual correlates are intrinsically correlated, which would introduce collinearity that may make the interpretation of regression models difficult. Therefore, we fitted the random forest model, which is less affected by collinearity, to better counteract the correlation between predictors.
To account for multiple comparisons, p-value adjustments were conducted using the Benjamini-Hochberg method within families defined by research questions to control the false discovery rate at q = .05. Tests for RQ1 (1 LMER model) were adjusted together, then for RQ2 (three correlation tests, five LMER models). We report both the raw p-values and adjusted p-values (referred to as BH-corrected p) in the following sections. Notably, the random forest model outputs (e.g., accuracy, variable importance, model metrics, etc.) are not p-values and were not considered in the correction.
3. Results
In the following sections, we discuss the results of our statistical modeling and how they address each of our research questions. Section 3.1 begins by examining the visual properties of the two sibilant classes, which are examined across the entire syllable instead of just during the frication period. In Section 3.2, we address the temporal relation between the frication and the realization of rounding gestures. In Section 3.3, we answer whether there is a cue-enhancing relationship between rounding and CoG and whether acoustic and visual properties together enhance the identification of retroflexes and alveolars.
3.1. Visual properties and variability
Our first research question asks about the visual properties of the two classes of sounds, as no previous studies (to our knowledge) have measured them. Figure 3 shows the mean values7 and distributional envelopes for horizontal aperture, vertical aperture, and roundness in alveolars and retroflexes. In general, alveolars show larger horizontal apertures than retroflexes (50.26 vs. 47.57 mm), whereas retroflexes are significantly more open in the vertical direction than alveolars (17.15 vs. 14.18 mm). Together, these differences yield a higher mean roundness score for retroflexes than for alveolars (0.08 vs. 0.04). The linear mixed-effects model predicting roundness also showed a significant main effect of property (F(1,32.35) = 15.56, p < .001, BH-corrected p = .001), with larger roundness scores for retroflexes than alveolars (b = 0.039, SE = 0.010, df = 33.41, t = 3.94, p < .001), indicating that retroflexes are indeed characterized by greater roundness.
The violin plots (Figure 3) indicate that alveolars and retroflexes differ in their overall lip configuration, but the size of the category separation is relatively modest when the raw aperture measures are considered separately, i.e., 2.69 mm and 2.97 mm for horizontal and vertical differences, respectively. Alveolars show a slightly larger mean horizontal aperture than retroflexes, whereas retroflexes show a slightly larger mean vertical aperture. In both dimensions, however, the two categories overlap broadly, and the mean differences are small relative to the amount of within-category spread. This suggests that the visual contrast is not carried by a clean separation in any single aperture measure. Rather, the categories are only partially differentiated in horizontal and vertical distance, with substantial token-to-token variability within each class.
The roundness plot shows a clearer contrast. Compared with the alveolars, the retroflexes’ distribution is shifted upward, with a longer upper tail. This means that retroflexes are generally produced with greater lip rounding, but also with much more variability (SD = 0.07) in how strongly that rounding is realized across tokens. By contrast, alveolars are concentrated in a narrower scope of low roundness scores (SD = 0.02), indicating a more stable articulatory pattern with relatively little rounding. In other words, the figure shows that retroflexes differ from alveolars more in the broader and right-skewed rounding distribution, less so by forming a separate category. This pattern is consistent with the idea that retroflexes employ the rounding gesture more strongly overall, while also allowing greater within-category variation in the magnitude of that gesture.
After examining the overall visual properties of alveolars and retroflexes, we next asked whether the visual contrast was shown uniformly across participants. Figure 4 therefore shows the variability in roundness among participants, measured by difference between alveolars and retroflexes in roundness for each sibilant pair, operationalized with Cohen’s d values. Most by-participant by-pair comparisons (i.e., 30 participants, with 3 pairs of sibilants each) still showed a robust visual contrast: Across the 90 comparisons, 75 fell in the stable-contrast range (d ≥ 0.8), 9 showed partial merger (0.2 ≤ d < 0.8), and 6 were categorized as full or reversed merger (d < 0.2). In raw values, the retroflex-alveolar difference in roundness ranged from 0.004 to 0.275 for male speakers and from -0.006 to 0.079 for female speakers. Thus, for most speakers, retroflexes remained more rounded than alveolars, although the strength of this visual contrast varied across individuals.
The visual merger cases included participants 15 and 29, both female. Participant 15 showed negative raw differences for all three sibilant pairs, with Cohen’s d values between –0.92 and –0.19, placing all three pairs in the full or reversed merger range. Participant 29 likewise showed three full or reversed merger pairs, with negative raw differences and Cohen’s d values between –0.79 and –0.39. These patterns indicate that, for these speakers, retroflexes were not produced with systematically greater roundness than alveolars, suggesting that rounding was not used as a reliable articulatory difference for the contrast.
By contrast, several speakers showed much stronger visual differentiation across pairs, including participants 4, 8, 22, and 28 among the female speakers, and participants 2 and 11 among the male speakers, all of whom displayed large positive raw differences and large Cohen’s d values. Other speakers, such as participants 7, 9, 13, 16, and 27, showed reduced, but still generally positive, contrasts in one or more pairs, indicating a weaker visual difference. However, because the number of speakers was unbalanced between female and male speakers, we do not interpret these patterns as systematic evidence for a gender effect here.
3.2. Trajectory of lip rounding
Figure 5 shows the articulation trajectory of the two classes of sounds by each individual onset along the three visual measurements. All timepoints were normalized relative to the duration of each syllable, with 0% and 100% being the stimuli onset and speech offset (i.e., the onset and offset of trials), respectively. The dotted green lines mark the median speech onset, and the dashed blue lines mark the median timepoint of the boundary between the sibilant and the nucleus. The sibilant-vowel boundary was at 19.901%, 38.504%, and 41.759% in normalized time of the syllable (i.e., speech onset to offset), corresponding to 79 ms, 183 ms, and 214 ms of mean onset duration for /ts-tʂ/, /tsʰ-tʂʰ/, and /s-ʂ/ pairs, respectively. Aperture values at the 0% and 100% indicate the average vertical and horizontal apertures of lips at the resting position.
Figure 5: Fitted trajectories of (a) horizontal and vertical apertures, and (b) roundness scores for syllables with alveolars and retroflexes, split by place and manner of articulation. The green vertical dotted lines mark the median speech onset. The blue vertical dashed lines mark the median sibilant-vowel boundary. The shaded ribbons show central 95% confidence intervals, i.e., 2.5%–97.5%.
Notably, compared to alveolars, retroflexes showed continued increases in vertical aperture and roundness after the initial lip opening, resulting in a higher vertical trajectory and a later peak in roundness. In the analysis of peak alignment, alveolar affricates and fricatives reached peak roundness at 45.6% and 43.0% of the normalized tracked trial interval (i.e., trial onset to speech offset), whereas retroflex affricates and fricatives peaked later, at 53.6% and 49.8%, respectively. However, because mean speech onset occurred at 60.1% of the interval, these retroflex peaks still preceded speech onset by 8.8 (time) percentage points for affricates and 5.5 percentage points for fricatives. Across all tokens, the overall peak occurred at 48.9%, or 11.2 percentage points before speech onset. Moreover, 67.0% to 81.4% of tokens reached peak roundness during the pre-speech interval, indicating that speakers begin preparing the rounding gesture before audible frication begins. This also further confirms that the two sibilant classes differ visually, with the contrast being articulatorily planned and becoming visible well before speech onset and remaining especially clear near the end of the syllabic period.
The graph also shows larger rounding values for retroflexes on the following vowel. Therefore, to fully analyze the time-course differences, we conducted a post-hoc analysis of rounding by syllable phase (consonant vs. vowel) by fitting a GAMM model8 during the speech onset-to-offset window to predict roundness score, implemented using bam in the mgcv package (Wood, 2023). We included property (alveolar vs. retroflex), phase (consonant vs. nucleus), manner of articulation (fricatives vs affricate) and their interactions as fixed effects, as well as smooth terms over normalized time, including by-property subsegment interaction and participant-specific smooths to capture individual variation, and an AR(1) error structure to account for autocorrelation with residuals from adjacent timepoints within each token. This model is intended to test the category and manner differences globally and the fine-grained differences over time, especially whether the trajectories diverge and at which phase, and what effects contribute to such divergence.
The GAMM showed a moderate fit that explained 48.7% of the deviance. No evidence of inadequate basis dimension was found in gam.check (all k-indices = 1.00, all p ≥ .49). The parametric terms showed a significant positive main effect of retroflex property (β = 0.0616, SE = 0.0030, t = 20.21, p < .001, BH-corrected p < .001), a negative coefficient for nucleus relative to consonant in the reference alveolar-affricate condition (β = –0.0084, SE = 0.0025, t = –3.42, p < .001, BH-corrected p = .001), and a significant interaction between property and phase (β = –0.0343, SE = 0.0034, t = –10.15, p < .001, BH-corrected p < .001), indicating that the retroflex advantage in roundness over alveolars was present, though it was smaller in the nucleus than in the consonant phase. There was no main effect of manner (β = –0.0019, SE = 0.0040, t = –0.48, p = .635, BH-corrected p = .750), and no interactions involving manner were significant. Smooth terms further showed no reliable time effect for alveolar consonants (edf = 1.001, F = 3.487, p = .062, BH-corrected p = .099), but significant time-varying effects for retroflex consonants (edf = 1.003, F = 140.748, p < .001, BH-corrected p < .001), alveolar nuclei (edf = 3.967, F = 18.008, p < .001, BH-corrected p < .001), and retroflex nuclei (edf = 7.456, F = 129.657, p < .001, BH-corrected p < .001).
The results indicate that the roundness difference emerges primarily in the consonant, though it is also present in the following vowel, and that retroflexes show a higher and more variable rounding trajectory over the whole portion of the syllable. This confirms that participants have employed extra articulatory gestures other than tongue retroflexion. Since the vowel qualities in the current study are tightly controlled (and, importantly, unround), we hypothesize that the additional rounding gesture may be temporally spread from the primary acoustic features in the course of a syllable, but we cannot generalize. More data is needed to test if the rounding difference still persists with more vowel qualities, particularly round vowels. Analyses in the study therefore use the average value of roundness across the syllable, instead of roundness during the frication phase alone.
The finding that rounding differences persisted into the syllable nucleus suggests that the labial gesture associated with retroflexes is not confined to the frication interval but carries over into the following vowel. This can be explained by a sequential combination of protrusion and rounding gestures: Lip protrusion may begin early to lengthen the front cavity and shape the frication acoustics, while the more visibly rounded configuration becomes strongest as the gesture stabilizes and continues into the vowel where rounding first fades away, followed by declining protrusion. This, in turn, provides a potential articulatory link between more retroflex rounding observed in the current study and the formant-onset differences for the following vowel in Mandarin, especially lower F2 onsets after retroflexes: If retroflexes are produced with greater and more persistent lip rounding, the following vowel begins from a more rounded labial configuration and therefore may contribute to a different formant onset. This hypothesis is partially corroborated by other production studies in Mandarin (e.g., Hauser, 2023), where F2 (especially F2 onsets) in the vowel period also differ between retroflexes and palatoalveolars. In addition, the aforementioned pre-speech peaks of roundness are also consistent with an anticipatory account rather than an unrelated gesture, which suggests that speakers often establish the rounding gesture before audible frication begins as part of speech planning for the upcoming CV sequence, with the same labial configuration continuing to shape both the onset consonant and the early portion of the vowel. However, ultrasound imaging and sagittal lip recordings will be needed to disentangle the relative timing of protrusion, rounding, tongue posture and transition, and frication.
3.3. The enhancing relationship between audio and visual correlates
We now turn to the relationship between acoustic and visual properties, beginning with showing that there is a significant acoustic contrast between the two categories among our subjects. However, since our participants came from across Mainland China, where the contrast may have interacted with local language varieties, we expected to see a range of individual variation. To further analyze this, CoG differences analyzed as Cohen’s d between retroflexes and alveolars were plotted for male and female speakers (Figure 6), which show that the retroflex-alveolar contrast is robustly maintained across speakers. Distance was also converted to Cohen’s d based on the difference in CoG between alveolars and retroflexes for each speaker-pair combination. At the level of speaker-pair combinations, 89 of 90 fell in the stable contrast range (d ≥ 0.8), and only one observation fell in the partially merged category. All 30 participants had mean acoustic distances above the maintained-contrast threshold when distances were averaged across the three pair types. In raw CoG units, speaker-pair distances ranged from 1168.95 to 7985.16 Hz, with a mean of 3861.22 Hz. Mean speaker-level Cohen’s d values ranged from 1.35 (Participant 16) to 15.03 (Participant 11). Overall, the figures indicate a clear-cut boundary between retroflexes and alveolars along CoG, with only one relatively attenuated speaker-pair combination showing partial merger, with no evidence for contrast collapse.
Figure 7 visualizes the four spectral moments and labels the mean of each distribution inside the boxplots. The clearest separation is in CoG values, with a mean for alveolars of 8453.80 Hz and 4879.35 Hz for retroflexes. The differences between the two categories are closer for spectral dispersion (2098.62 vs. 2082.00 Hz), skewness (1.05 vs. 1.21), and kurtosis (4.36 vs. 4.34). We fitted separate linear mixed-effects models with property, manner, and their interaction as fixed effects, and by-participant random intercepts, by-participant random slope for property, manner, and their interaction, as well as by-item random intercepts. For kurtosis, the maximal model failed to converge, so the final model retained a by-participant random intercept and slope for property, along with a by-item random intercept. The results showed that only CoG showed a significant effect of property (F(1,40.17) = 241.54, p < .001, BH-corrected p < .001) with lower CoG for retroflexes than alveolars (b = –3779.73, SE = 246.74, df = 42.38, t = –15.32, p < .001). By contrast, the effect of property was not reliable for spectral dispersion (F(1,35.07) = 0.49, p = .490, BH-corrected p = .490), skewness (F(1,33.48) = 3.52, p = .069, BH-corrected p = .089) or kurtosis (F(1,32.76) = 0.72, p = .403, BH-corrected p = .454). These results indicate that CoG retains the robust acoustic cue distinguishing alveolars from retroflexes. Although other spectral moments failed to reach significance, the strongest energy for retroflexes is still more concentrated in the lower portion of the spectrum, in comparison to alveolars.
We further visualized and tested the correlation between changes in roundness scores and changes in CoG values (Figure 8). Across the 90 speaker-pair combinations, the raw CoG difference and the raw roundness difference were positively correlated, but the association across both genders was not significant (r = 0.139, t(88) = 1.32, p = .190, BH-corrected p = .190). When examined separately by gender, the positive association was numerically larger and significant among female speakers (r = 0.441, t(55) = 3.64, p = .0006, BH-corrected p = .0018) than among male speakers (r = 0.296, t(31) = 1.72, p = .095, BH-corrected p = .142), which shows a modest but non-significant positive correlation between raw CoG differences and roundness differences. This indicates that in the current sample, the overall visual-acoustic association was driven more strongly by the female speakers.
Although the male regression line in Figure 8 appears steeper, which may indicate a larger effect of rounding per unit on CoG, the fitted regression slope and the Pearson correlation coefficient do not measure the same thing. The slope reflects the expected change in roundness for a one-unit change in CoG in the original measurement units, whereas the correlation coefficient is unit-free and reflects how tightly the observations cluster around a linear trend. Thus, a steeper line for male speakers can coexist with a larger correlation coefficient for female speakers.
Due to the unbalanced gender numbers, we refrained from interpreting the larger correlation coefficient or slope as any main gender effect. Generally, the plot and the tests indicate that larger acoustic differences between alveolars and retroflexes were associated with larger differences in lip roundness, such that speaker-pair contrasts with greater distinction in CoG also tended to show more roundness contrast. This indicates that, on average, participants who were more distinct in the acoustics were also more distinct in the visual dimension. Also, note that when the difference in roundness is close to 0, the difference in raw CoG is still substantial (i.e., 1208 Hz), which indicates the effect of tongue retroflexion gesture independent of lip rounding. Setting tongue retroflexion aside, this indicates that changes in CoG may be at least partially attributed to increased vertical aperture. These results further corroborate our expectation that there is an enhancing relationship between the primary acoustic correlate and visual property of roundness.
To further answer our second research question, we fitted a linear mixed-effects model predicting CoG with property and roundness, with by-participant random intercepts and random slopes for roundness, and by-item random intercepts (Table 1). We expect rounding to contribute to the distinction along CoG by a negative regression coefficient, due to its lengthening of the front resonating tube that lowers the filtered frequency. The decision to include roundness instead of horizontal and vertical apertures was made based on the possibility of non-independence between horizontal and vertical apertures, as both apertures may be affected by each other. A larger horizontal aperture may limit the range in which vertical aperture may vary. Although this may be solved by including interaction terms between the two apertures, using a single variable of roundness score best answers the effects of rounding on acoustics.
Table 1: Summary of the fixed effects in the linear mixed-effects regression model predicting CoG. Formula: cog ~ property + roundness_z + (1 + roundness_z | participant) + (1 | item). Retroflex was set as the reference level.
| term | est. | std. error | df | t_value | p_value |
| Intercept | 4819.216 | 159.095 | 33.523 | 30.291 | <.001*** |
| Property: alveolar | 3445.654 | 100.200 | 24.167 | 34.388 | <.001*** |
| Roundness (z-scored) | –714.525 | 193.527 | 31.769 | –3.692 | <.001*** |
Table 1 summarizes the fixed effects in the model predicting CoG from property and z-score standardized roundness. We assumed that the effect of roundness on CoG could vary across speakers, and therefore we fit a model with a by-participant random slope for roundness, which was the maximal structure that still converged without singular fit. To assess the contribution of roundness itself, we additionally compared two random intercepts-only models that differed only in whether roundness was included as a fixed effect. Model performance worsened significantly when roundness was removed (χ²(1) = 22.63, p < .001), with higher AIC and BIC values (AIC: 32305.43 vs. 32284.80; BIC: 32333.32 vs. 32318.27) and a lower log-likelihood (–16147.72 vs. –16136.40). We therefore interpret the model that includes roundness and the participant-specific random slope. Property was strongly significant: At mean roundness, alveolars had a significantly higher CoG than retroflexes (β = 3445.654, SE = 100.200, t = 34.388, p < .001, BH-corrected p < .001), consistent with our previous findings that the contrast is still maintained. The coefficient for roundness was also negative and significant (β = –714.526, SE = 193.527, t = –3.692, p = .0008, BH-corrected p = .0008), indicating that greater roundness was associated with lower CoG values. The omnibus test for roundness was likewise significant in the model ANOVA (F = 13.632, p = .0008), showing that roundness contributed to CoG differences overall. Taken together, these results confirm that CoG is the primary correlate to the sibilant contrast, and they support our hypothesis that the visual property of rounding effectively enhances sibilant contrast together with the acoustic property of CoG.
Recall that the acoustic and visual properties showed different variability among participants, creating different continua of less to more distinct. Individual speakers may thus have employed rounding to different degrees to achieve the contrast in CoG, essentially assigning different loadings to acoustic (i.e., CoG) and visual (i.e., roundness and apertures) correlates in production. Therefore, we further tested whether the inclusion of visual correlates (i.e., roundness) improves the identification of sibilant category overall. Existing studies in cue weighting (e.g., Jiang et al., 2020; Yang & Sundara, 2019; Shultz et al., 2012; Holt & Lotto, 2006; Berg, 1989; Christensen & Humes, 1996; Doherty & Turner, 1996) have employed inferential methods such as ANOVA and generalized linear mixed-effects models with different acoustic parameters to assess the contribution of individual acoustic correlates to the contrast, in both production and perception. However, as mentioned, the current study is limited by the collinearity between roundness and CoG. Therefore, we performed a conditional random forest model, which is less affected by collinearity issues, to categorize retroflexes from alveolars using both CoG and roundness. Using party::cforest (Hothorn et al., 2006) in R, the forest was grown with 500 trees with two predictors considered at each split. Model performance was evaluated with out-of-bag predictions rather than a separate training and test partition: For each tree, 63.2% of the data was randomly sampled without replacement for training, and the remaining cases served as out-of-bag test cases. This sampling was random and not additionally stratified by speaker or contrast.
The model showed high classification performance, with an out-of-bag accuracy of 97.6% and similarly strong precision (0.9695), recall (0.9868), and F1 (0.9781) values. It correctly classified 859 out of 892 alveolar tokens and 1,048 out of 1,062 retroflex tokens. Conditional variable importance indicated that classification was driven overwhelmingly by CoG (conditional importance = 0.381), whereas roundness slightly contributed unique information after CoG was taken into account (conditional importance = 0.000762). This indicates that CoG contributed primarily to the identification of the contrast among the speakers in the current study. We interpret this in relation to the fact that most speakers in our dataset maintained significant differences between alveolars and retroflexes in acoustics. In the context of where the loss of retroflexion is more prevalent (e.g., Taiwan Mandarin), visual measures may contribute more substantially than the current study shows.
4. Discussion
To summarize, the results of our experiment show that the additional rounding gesture is significantly different between the two categories. The rounding is found to be most aligned with the start of the onset portion of the syllable, or before the consonant onset at the end of the pre-speech interval, which suggests that the rounding gesture is initiated anticipatorily and is already well underway when frication begins. Rounding differences also carry over into the nucleus. Retroflexes are characterized with larger vertical apertures and smaller horizontal apertures than alveolars, leading to more roundness. The difference in apertures and roundness is consistent across all sibilants. Acoustically, although there is an ongoing merger from retroflexes to alveolars among some Mandarin speakers, especially in Taiwan, the participants that we sampled did not exhibit substantial acoustic overlap between the two categories. The primary acoustic correlate to this contrast (i.e., CoG) was found to be distinct between retroflexes and alveolars. Participants also vary in the reliance on visual correlates in production, showing significant individual variation in roundness. We further found an enhancing relationship between CoG and roundness. In a linear mixed-effects regression model, more rounding significantly predicted smaller CoG values, which is consistent with the source-filter articulation model that predicts a longer resonating tube lowers the filtered frequency in the oral cavity. More changes in CoG are also positively correlated with more changes in roundness, especially for female speakers. A classification model (RF) was fitted and found that both CoG and roundness contributed to the contrast, though the importance of roundness is minor. We now turn to implications for sound change and finish by remarking on multimodal correlates in a phonological contrast.
4.1. Articulatory markedness of Mandarin retroflexes
As mentioned previously, certain varieties of Mandarin have an ongoing sibilant merger from retroflexes to alveolars. The merger is notable for two reasons. First, broadly speaking, retroflex neutralization tends to be unidirectional, where retroflexes merge towards the alveolars, but not the other way around.9 Second, since Mandarin has a three-way contrast, the merger in Mandarin is not only unidirectional but also asymmetrical, as retroflexes almost never merge with palatoalveolars. This is perhaps because alveolars and retroflexes are both produced with the tongue tip, while palatoalveolars are produced with the tongue blade, and furthermore palatoalveolars use F2 at the onset of the following vowel as a distinct acoustic property from alveolars and retroflexes (Hauser, 2023). But this still does not explain why retroflexes are the sibilant class that merges with alveolars rather than vice versa.
The facts of this merger suggest that some unique property of retroflexes may be driving the merger. The unidirectionality can be explained regarding the differences in general markedness (Maddieson, 1984; Greenberg, 1987). Retroflexes are produced by curling the tongue tip backwards while creating a stricture (fricatives) and a complete blockage of airflow followed by the narrow construction (affricates). In comparison to alveolars, which do not possess tongue retroflexion, the required articulatory complexity and efforts are generally higher.
The current study further expands on the articulatory complexity by finding an additional lip rounding gesture for retroflexes, adding to the existing gestural efforts employed inside the oral cavity. The general tendency towards the reduction of articulatorily complex or “marked” segments may have motivated the sibilant merger, among other factors, such as contact with local Chinese varieties (Ing, 1984; Kubler, 1985).
The findings also align with existing studies in multiple articulatory configurations where, phonetically, the same contrast and acoustic properties can be achieved through different articulatory gestures, such as the case of bunched vs. retroflex /ɹ/ (Delattre & Freeman, 1968). In our case, the Mandarin sibilant contrast can be achieved through differences in tongue position, lip rounding, or both.
4.2. Implications for sound change
It is well understood that the dynamic nature of cue relationships can lead to sound change over time, as seen in recent cases of tonogenesis in Afrikaans (Coetzee et al., 2018; Pfiffner, 2021, 2025) and Seoul Korean (Kang & Han, 2013; Kang, 2014; Bang et al., 2018), or in registrogenesis in Eastern Cham (Brunelle, 2009) and Chru (Brunelle et al., 2020). In these cases, a word-initial voicing contrast is changing towards a primary cue of f0 (tonogenesis) or combinations of f0, F1, and phonation type (registrogenesis; Brunelle & Kirby, 2016). Several recent studies have found that changes in cue strength of the same contrast correlate with age, where younger speakers’ perception and/or production of the same contrast is significantly different from older speakers’ perception and/or production, indicating a population-level sound change (Gao, 2016; Bang et al., 2018; Brunelle et al., 2020; Jiang et al., 2020).
Within accounts that relate sound change to the distribution of phonetic cues, the addition of visual correlates in production may increase the likelihood for listeners to fail in perceptual reanalysis and not recover the speaker’s intended category. For example, the listener-oriented model proposed by Ohala (1981, 1983, 1989, 1993) views the exchange of speech as a series of mappings of acoustic cues onto phonological categories. Sound change arises from the listener’s “innocent misinterpretation” along the noisy channel between the speaker and the listener, which can introduce biases against certain acoustic cues and make listeners match acoustic cues onto other phonological categories that were not originally intended by the speaker. Later studies found that the reanalysis is performed not only on the phonological categories but also gestures (e.g., French nasalized vowels, see Beddor, 2009, 2012; Carignan, 2018) and in the visual domain of cues (e.g., Havenhill, 2019; Havenhill & Do, 2018).
The articulatory variation may take on different visual properties, and this variability in visual cues may differentially affect variants in a sound change in progress, such as the case of [θ] and [f] mentioned above (McGuire & Babel, 2012), which could be why the sound change progresses in this direction. Similarly, in regards to the length of resonating cavities, Havenhill (2019) and Havenhill and Do (2018) found that there are at least two articulatory patterns for back vowels used by American English speakers who partake in the Northern Cities Vowel Shift (NCVS); although lip rounding and tongue backing were found to have similar effects on F2, variants of /ɔ/ that were produced with more tongue backing but less lip rounding were both visually and perceptually more similar to variants of /ɑ/.
The source of such individual variability may lie in anatomical differences and speaker’s active control for articulatory configurations to maintain constant acoustic output. Brunner et al. (2009) found that individual differences in palatal morphology can shape speaker-specific tongue configurations. In an electropalatographic and acoustic study of 32 speakers, those with flatter palates showed reduced variability in tongue height, likely because a given change in tongue height produces a larger change in constriction area for flatter palates than for more domed palates. On the other hand, acoustic variability did not increase relative to speakers with more domed palates. This suggests that anatomical differences in palatal shape may have pushed speakers toward different lingual configurations, even when they converged on similar acoustic targets.
In the case of Mandarin, the contrast between alveolars and retroflexes is primarily in CoG, which can be manipulated by tongue position and/or the degree of lip rounding, both of which affect the length of the front cavity and therefore the resonant frequencies. Lip rounding provides both a visual cue and an additional source of variability in acoustics that may influence the possible trajectory of change. Chiu and colleagues (2020) found that there is a continuum of articulatory gestures for Mandarin retroflexes, ranging from tongue retroflexion to tongue fronting. This partially resulted in variable acoustic cues in the frication portion of the sibilants. Such variability in tongue gestures, along with the lip rounding found in the current study, may also arise due to anatomical differences in tongue shape or the height or shape of the alveolar/post-alveolar regions. The variability of retroflexion is then further coupled with the variability of lip gestures that we report among retroflexes, leading to an even less robust contrast and resulting in the total cues available to listeners being less stable across the board.
Therefore, considering the additional individual variation in audiovisual correlates observed in the production study, it is possible for us to hypothesize potential pathways to prevent or induce this merger in perception. Mandarin speakers may fall into any of the four categories of articulatory gestures and resulting cue weighting relationship for retroflex sibilants: One may employ both lip rounding and tongue retroflexion, leading to comparable cue weights between different acoustic cues and between acoustic and visual cues. As rounding cues are discovered to be in an enhancing relationship with acoustic cues, both primary and secondary cues contribute to the contrast. Therefore, this articulatory configuration may be the most robust and least prone to sound change, as it offers the clearest distributional separation along multiple cues for listeners. Alternatively, one may primarily rely on tongue retroflexion and employ little rounding, resulting in little to no visual and less acoustic contrast. As tongue retroflexion tends to have more variable articulation (see, e.g., Tabain, 2009, on variability in the place of articulation of retroflexes mediated by stress in Arrente), this second configuration is more likely to be incorrectly recovered by perceptual reanalysis in perception, possibly most prone to undergo a merger. A third option would be for speakers to resort to rounding as a primary cue; however, this would involve a reweighting of cues that is not currently supported by our overall production results, and, at the same time, the availability of visual cues is not always consistent and can be influenced by the noisy channel along the visual dimension, such as attention level (e.g., Andersen et al., 2009; Tiippana et al., 2004) or the medium of speech exchanges (i.e., auditory or audiovisual), which may also make rounding-only configurations more susceptible to merger. Finally, if both visual and acoustic cues are no longer available, younger generations or language acquirers will merge the two categories.
5. Conclusion
This paper investigates visual properties and individual variation in the production of Mandarin sibilants and how these factors may provide evidence to the development of the Mandarin sibilant merger. An audiovisual production experiment showed differences in lip gestures and overall roundness between alveolars and retroflexes, including larger vertical lip apertures and overall higher aperture variability for retroflexes in comparison to alveolar sibilants. The distinction was also found to surface during the vowel phase. Our analysis shows that the acoustic properties and lip gestures are strongly correlated with each other, indicating an enhancing relationship between audio and visual dimensions. We also employed a classification method (conditional random forest) to assess the contributions of factors that would introduce collinearity issues in regression models, which will be helpful for follow-up studies on how tongue retroflexion interacts with rounding and other visual characteristics, or for other studies where multiple predictors are correlated. These findings suggest four different possible configurations of articulatory gestures, each of which may have different visual and acoustic properties that are more or less likely to result in a maintenance of the contrast or alternatively, a merger.
This study has several limitations and open avenues for future research. The variability in roundness and less variable CoG across the sample imply that various configurations of rounding may have different effects on the spectral properties. How exactly apertures and roundness interact with tongue gestures and together impact the properties is unanswered. Additionally, the hypotheses that were supported in the production experiment have not been tested in perception, especially the cross-modal cue integration and weighting between audio and visual cues. In other words, the production experiment has been unable to establish whether the visual cues of aperture and roundness influence perception in the Mandarin sibilant merger. In future work, we will create a perception experiment with audiovisual continua to investigate cross-modality cue weighting between audio and visual cues, which will allow us to fully account for the directionality of the merger in both production and perception.
One final consideration is that this merger is stigmatized and associated with Southern China/Taiwan speakers (see Lee-Kim & Chou, 2022, and references therein), and thus the role of social factors and attitudes towards this variable will undoubtedly affect the sound change process (Lee-Kim & Tung, 2025; Baranowski, 2013; Regan, 2020; Duncan, 2022). This may play a large role in determining when and how the merger is actuated, whether it can be completed, what the ultimate status will be, and act as an additional factor to the phonological and phonetic motivations considered here.
Appendix A: The wordlist for the production experiment
As mentioned in Section 2.1, /i/ in this table is an apical vowel. Tones are denoted following Pinyin conventions, i.e., high level T1, rising T2, low dipping T3, falling T4.
| Item | Group | Target | Glosses | Onset | Vowel | Tone | Pinyin | IPA | Chr per M* | Percentile |
| 1 | alveolar | 资 | Capital | ts | i | 1 | zi1 | tsi1 | 168.23 | 87.3 |
| 2 | 紫 | Purple | ts | i | 3 | zi3 | tsi3 | 15.82 | 64.2 | |
| 3 | 自 | Self | ts | i | 4 | zi4 | tsi4 | 2107.08 | 98.6 | |
| 4 | 玼 | Flaw of jade | tsʰ | i | 1 | ci1 | tsʰi1 | 0 | 0 | |
| 5 | 词 | Word | tsʰ | i | 2 | ci2 | tsʰi2 | 153.24 | 86.3 | |
| 6 | 此 | This | tsʰ | i | 3 | ci3 | tsʰi3 | 598.83 | 95.3 | |
| 7 | 次 | Time+ | tsʰ | i | 4 | ci4 | tsʰi4 | 1240.53 | 97.6 | |
| 8 | 丝 | Silk | s | i | 1 | si1 | si1 | 142.99 | 85.8 | |
| 9 | 死 | Death | s | i | 3 | si3 | si3 | 1521.23 | 98.1 | |
| 10 | 四 | Four | s | i | 4 | si4 | si4 | 282.7 | 91.5 | |
| 11 | retroflex | 支 | Branch | ʈʂ | i | 1 | zhi1 | ʈʂi1 | 249.72 | 90.5 |
| 12 | 直 | Straight | ʈʂ | i | 2 | zhi2 | ʈʂi2 | 985.52 | 97 | |
| 13 | 纸 | Paper | ʈʂ | i | 3 | zhi3 | ʈʂi3 | 111.46 | 83.7 | |
| 14 | 至 | Till | ʈʂ | i | 4 | zhi4 | ʈʂi4 | 392.69 | 93.6 | |
| 15 | 吃 | Eat | ʈʂʰ | i | 1 | chi1 | ʈʂʰi1 | 722.21 | 96 | |
| 16 | 池 | Pond | ʈʂʰ | i | 2 | chi2 | ʈʂʰi2 | 50.64 | 76.4 | |
| 17 | 尺 | Ruler | ʈʂʰ | i | 3 | chi3 | ʈʂʰi3 | 56.6 | 77.6 | |
| 18 | 赤 | Red | ʈʂʰ | i | 4 | chi4 | ʈʂʰi4 | 14.82 | 63.6 | |
| 19 | 湿 | Moist | ʂ | i | 1 | shi1 | ʂi1 | 27.35 | 69.7 | |
| 20 | 时 | Time | ʂ | i | 2 | shi2 | ʂi2 | 3471.55 | 99.2 | |
| 21 | 使 | Make | ʂ | i | 3 | shi3 | ʂi3 | 361.07 | 93.1 | |
| 22 | 市 | Market | ʂ | i | 4 | shi4 | ʂi4 | 223.16 | 89.5 | |
| 23 | palato-alveolar | 积 | Accumulate | tɕ | i | 1 | ji1 | tɕi1 | 34.54 | 72.2 |
| 24 | 级 | Level | tɕ | i | 2 | ji2 | tɕi2 | 211.2 | 89.1 | |
| 25 | 几 | Several | tɕ | i | 3 | ji3 | tɕi3 | 622.36 | 95.5 | |
| 26 | 记 | Remember | tɕ | i | 4 | ji4 | tɕi4 | 832.77 | 96.4 | |
| 27 | 七 | Seven | tɕʰ | i | 1 | qi1 | tɕʰi1 | 87.4 | 81.7 | |
| 28 | 骑 | Ride | tɕʰ | i | 2 | qi2 | tɕʰi2 | 72.42 | 80 | |
| 29 | 起 | Rise | tɕʰ | i | 3 | qi3 | tɕʰi3 | 2723.57 | 99 | |
| 30 | 气 | Gas | tɕʰ | i | 4 | qi4 | tɕʰi4 | 604.9 | 95.4 | |
| 31 | 西 | West | ɕ | i | 1 | xi1 | ɕi1 | 1260.09 | 97.6 | |
| 32 | 袭 | Attack | ɕ | i | 2 | xi2 | ɕi2 | 69.02 | 79.4 | |
| 33 | 洗 | Wash | ɕ | i | 3 | xi3 | ɕi3 | 175.49 | 87.8 | |
| 34 | 细 | Thin | ɕ | i | 4 | xi4 | ɕi4 | 132.23 | 85.2 | |
| 35 | fillers | 发 | Send | f | a | 1 | fa1 | fa1 | 1870.79 | 98.6 |
| 36 | 蓝 | Blue | l | an | 2 | lan2 | lan2 | 76.09 | 80.4 | |
| 37 | 无 | Null | u | 2 | wu2 | u2 | 1086.82 | 97.2 | ||
| 38 | 饿 | Hungry | ɤ | 4 | e4 | ɤ4 | 71.22 | 79.9 | ||
| 39 | 有 | Have | j | oʊ | 3 | qi1 | joʊ3 | 11392.37 | 99.8 | |
| 40 | 按 | Press | an | 4 | an4 | an4 | 145.83 | 85.9 | ||
| 41 | 慢 | Slow | m | an | 3 | qi3 | tɕʰi3 | 188.08 | 88.3 | |
| 42 | 看 | Look | k | an | 4 | kan4 | kʰan4 | 4568.08 | 99.4 | |
| 43 | 欧 | Europe | oʊ | 1 | ou1 | oʊ1 | 78.5 | 80.7 | ||
| 44 | 半 | Half | b | an | 4 | ban4 | ban4 | 232.92 | 89.8 |
-
* Characters per million.
+ Time, as in “the first time”.
Appendix B. Descriptive statistics for the frequency of individual members of the contrast
| Onset | Place | n | Mean frequency (SD) | Mean log frequency* (SD) | Mean characters/million (SD) |
| s | Alveolar | 3 | 30,399 (35,534) | 9.82 (1.22) | 649 (759) |
| ʂ | Retroflex | 4 | 47,814 (76,799) | 9.54 (1.99) | 1,021 (1,640) |
| ts | Alveolar | 3 | 35,773 (54,611) | 9.03 (2.45) | 764 (1,166) |
| ʈʂ | Retroflex | 4 | 20,369 (18,018) | 9.62 (0.91) | 435 (385) |
| tsʰ | Alveolar | 4 | 23,334 (26,057) | 7.52 (5.09) | 498 (556) |
| ʈʂʰ | Retroflex | 4 | 9,886 (15,985) | 8.16 (1.63) | 211 (341) |
* Log(frequency+1).
Ethics and consent
The study was approved by University of California-Berkeley’s Committee for Protection of Human Subjects under protocol ID 2024-03-16186. All participants gave their informed consent to take part in this study.
Acknowledgements
We thank all speakers who participated in this study, as well as the Department of Linguistics at University of California-Berkeley for the funding to pay participants. We are grateful to Keith Johnson, Yao Yao, Megha Sundara, University of California-Los Angeles phonetics seminar attendees, University of California-Berkeley PhonLab members, Associate Editor Bettina Braun, two anonymous reviewers, and audiences at NWAV51 and LSA2024 for their valuable comments and suggestions on the paper.
Competing interests
The authors have no competing interests to declare.
Authors’ contributions
Baichen Du was responsible for conceptualization, study design, methodology, stimuli creation, data collection, acoustic and visual data analysis, writing of the original manuscript, reviewing, and editing.
Alexandra Pfiffner was responsible for conceptualization, study design, methodology, writing, reviewing, and editing.
Notes
- The sibilant merger has traditionally been associated with Taiwan Mandarin, but there is variation present in both Taiwan speakers and Mainland China speakers (Hauser, 2023, and references therein). See, e.g., Chang (2012) showing overlapping alveolars and retroflexes in Beijing speakers, and Chiu et al. (2020) showing Taiwan speakers ranging from merged to unmerged. [^]
- Pinyin does show the distinction between the classes of sibilants, but there were only two characters that had multiple readings, i.e., 骑 (Pinyin: qi2/ji4, IPA: /tɕʰi35/ or /tɕi51/), and 几 (Pinyin: ji3/ji1, IPA: /tɕi214/, /tɕi55/). In the former case, the ji4 reading only appears in the word 铁骑 calvary and is considered an antiquated reading. In the second case, the multiple reading problem only concerns tone contour, not necessarily the contrast at the onset position. Therefore, we consider the influence of Pinyin to be minimal. [^]
- The character was still selected due to a lexical gap. For additional context, it is a word used more commonly in old Chinese literature. It also has largely available orthographic and phonological neighbors, such as 此 (“this,” tone 3), 呲 (“grimace,” same tone but colloquial), 疵 (“flaw,” which only appears in the second character position in a limited number of words, such as “瑕疵”, also meaning “flaw”), which have frequency values as follows, 此: Count = 28050, CHR/million = 599; 呲: Count = 5, CHR/million = 0.1; 疵: Count = 130, CHR/million = 2.78. The fact that we supplemented the romanization (i.e., Pinyin) may also help the character recognition. [^]
- It should be noted that preceding consonants do have an effect on the following vowels. In our case, the vowel following retroflexes can be different from those following palatoalveolars and alveolars; for example, there are some studies that categorize this class of vowel as [ɨ] (Chiu et al., 2020; though see discussion in Wu & Shih, 2009). This is due to the effect of tongue retroflexion and the more centered tongue position while vowels are being articulated. However, both variants are unrounded, and studies have suggested that the mental representation of the two classes of vowels are the same for Mandarin speakers (Wu & Shih, 2009). Previous studies also treated these vowels to be in minimal pairs (e.g., Lee-Kim & Chou, 2022); we do not expect this to affect the visual property of preceding onsets. In any case, all onset sibilants have the same place of articulation as their following vocalic segments. [^]
- The four parameters include fx (focal horizontal length), fy (focal vertical length), cx, (center horizontal length), and cy (center vertical length). See the OpenCV Camera Calibration documentation for further details: https://docs.opencv.org/4.x/dc/dbb/tutorial_py_calibration.html. [^]
- The R code for our analysis can be found at: https://osf.io/473fq/overview?view_only=14838538df814b30b2dcd7def70985bf. [^]
- We use whole-syllable averages for roundness because the GAMM model shows that visible rounding differences also emerge in the nucleus. We interpret this as evidence that the visible lip-rounding gesture is temporally spread relative to the primary acoustic correlate, likely because visible rounding reflects a sequence of protrusion and rounding. By contrast, CoG remains measured from the consonant/frication interval. [^]
- GAMM formula: roundness ~ property * phase * manner + s(timepoint, by = cond, bs = “tp”, k = 20) + s(timepoint, participant, bs = “fs”, m = 1, k = 15) + s(item, bs = “re”). See the supplementary HTML file at footnote 5 for model details and the accompanying tables for parametric and smooth terms. [^]
- Although rare, it should be noted that the neutralization between dentals/alveolars and retroflexes does not always happen in this direction, such as the merger between implosives that resulted in the relatively high frequency of retroflexes in Sindhi and Indo-Aryan languages in general (Hussain & Mielke, 2023). [^]
References
Andersen, T. S., Tiippana, K., Laarni, J., Kojo, I., & Sams, M. (2009). The role of visual spatial attention in audiovisual speech perception. Speech Communication, 51(2), 184–193. http://doi.org/10.1016/j.specom.2008.07.004
Baltrušaitis, T. (2019). OpenFace (Github repository). Available from https://github.com/TadasBaltrusaitis/OpenFace
Baltrušaitis, T., Zadeh, A., Lim, Y. C., & Morency, L.-P. (2018). OpenFace 2.0: Facial behavior analysis toolkit. Proceedings of the 13th IEEE International Conference on Automatic Face & Gesture, 59–66. http://doi.org/10.1109/FG.2018.00019
Bang, H.-Y. B., Sonderegger, M., Kang, Y., Clayards, M., & Yoon, T.-J. (2018). The emergence, progress, and impact of sound change in progress in Seoul Korean: Implications for mechanisms of tonogenesis. Journal of Phonetics, 66, 120–144. http://doi.org/10.1016/j.wocn.2017.09.005
Baranowski, M. (2013). Ethnicity and sound change: African American English in Charleston, SC. University of Pennsylvania Working Papers in Linguistics, 19(2), 2. https://repository.upenn.edu/handle/20.500.14332/44925
Beddor, P. S. (2009). A coarticulatory path to sound change. Language, 85(4), 785–821. https://www.jstor.org/stable/40492954
Beddor, P. S. (2012). Perception grammars and sound change. In M.-J. Solé & D. Recasens (Eds.), The Initiation of Sound Change: Perception, production, and social factors (pp. 37–56). John Benjamins. http://doi.org/10.1075/cilt.323.06bed
Berg, B. G. (1989). Analysis of weights in multiple observation tasks. Journal of the Acoustical Society of America, 86(5), 1743–1746. http://doi.org/10.1121/1.398605
Bladon, R. A. W., & Nolan, F. J. (1977). A video-fluorographic investigation of tip and blade alveolars in English. Journal of Phonetics, 5(2), 185–193. http://doi.org/10.1016/S0095-4470(19)31128-3
Boersma, P., & Weenink, D. (2023). Praat: Doing phonetics by computer (Version 6.3.02) [Computer program]. https://www.fon.hum.uva.nl/praat/
Broś, K., & Krause, P. A. (2024). Stop lenition in Canary Islands Spanish – a motion capture study. Laboratory Phonology, 15(1). http://doi.org/10.16995/labphon.9934
Brunelle, M. (2009). Contact-induced change? Register in three Cham dialects. Journal of the Southeast Asian Linguistics Society, 2, 1–22.
Brunelle, M., & Kirby, J. (2016). Tone and phonation in Southeast Asian Languages. Language and Linguistics Compass, 10(4), 191–207. http://doi.org/10.1111/lnc3.12182
Brunelle, M., Tấn, T., Kirby, J., & Giang, Đ. (2020). Transphonologization of voicing in Chru: Studies in production and perception. Laboratory Phonology, 11(1), 15. http://doi.org/10.5334/labphon.278
Brunner, J., Fuchs, S., & Perrier, P. (2009). On the relationship between palate shape and articulatory behavior. Journal of the Acoustical Society of America, 125(6), 3936–3949. http://doi.org/10.1121/1.3125313
Cai, Q., & Brysbaert, M. (2010). SUBTLEX-CH: Chinese word and character frequencies based on film subtitles. Plos ONE, 5(6), e10729. http://doi.org/10.1371/journal.pone.0010729
Calvert, G. A., Bullmore, E. T., Brammer, M. J., Campbell, R., Williams, S. C., McGuire, P. K., … & David, A. S. (1997). Activation of auditory cortex during silent lipreading. Science, 276(5312), 593–596. http://doi.org/10.1126/science.276.5312.593
Calvert, G. A., Hansen, P. C., Iversen, S. D., & Brammer, M. J. (2001). Detection of audio-visual integration sites in humans by application of electrophysiological criteria to the BOLD effect. Neuroimage, 14(2), 427–438. http://doi.org/10.1006/nimg.2001.0812
Carignan, C. (2018). Using naïve listener imitations of native speaker productions to investigate mechanisms of listener-based sound change. Laboratory Phonology, 9(1), 18. http://doi.org/10.5334/labphon.136
Chang, Y.-H. (2012). Variability in cross-dialectal production and perception of contrasting phonemes: The case of the alveolar-retroflex contrast in Beijing and Taiwan Mandarin(Publication no. 3632865). [Doctoral dissertation, University of Illinois, Urbana-Champagne]. ProQuest Dissertations and Theses Global.
Chang Y.-H., Shih C. (2015). Place contrast enhancement: The case of the alveolar and retroflex sibilant production in two dialects of Mandarin. Journal of Phonetics, 50, 52–66. http://doi.org/10.1016/j.wocn.2015.02.001
Chang Y.-H., Shih C., Allen J. B. (2013). Dialectal variation in the perception of phonological contrasts. In Lee W. S. (Ed.), Proceedings of the International Conference on Phonetics of the Languages in China (pp. 115–118). City University of Hong Kong.
Cheng, R. L. (1985). A comparison of Taiwanese, Taiwan Mandarin, and Peking Mandarin. Language, 61(2), 352–377. http://doi.org/10.2307/414149
Chiu, C., Wei, P.-C., Noguchi, M., & Yamane, N. (2020). Sibilant fricative merging in Taiwan Mandarin: An investigation of tongue postures using ultrasound imaging. Language and Speech, 63(4), 877–897. http://doi.org/10.1177/0023830919896386
Christensen, L. A., & Humes, L. E. (1996). Identification of multidimensional complex sounds having parallel dimension structure. Journal of the Acoustical Society of America, 99(4), 2307–2315. http://doi.org/10.1121/1.415418
Coetzee, A. W., Beddor, P. S., Shedden, K., Styler, W., & Wissing, D. (2018). Plosive voicing in Afrikaans: Differential cue weighting and tonogenesis. Journal of Phonetics, 66, 185–216. http://doi.org/10.1016/j.wocn.2017.09.009
Cohen, J. (1988). The effect size index: d. Routledge. http://doi.org/10.4324/9780203771587
De Decker, P. M., & Nycz, J. R. (2012). Are tense [æ]s really tense? The mapping between articulation and acoustics. Lingua, 122(7), 810–821. http://doi.org/10.1016/j.lingua.2012.01.003
Delattre, P., & Freeman, D. C. (1968). A dialect study of American r’s by x-ray motion picture. Linguistics, 6(44), 29–68. http://doi.org/10.1515/ling.1968.6.44.29
Doherty, K. A., & Turner, C. W. (1996). Use of a correlational method to estimate a listener’s weighting function for speech. Journal of the Acoustical Society of America, 100(6), 3769–3773. http://doi.org/10.1121/1.417336
Duanmu, S. (2007). The phonology of standard Chinese. Oxford University Press. http://doi.org/10.1093/oso/9780199215782.001.0001
Dubois, S., & Horvath, B. M. (1998). Let’s tink about dat: Interdental fricatives in Cajun English. Language Variation and Change, 10(3), 245–261. http://doi.org/10.1017/S0954394500001320
Duncan, D. (2022). Merger reversal in St. Louis: Implementation and implications. Journal of English Linguistics, 50(1), 72–105. http://doi.org/10.1177/00754242221083648
Fant, G. (1960). Acoustic Theory of Speech Production. Mouton.
Fant, G. (1971). Acoustic theory of speech production: with calculations based on X-ray studies of Russian articulations (No. 2). Walter de Gruyter. http://doi.org/10.1515/9783110873429
Fitch, H. L., Halwes, T., Erickson, D. M., & Liberman, A. M. (1980). Perceptual equivalence of two acoustic cues for stop-consonant manner. Perception & Psychophysics, 27, 343–350. http://doi.org/10.3758/BF03206123
Francis, A. L., Baldwin, K., & Nusbaum, H. C. (2000). Effects of training on attention to acoustic cues. Perception & Psychophysics, 62, 1668–1680. http://doi.org/10.3758/BF03212164
Fraser, C. S. (2001). Photogrammetric camera component calibration: A review of analytical techniques. In A. Gruen & T. S. Huang (Eds.), Calibration and orientation of cameras in computer vision (pp. 95–121). Springer. http://doi.org/10.1007/978-3-662-04567-1_4
Gao, J. (2016). Sociolinguistic motivations in sound change: On-going loss of low tone breathy voice in Shanghai Chinese. Papers in Historical Phonology, 1, 166–186. http://doi.org/10.2218/pihph.1.2016.1698
Gick, B., & Derrick, D. (2009). Aero-tactile integration in speech perception. Nature, 462(7272), 502–504. http://doi.org/10.1038/nature08572
Grant, K. W., & Seitz, P.-F. (2000). The use of visible speech cues for improving auditory detection of spoken sentences. Journal of the Acoustical Society of America, 108(3), 1197–1208. http://doi.org/10.1121/1.1288668
Greenberg, J. H. (1987). The present status of markedness theory: A reply to Scheffler. Journal of Anthropological Research, 43(4), 367–374. http://doi.org/10.1086/jar.43.4.3630547
Greenlee, M. (1980). Learning the phonetic cues to the voiced-voiceless distinction: A comparison of child and adult speech perception. Journal of Child Language, 7(3), 459–468. http://doi.org/10.1017/S0305000900002786
Guion, S. G. (1998). The role of perception in the sound change of velar palatalization. Phonetica, 55, 18–52. http://doi.org/10.1159/000028423
Hauser, I. (2023). Differential cue weighting in Mandarin sibilant production. Language and Speech, 00238309231152495. http://doi.org/10.1177/00238309231152495
Havenhill, J., & Do, Y. (2018). Visual speech perception cues constrain patterns of articulatory variation and sound change. Frontiers in Psychology, 9, 728. http://doi.org/10.3389/fpsyg.2018.00728
Havenhill, J. E. (2019). Constraints on articulatory variability: Audiovisual perception of lip rounding (Publication No. 13812740). [Doctoral dissertation, Georgetown University]. ProQuest Dissertations & Theses Global.
Holt, L. L., & Lotto, A. J. (2006). Cue weighting in auditory categorization: Implications for first and second language acquisition. Journal of the Acoustical Society of America, 119(5), 3059–3071. http://doi.org/10.1121/1.2188377
Hothorn, T., Hornik, K., & Zeileis, A. (2006). Unbiased recursive partitioning: A conditional inference framework. Journal of Computational and Graphical Statistics, 15(3), 651–674. http://doi.org/10.1198/106186006X133933
Hussain, Q., & Mielke, J. (2023). Place typology and evolution of implosives in Indo-Aryan languages. Linguistic Typology, 27(2), 429–453. http://doi.org/10.1515/lingty-2022-0040
Ing, R. O. (1984). Issues on the pronunciations of Mandarin. The World of Chinese Language, 35, 6–16.
Iverson, P., Hazan, V., & Bannister, K. (2005). Phonetic training with acoustic cue manipulations: A comparison of methods for teaching English /r/-/l/ to Japanese adults. Journal of the Acoustical Society of America, 118, 3267–3278. http://doi.org/10.1121/1.2062307
Jiang, B., Clayards, M., & Sonderegger, M. (2020). Individual and dialect differences in perceiving multiple cues: A tonal register contrast in two Chinese Wu dialects. Laboratory Phonology, 11(1), 11. http://doi.org/10.5334/labphon.266
Jones, M. J. (2002). More on the “instability” of interdental fricatives: Gothic þliuhan ‘flee’ and Old English flēon ‘flee’ revisited. WORD, 53(1), 1–8. http://doi.org/10.1080/00437956.2002.11432521
Kang, Y. (2014). Voice onset time merger and development of tonal contrast in Seoul Korean stops: A corpus study. Journal of Phonetics, 45, 76–90. http://doi.org/10.1016/j.wocn.2014.03.005
Kang, Y., & Han, S. (2013). Tonogenesis in early Contemporary Seoul Korean: A longitudinal case study. Lingua, 134, 62–74. http://doi.org/10.1016/j.lingua.2013.06.002
Kingston, J., Diehl, R. L., Kirk, C. J., & Castleman, W. A. (2008). On the internal perceptual structure of distinctive features: The [voice] contrast. Journal of Phonetics, 36(1), 28–54. http://doi.org/10.1016/j.wocn.2007.02.001
Kjellmer, G. (1995). Unstable fricatives: On Gothic pliuhan and Old English flēon. WORD, 46(2), 207–223. http://doi.org/10.1080/00437956.1995.11435942
Krause, P. A., Kay, C. A., & Kawamoto, A. H. (2020). Automatic motion tracking of lips using digital video and OpenFace 2.0. Laboratory Phonology, 11(1), 9. http://doi.org/10.5334/labphon.232
Krause, P. A., Pili, R. J., & Hunt, E. (2024). A process for measuring lip kinematics using participants’ webcams during linguistic experiments conducted online. Laboratory Phonology, 15(1). http://doi.org/10.16995/labphon.10483
Krause, S. E. (1982). Developmental use of vowel duration as a cue to postvocalic stop consonant voicing. Journal of Speech, Language, and Hearing Research, 25, 388–393. http://doi.org/10.1044/jshr.2503.388
Kubler, C. C. (1985). The influence of Southern Min on the Mandarin of Taiwan. Anthropological Linguistics, 27(2), 156–176. https://www.jstor.org/stable/30028064
Labov, W. (1990). The intersection of sex and social class in the course of linguistic change. Language Variation and Change, 2(2), 205–254. http://doi.org/10.1017/S0954394500000338
Labov, W. (2001). Applying our knowledge of African American English to the problem of raising reading levels in inner-city schools. In S. L. Lanehart (Ed.), Sociocultural and historical contexts of African American English (pp. 299–317). John Benjamins. http://doi.org/10.1075/veaw.g27.19lab
Ladefoged, P. (1980). Articulatory parameters. Language and Speech, 23(1), 25–30. http://doi.org/10.1177/002383098002300103
Lee-Kim, S.-I., & Chou, Y.-C. (2022). Unmerging the sibilant merger among speakers of Taiwan Mandarin. Laboratory Phonology, 13(1), 10. http://doi.org/10.16995/labphon.6446
Lee-Kim, S. I., & Tung, H. Y. (2025). Linguistic experience and social factors in speech perception: The case of merged speakers of Mandarin sibilants. Laboratory Phonology, 16(1). http://doi.org/10.16995/labphon.15384
Lin, Y. H. (2007). The sounds of Chinese. Cambridge University Press.
Lisker, L. (1986). “Voicing” in English: A catalogue of acoustic features signaling /b/ versus /p/ in trochees. Language and Speech, 29, 3–11. http://doi.org/10.1177/002383098602900102
Lu, X. X. (2018). A review of solutions for perspective-n-point problem in camera pose estimation. Journal of Physics: Conference Series, 1087, 052009. http://doi.org/10.1088/1742-6596/1087/5/052009
Maddieson, I. (1984). Patterns of sounds. Cambridge University Press. http://doi.org/10.1017/CBO9780511753459
Marks, L. E. (2004). Cross-modal interaction in speeded classification. In G. Calvert, C. Spence, & B. Stein (Eds.), The handbook of multisensory processes (pp. 85–105). MIT Press. http://doi.org/10.7551/mitpress/3422.003.0009
Mayer, K. M., Yildiz, I. B., Macedonia, M., & von Kriegstein, K. (2015). Visual and motor cortices differentially support the translation of foreign language words. Current Biology, 25(4), 530–535. http://doi.org/10.1016/j.cub.2014.11.068
McGuire, G., & Babel, M. (2012). A cross-modal account for synchronic and diachronic patterns of/f/and/θ/in English. Laboratory Phonology, 3(2), 251–272. http://doi.org/10.1515/lp-2012-0014
McGurk, H., & MacDonald, J. (1976). Hearing lips and seeing voices. Nature, 264(5588), 746–748. http://doi.org/10.1038/264746a0
Ménard, L., Cathiard, M.-A., Troille, E., & Giroux, M. (2016). Effects of congenital visual deprivation on the auditory perception of anticipatory labial coarticulation. Folia Phoniatrica et Logopaedica, 67(2), 83–89. http://doi.org/10.1159/000434719
Montero, R.S., & Bribiesca, E. (2009). State of the art of compactness and circularity measures. International Mathematical Forum, 4, 1305–1335.
Musacchia, G., Sams, M., Nicol, T., & Kraus, N. (2006). Seeing speech affects acoustic information processing in the human brainstem. Experimental Brain Research, 168, 1–10. http://doi.org/10.1007/s00221-005-0071-5
Nittrouer, S. (2005). Age-related differences in weighting and masking of two cues to word-final stop voicing in noise. Journal of the Acoustical Society of America, 118(2), 1072–1088. http://doi.org/10.1121/1.1940508
Ohala, J. J. (1981). The listener as a source of sound change. In C. S. Masek, R. A. Hendrick, & M. F. Miller (Eds.), Papers from the parasession on language and behavior (pp. 178–203). Chicago Linguistic Society.
Ohala, J. J. (1983). The origin of sound patterns in vocal tract constraints. In P. F. MacNeilage (Ed.), The production of speech (pp. 189–216). Springer. http://doi.org/10.1007/978-1-4613-8202-7_9
Ohala, J. J. (1989). Sound change is drawn from a pool of synchronic variation. In L. E. Breivik & E. H. Jahr (Eds.), Language change: Contributions to the study of its causes (pp. 173–198). De Gruyter Mouton. http://doi.org/10.1515/9783110853063.173
Ohala, J. J. (1993). The phonetics of sound change. In C. Jones (Ed.), Historical linguistics: Problems and perspectives (pp. 237–278). Longman.
Pfiffner, A. M. (2021). Cue-based features: Modeling change and variation in the voicing contrasts of Minnesotan English, Afrikaans, and Dutch (Publication No. 28866585). [Doctoral dissertation, University of Minnesota]. ProQuest Dissertations and Theses Global.
Pfiffner, A. M. (2025). Women of all ages lead Tonogenesis in Afrikaans. Journal of Germanic Linguistics, 37(3), 270–286. http://doi.org/10.1017/S1470542725100056
Quan, L., & Lan, Z. (1999). Linear N-point camera pose determination. IEEE Transactions on Pattern Analysis and Machine Intelligence, 21(8), 774–780. http://doi.org/10.1109/34.784291
R Core Team. (2024). R: A language and environment for statistical computing [Computer software]. R Foundation for Statistical Computing, Vienna, Austria. https://www.R-project.org/
Rao, D., & Shaw, J. A. (2021). The role of gestural timing in non-coronal fricative mergers in Southwestern Mandarin: Acoustic evidence from a dialect island. Journal of Phonetics, 89, 101112. http://doi.org/10.1016/j.wocn.2021.101112
Regan, B. (2020). The split of a fricative merger due to dialect contact and societal changes: A sociophonetic study on Andalusian Spanish read-speech. Language Variation and Change, 32(2), 159–190. http://doi.org/10.1017/S0954394520000113
Repp, B. H. (1982). Phonetic trading relations and context effects: New experimental evidence for a speech mode of perception. Psychological Bulletin, 92, 81–110. http://doi.org/10.1037/0033-2909.92.1.81
Sekiyama, K., & Tohkura, Y. (1991). McGurk effect in non-English listeners: Few visual effects for Japanese subjects hearing Japanese syllables of high auditory intelligibility. Journal of the Acoustical Society of America, 90(4), 1797–1805. http://doi.org/10.1121/1.401660
Shadle, C. H. (2023). Alternatives to moments for characterizing fricatives: Reconsidering Forrest et al. (1988). Journal of the Acoustical Society of America, 153(2), 1412–1426. http://doi.org/10.1121/10.0017231
Shultz, A. A., Francis, A. L., & Llanos, F. (2012). Differential cue weighting in perception and production of consonant voicing. Journal of the Acoustical Society of America, 132(2), EL95–EL101. http://doi.org/10.1121/1.4736711
Smith, B. (2009). Dental fricatives and stops in Germanic: Deriving diachronic processes from synchronic variation. In M. Dufresne, F. Dupuis, & E. Vocaj (Eds.), Historical linguistics 2007: Selected papers from the 18th international conference on historical linguistics, Montreal, 6–11 August 2007 (pp. 19–36). John Benjamins. http://doi.org/10.1075/cilt.308.02smi
Sumby, W. H., & Pollack, I. (1954). Visual contribution to speech intelligibility in noise. Journal of the Acoustical Society of America, 26(2), 212–215. http://doi.org/10.1121/1.1907309
Tabain, M. (2009). An EPG study of the alveolar vs. retroflex apical contrast in Central Arrernte. Journal of Phonetics, 37(4), 486–501. http://doi.org/10.1016/j.wocn.2009.08.002
Tiippana, K., Andersen, T. S., & Sams, M. (2004). Visual attention modulates audiovisual speech perception. European Journal of Cognitive Psychology, 16(3), 457–472. http://doi.org/10.1080/09541440340000268
Trudgill, P. (1974). The social differentiation of English in Norwich. Cambridge University Press.
Tseng, S. C., Kuei, K., & Tsou, P. C. (2011). Acoustic characteristics of vowels and plosives/affricates of Mandarin-speaking hearing-impaired children. Clinical Linguistics & Phonetics, 25(9), 784–803. http://doi.org/10.3109/02699206.2011.565906
Wardrip-Fruin, C., & Peach, S. (1984). Developmental aspects of the perception of acoustic cues in determining the voicing feature of final stop consonants. Language and Speech, 27(4), 367–379. http://doi.org/10.1177/002383098402700407
Wood, S. (2023). mgcv: Mixed GAM computation vehicle with automatic smoothness estimation (Version 1.9-1) [Computer software, R package]. CRAN. https://cran.r-project.org/web/packages/mgcv/mgcv.pdf
Wu, C.-H., & Shih, C. (2009). Mandarin vowels revisited: Evidence from electromagnetic articulography. Berkeley Linguistics Society, 35(1), 329–340. http://doi.org/10.3765/bls.v35i1.3622
Yang, M., & Sundara, M. (2019). Cue-shifting between acoustic cues: Evidence for directional asymmetry. Journal of Phonetics, 75, 27–42. http://doi.org/10.1016/j.wocn.2019.04.002








