1. Introduction
Many languages exhibit longer vowel durations before voiced consonants than before voiceless consonants. A substantial body of work exists on the topic of voicing-conditioned vowel duration, including some of the other factors with which it interacts (e.g., Crystal & House, 1988; de Jong & Zawaydeh, 2002; Laeufer, 1992; Port, 1981; Umeda, 1975) and what might cause it (e.g., Chen, 1970; Chomsky & Halle, 1968; Coretta, 2019; Fowler, 1981, 1992; Halle & Stevens, 1967; Kluender et al., 1988; Öhman, 1967; Sanker, 2020).
Work on voicing-conditioned vowel duration often mentions that this effect has been observed in many languages and that there are some languages without a significant effect (e.g., Coretta, 2019; Durvasula & Luo, 2014; Sanker, 2020; Schwarz, 2024). However, previous work has not provided a systematic comparison across a large number of languages to examine patterns in the size of voicing-conditioned vowel duration effects. A typological overview of voicing-conditioned vowel duration can help establish its typical patterns, highlight gaps in our knowledge, and provide a stronger foundation for discussions of what underlies the effect.
This paper provides a summary of existing literature which reports vowel duration based on following voicing environment. It also adds data from languages with no prior literature describing the relationship between voicing and preceding vowel duration.
1.1. English versus other languages
Much of the work on voicing-conditioned vowel duration is based on English (e.g., Crystal & House, 1988; Fowler, 1981; Kluender et al., 1988; Luce & Charles-Luce, 1985; Port, 1981). However, English is not necessarily representative. The pervasiveness of voicing-conditioned vowel duration across different languages indicates that strong pressures drive it, but the differences between languages suggest that there is a learned component determining the size of the effect. That is, the vowel durations measured within individual languages might reflect learned language-specific duration targets rather than always reflecting purely phonetic factors.
Languages like English may exaggerate vowel duration as a cue to following consonant voicing (Chen, 1970; Cho, 2015; Solé, 2007). Thus, the voicing-conditioned vowel duration differences in English might be larger than what occurs naturally as an automatic consequence of voicing, in which case it is not an ideal language for testing universal effects of voicing on preceding vowels. Universal articulatory, aerodynamic, or perceptual effects of voicing on preceding vowels are likely to be clearest in a language without language-specific vowel duration targets specified as components of voicing. Typological patterns can provide evidence for evaluating the likely size of mechanically-driven differences in voicing-conditioned vowel duration and, consequently, which languages are most likely to be exhibiting voicing-conditioned vowel duration as it arises automatically rather than being specified as a representational target.
Some factors which interact with voicing-conditioned vowel duration have been observed primarily or solely in English. It is unclear whether the same effects would be observed in other languages. Some of the work which does compare such interactions across languages finds that different languages do not behave the same way, e.g., the ratio of vowel duration before voiced versus voiceless consonants when combined with other factors such as stress, speech rate, and inherent vowel duration (de Jong & Zawaydeh, 2002; Solé, 2007). Potential differences across languages are particularly important for patterns which have been used as evidence in evaluating what causes voicing-conditioned vowel duration, such as the size of the vowel duration difference (Chen, 1970; de Jong, 1991), speed of the articulatory transitions from the vowel to the following consonant (Chen, 1970; Raphael, 1975), and the existence of voicing-conditioned vowel duration differences even when there are intervening sonorants (Chen, 1970; Maddieson, 1999).
Previous work has noted cross-linguistic variation in the size of voicing-conditioned vowel duration effects. However, these descriptions often focus on a divide between languages with a large effect and languages with negligible effects. Some work discusses English as an outlier with a much larger effect of voicing on preceding vowel duration than other languages (e.g., Chen, 1970; Sóskuthy, 2013). Other studies focus on significance, noting that many languages have a significant effect (e.g., Coretta, 2019; Durvasula & Luo, 2014; Sanker, 2020; Schwarz, 2024). The same three languages are often cited as having a small or non-significant effect: Polish (Keating, 1980), Czech (Keating, 1985) and Arabic (de Jong & Zawaydeh, 2002; Flege & Port, 1981; Mitleb, 1984).
1.2. Potential causes
A range of different explanations have been proposed for what causes voicing-conditioned vowel duration differences, i.e., the phonetic pressures or biases which could contribute to how voicing-conditioned vowel duration differences arise. No conclusive evidence has led to a consensus about the cause. Some work proposes that there are multiple contributing mechanisms, which helps resolve these disagreements (e.g., Beguš, 2017; Coretta, 2020).
One proposed account for voicing-conditioned vowel duration is that it arises from compensatory timing in production (e.g., Coretta, 2019; Fowler, 1981; Lehiste, 1970) or in perception (e.g., Kluender et al., 1988), given that voiced obstruents are usually shorter than voiceless obstruents. Proposals that C/V ratio is a cue to voicing (e.g., Kohler, 1979; Port & Dalby, 1982) lend themselves to a similar account based on perceptual compensation. The compensatory timing account has also been explained gesturally: Segmental durations are the result of how gestures align, so a longer coda consonant might start earlier and result in a shorter preceding vowel due to greater overlap (e.g., Lehiste, 1970; Lindblom, 1967; Maddieson, 1999). In some languages, the combined V+C duration is nearly consistent with voiced and voiceless consonants (e.g., Dutch: Slis & Cohen, 1969). One criticism of compensation accounts is that syllable duration is not constant in languages like English, so the differences in vowel duration cannot simply reflect differences in gesture overlap or duration modulation to maintain consistent duration of the entire syllable (e.g., Chen, 1970; de Jong, 1991). However, compensatory timing does not need to target consistent syllable duration. Listeners normalize for speech rate, perceiving sounds as longer when surrounding sounds are shorter (e.g., Maslowski et al., 2019; Mitterer, 2018); similar relative perception of duration could result in voicing-conditioned vowel duration. Fowler (1992) argues against a perceptual compensation account because longer stop closures increase perceived vowel duration rather than decreasing it. That finding might be explained as the result of closure duration being a cue to consonant voicing (e.g., Port, 1979) and listeners identifying vowel duration relative to the duration expected in the perceived voicing context (e.g., Sanker, 2020).
A related possibility is that the glottal opening gesture might occur before the consonant constriction gesture. An early onset of the devoicing gesture would produce pre-aspiration or a devoiced portion at the end of the vowel (Hansson, 2003; Lisker, 1974), which could be reinterpreted as part of the consonant, thus shortening the vowel. Early devoicing might occur as overcompensation to avoid continuing voicing from a preceding vowel into an intervocalic consonant (cf. “shielding” as a source of postoralized nasal consonants, based on preserving the orality of a following vowel, Herbert, 1986, Ch. 7).
Another potential explanation is that more force is required to maintain the constriction in a voiceless obstruent, in which there is more pressure build-up, and that the high level of motor excitation necessary to produce such a closure results in the closure being produced more quickly (e.g., Chen, 1970; Öhman, 1967). The large vowel duration differences in languages like English could be a confound in measuring aspects of timing if they are specified as representational targets rather than directly reflecting aerodynamic requirements. The measurements given by Öhman (1967) come from a Swedish speaker, providing a useful comparison with a language other than English. Explanations based on how the transition into the obstruent is achieved seem to predict that effects of obstruent voicing should be local, i.e., should affect the immediately preceding segment but have little effect on more distant segments. In a sequence of vowel + sonorant + obstruent, articulation of the obstruent does not begin during the vowel (Chen, 1970; Maddieson, 1999), so the transition to the obstruent should primarily affect the adjacent sonorant and not the vowel preceding the sonorant. However, vowels in these environments can exhibit voicing-conditioned duration differences comparable to what is observed in vowels immediately preceding obstruents (e.g., English: Chen, 1970; Van Santen, 1992). One potential explanation is that vowel duration targets in these languages extended from vowels with immediately following obstruents to also include vowels with a sonorant intervening between them and a following obstruent.
A final possibility is that obstruent voicing requires more precision in planning the state of the larynx, so the vowel is longer to accommodate the necessary adjustments (e.g., Chomsky & Halle, 1968; Halle & Stevens, 1967). A limitation of this account is that direct articulatory observations do not provide evidence for laryngeal adjustments occurring in vowels before voiced obstruents (Lisker, 1974). Under this account, lengthening is specifically caused by voiced obstruents rather than voiced consonants more broadly, because voicing in sonorant consonants does not have the same aerodynamic challenges as voicing in obstruents. This lengthening pattern predicts that vowel duration before sonorants should be comparable to vowel duration before voiceless obstruents. However, vowel durations before sonorants are actually more similar to vowel durations before voiced obstruents (e.g., House & Fairbanks, 1953).
When a vowel duration difference exists due to phonetic pressures, it might become part of what is specified as a learned target for vowel duration in each environment. Thus, some languages might encode voicing-conditioned vowel duration as part of the representation (e.g., Chen, 1970; Fromkin, 1976; Solé, 2007).1 Once vowel durations are encoded as targets under speaker control rather than being produced automatically, they could subsequently shift in ways which are not driven by the original phonetic cause(s). One proposed piece of evidence for vowel duration being encoded as a representational target is that some languages have a language-specific consistent proportional size of the effect across different conditions which also impact vowel duration, e.g., stress (de Jong & Zawaydeh, 2002) and speech rate (Solé, 2007). The interpretation of this evidence might depend on what we propose as the phonetic cause of voicing-conditioned vowel duration. However, there are other languages in which the evidence for reflecting a phonological process is clear. In Breton and Friulian, vowel length differences conditioned by underlying voicing of the following consonants are categorically preserved with word-final consonants despite final devoicing (Baroni & Vanelli, 2000; Le Dû, 1986). Cho (2015) proposes three types of languages based on the size of the voicing effect and whether it scales proportionally when interacting with other factors that also impact vowel duration: (1) the mechanically motivated size, e.g., Arabic and Catalan; (2) representationally encoded large effect, e.g., English; (3) representationally encoded lack of effect, e.g., Polish and Czech. While this captures more patterns than a binary divide, it does not account for the wide range of gradient sizes which are observed across languages.
2. Data
One limitation for evaluating cross-linguistic variation in voicing-conditioned vowel duration is that data on voicing-conditioned vowel duration is only available for a relatively small number of languages. An even smaller subset of those languages are frequently cited as examples of how voicing-conditioned vowel duration behaves across languages.
This section presents a summary of vowel durations by voicing environment for 76 languages. Most of the data comes from previously published work, listed in Table 1. See the Supplementary Materials for a table of measurements from each study and information about the methods and materials used in each study.
Table 1: References for measurements from each language.
Not all of these studies were specifically aimed at testing voicing-conditioned vowel duration; some studies provide the relevant duration measurements even though they were aimed at addressing different questions. The cited works here are not an exhaustive list of all studies which report vowel duration compared across voicing environments. This phenomenon has been widely studied in a few languages (e.g., English); for these languages, a representative subset of papers are reported.
In addition to previously published duration measurements, this paper adds data based on recordings from the UCLA Phonetics Lab Archive (UCLA Department of Linguistics, 2007) for eight languages: Degema, Gujarati, Ibibio, Igbo, Kapampangan, Q’eqchi’, Tagalog, and Toda. They each have recordings from at least three speakers and at least eight total voiced-voiceless minimal pairs (16 tokens), with at least two pairs to represent each combination of position and manner of articulation being considered. While the amount of data is quite limited for some of these languages, they are included with the hope that having a small amount of data from additional languages is better than having none. To my knowledge, there are no prior studies describing vowel duration based on following consonant voicing in any of these languages.
Although the results here are grouped by language, the effect of voicing on vowel duration can vary within dialects of the same language (e.g., English: Tanner et al., 2019). Some of the variable results within a language may be due to dialect differences, though variation in study design also contributes to differences in results.
The results are separated by the consonant’s position in the word and the manner of articulation of the consonant. They are pooled across other variables (e.g., place of articulation, vowel quality), since very few studies separate the data along these dimensions. Tables 2, 3, 4, 5, 6, 7, 8 pool results for phonologically long and short vowels and for geminates and singletons. Tables 9, 10, 11, 12 separate results based on phonological length contrasts for studies which provide data separated along these dimensions.
The data tables which will be presented here do not always use the full set of items reported in each study. Data was restricted to exact minimal pairs when it was possible to calculate mean duration just using exactly paired items. When the results were partially grouped, e.g., by place of articulation, the table uses the maximal set of values for which those parameters match. The data excludes consonants which are contrastively aspirated, breathy, ejective, or implosive.
If a study did not provide numeric summaries of duration in each voicing environment, the durations were estimated based on visualizations illustrating the duration of vowels in each environment. The data table in the Supplementary Materials indicates each study which required estimates based on a figure. If a study did not provide enough information to estimate vowel durations, it was not included (e.g., Mitleb, 1984).
Table 2 presents vowel duration as conditioned by voicing of word-final stops. For some languages, multiple studies were available. Studies on the same language are pooled into a single summary measurement; each study was given equal weight for the by-language average. The columns indicate the language, number of studies which data in the row comes from, mean vowel duration before voiced consonants, mean vowel duration before voiceless consonants, mean difference between vowel duration before voiceless consonants versus voiced consonants, and mean ratio of vowel duration before voiceless consonants versus voiced consonants. For languages represented by multiple studies, the average for each measurement is followed by the range across studies.
Table 2: Vowel duration in milliseconds by language and voicing of the following word-final stop, for languages with contrastive word-final voicing.
| Language | # stu-dies | before voiced mean (range) | before vcls mean (range) | V duration diff mean (range) | V duration ratio mean (range) |
| Arabic | 6 | 138 (101–193) | 133 (96–188) | –5 (–14 to 0.5) | 0.96 (0.89–1.0) |
| Armenian | 2 | 127 (123–131) | 116 (114–117) | –11 (–14 to –9) | 0.91 (0.89–0.93) |
| Assamese | 2 | 107 (63–151) | 89 (53–124) | –18 (–27 to –10) | 0.83 (0.82–0.85) |
| Bengali | 2 | 119 (118–121) | 106 (101–111) | –13 (–17 to –10) | 0.89 (0.86–0.92) |
| English | 14 | 222 (98–335) | 149 (69–242) | –73 (–131 to –17) | 0.67 (0.54–0.89) |
| Farsi | 3 | 186 (146–225) | 160 (108–207) | –26 (–37 to –19) | 0.86 (0.74–0.92) |
| French | 8 | 159 (74–415) | 137 (69–396) | –22 (–49 to –4) | 0.86 (0.73–0.96) |
| Gujarati | 1 | 223 | 217 | –6 | 0.97 |
| Hindi | 3 | 176 (158–185) | 153 (145–160) | –23 (–30 to –13) | 0.87 (0.84–0.92) |
| Hungarian | 4 | 131 (99–148) | 116 (94–130) | –15 (–28 to –5) | 0.89 (0.81–0.95) |
| Javanese | 1 | 132 | 117 | –15 | 0.89 |
| Kabyle | 1 | 147 | 94 | –53 | 0.64 |
| Kanashi | 1 | 159 | 127 | –32 | 0.8 |
| Kapampangan | 1 | 129 | 115 | –14 | 0.89 |
| Lezgian | 1 | 212 | 197 | –15 | 0.93 |
| Maithili | 1 | 185 | 159 | –26 | 0.86 |
| Marathi | 2 | 150 (120–179) | 130 (105–154) | –20 (–25 to –15) | 0.87 (0.86–0.88) |
| Nepali | 1 | 127 | 109 | –18 | 0.86 |
| Ozolotepec Zapotec | 1 | 161 | 97 | –64 | 0.6 |
| Pashto | 1 | 71 | 62 | –9 | 0.87 |
| Q’eqchi’ | 1 | 140 | 103 | –37 | 0.74 |
| Serbian | 3 | 168 (123–224) | 146 (102–198) | –22 (–26 to –20) | 0.87 (0.83–0.88) |
| Swedish | 2 | 167 (140–194) | 147 (126–168) | –20 (–26 to –14) | 0.88 (0.86–0.9) |
| Tagalog | 1 | 148 | 114 | –34 | 0.77 |
| Tashlhiyt Berber | 1 | 111 | 87 | –24 | 0.78 |
| Toda | 1 | 135 | 112 | –23 | 0.83 |
| Welsh | 1 | 151 | 111 | –40 | 0.74 |
| Yalálag Zapotec | 1 | 154 | 125 | –29 | 0.81 |
Previous work varies in whether voicing-conditioned vowel duration is measured as a difference (e.g., Coretta, 2019; Maddieson, 1977; Mitleb, 1984) or a ratio (e.g., Chen, 1970; Mack, 1982; Peterson & Lehiste, 1960). It is not clear whether one measure is better motivated than another. The two approaches differ in how they reflect interactions between voicing-conditioned vowel duration and other variation in vowel duration. In some languages, the vowel duration ratio between voiced and voiceless environments is relatively consistent when combined with other factors that influence vowel duration, while duration differences are variable (e.g., English: de Jong, 2004; House, 1961; Peterson & Lehiste, 1960; Solé, 2007; German: Braunschweiler, 1997). In other languages, the vowel duration difference is more consistent (e.g., Arabic: de Jong & Zawaydeh, 2002; Catalan: Solé, 2007). For these reasons, both measurements of voicing-conditioned vowel duration are given here: the difference between vowel duration before voiceless consonants versus voiced consonants, and the ratio of vowel duration before voiceless consonants versus voiced consonants. A duration difference of 0 indicates that vowels have the same duration in both voicing environments. Difference values more distant from 0 indicate a larger effect of the environment. A duration ratio of 1.0 indicates that vowels have the same duration in both voicing environments. A smaller vowel duration ratio indicates a larger effect of the environment.
The focus here is on evaluating patterns in the size of the effect rather than imposing a binary divide, so the tables do not include the number of studies finding a significant effect. The threshold for significance varies widely across studies based on differences in sample size. In addition, some studies do not test significance at all.
Table 3 presents vowel duration as conditioned by voicing of word-final fricatives. Note that the different tables do not always include the same languages; many studies examine only one position and one manner of articulation.
Table 3: Vowel duration in milliseconds by language and voicing of the following word-final fricative, for languages with contrastive word-final voicing.
| Language | # stu-dies | before voiced mean (range) | before vcls mean (range) | V duration diff mean (range) | V duration ratio mean (range) |
| Arabic | 2 | 181 (122–240) | 160 (104–216) | –21 (–24 to –19) | 0.88 (0.85–0.9) |
| English | 7 | 282 (169–387) | 203 (128–313) | –79 (–125 to –38) | 0.72 (0.65–0.81) |
| Farsi | 1 | 245 | 192 | –53 | 0.78 |
| French | 7 | 211 (113–458) | 150 (90–357) | –61 (–117 to –16) | 0.71 (0.52–0.86) |
| Hebrew | 1 | 131 | 82 | –49 | 0.63 |
| Hungarian | 4 | 152 (114–226) | 129 (100–206) | –23 (–44 to –14) | 0.85 (0.7–0.91) |
| Kabyle | 1 | 182 | 148 | –34 | 0.81 |
| Ozolotepec Zapotec | 1 | 177 | 120 | –57 | 0.68 |
| Serbian | 1 | 152 | 125 | –27 | 0.82 |
| Tashlhiyt Berber | 1 | 103 | 79 | –24 | 0.77 |
| Turkish | 1 | 121 | 93 | –28 | 0.77 |
Table 4 presents vowel duration as conditioned by voicing of word-final affricates. Very few studies report measurements for voicing-conditioned vowel duration with affricates; some studies pool stops and affricates (e.g., Mikuteit & Reetz, 2007 for Bengali).
Table 4: Vowel duration in milliseconds by language and voicing of the following word-final affricate, for languages with contrastive word-final voicing.
| Language | # stu-dies | before voiced mean | before voiceless mean | V duration diff mean | V duration ratio mean |
| English | 1 | 246 | 172 | –74 | 0.7 |
| Maithili | 1 | 170 | 128 | –42 | 0.75 |
| Ozolotepec Zapotec | 1 | 172 | 110 | –62 | 0.64 |
Table 5 presents the effect of underlying voicing word-finally in languages with word-final devoicing. In these languages, any difference in voicing-conditioned vowel duration indicates incomplete neutralization of the voicing contrast. Many languages with contrastive word-final voicing do nonetheless exhibit some degree of final devoicing; the languages here are conventionally described as (largely) neutralizing the voicing contrast word-finally.
Table 5: Vowel duration in milliseconds by language and phonological voicing of the following word-final consonant, for languages with word-final devoicing.
| Language | # stu-dies | manner | before voiced mean (range) | before vcls mean (range) | V duration diff mean (range) | V duration ratio mean (range) |
| Afrikaans | 2 | stops | 158 (150–166) | 147 (136–159) | –11 (–15 to –8) | 0.93 (0.9–0.95) |
| Catalan | 1 | stops | 75 | 74 | –1 | 0.99 |
| Dutch | 2 | stops | 150 (149–151) | 146 (144–148) | –4 (–5 to –4) | 0.97 (0.97–0.98) |
| Friulian | 1 | stops | 255 | 119 | –136 | 0.47 |
| German | 4 | stops | 178 (146–239) | 157 (119–213) | –21 (–50 to –3) | 0.88 (0.7–0.98) |
| German | 2 | fricatives | 185 (180–189) | 148 (110–186) | –37 (–70 to –3) | 0.8 (0.61–0.98) |
| Kazakh | 1 | stops | 115 | 108 | –7 | 0.94 |
| Kazakh | 1 | fricatives | 141 | 117 | –24 | 0.83 |
| Lithuanian | 1 | stops | 166 | 145 | –21 | 0.87 |
| Polish | 4 | stops | 114 (102–127) | 107 (94–117) | –7 (–10 to –3) | 0.94 (0.92–0.97) |
| Polish | 2 | fricatives | 125 (114–136) | 115 (113–116) | –10 (–20 to –1) | 0.92 (0.85–0.99) |
| Polish | 2 | affricates | 120 (103–137) | 111 (93–128) | –9 (–10 to –9) | 0.93 (0.9–0.93) |
| Russian | 5 | stops | 143 (115–172) | 135 (109–169) | –8 (–22 to –3) | 0.94 (0.87–0.98) |
| Russian | 2 | fricatives | 145 (123–167) | 141 (120–162) | –4 (–5 to –3) | 0.97 (0.97–0.98) |
| Slovak | 2 | stops | 72 (67–77) | 67 (60–74) | –5 (–8 to –3) | 0.93 (0.89–0.96) |
| Slovak | 1 | fricatives | 102 | 94 | –8 | 0.92 |
| Slovenian | 1 | stops | 89 | 78 | –11 | 0.88 |
| Sorani Kurdish | 1 | stops | 112 | 91 | –21 | 0.81 |
| Turkish | 1 | stops | 93 | 90 | –3 | 0.97 |
Table 6 presents vowel duration as conditioned by voicing of word-medial stops. Table 7 presents vowel duration as conditioned by voicing of word-medial fricatives. Table 8 presents vowel duration as conditioned by voicing of word-medial affricates. In most of the studies with word-medial consonants, the consonants were intervocalic. However, a few studies looked at medial consonants followed by other consonants: Wissing (1992) on Afrikaans, Podlipský and Chládková (2007) on Czech, Campos-Astorkiza (2012) on Lithuanian, and Campos-Astorkiza (2019) on Spanish. In other studies, the following segment was included as a variable (e.g., Pajak, 2013 for Arabic), or was a consonant only in some items (e.g., Braunschweiler, 1997 for German). Sometimes the consonantal environment is crucial for avoiding manner differences which arise intervocalically due to lenition of voiced obstruents, e.g., in Spanish (see Delattre, 1962, for a discussion). Even in languages without categorical lenition processes, phonetic lenition may be present (Katz, 2021).
Table 6: Vowel duration in milliseconds by language and voicing of the following word-medial stop
| Language | # stu-dies | before voiced mean (range) | before vcls mean (range) | V duration diff mean (range) | V duration ratio mean (range) |
| Amharic | 1 | 98 | 81 | –17 | 0.83 |
| Arabic | 7 | 94 (61–137) | 83 (48–115) | –11 (–24 to 1) | 0.88 (0.78–1.01) |
| Argobba | 1 | 100 | 95 | –5 | 0.95 |
| Assamese | 1 | 86 | 63 | –23 | 0.73 |
| Balami | 1 | 123 | 100 | –23 | 0.81 |
| Bengali | 2 | 103 (84–121) | 90 (69–111) | –13 (–16 to –10) | 0.87 (0.81–0.92) |
| Breton | 1 | 108 | 58 | –50 | 0.54 |
| Bulgarian | 1 | 80 | 64 | –16 | 0.8 |
| Chamoru | 1 | 104 | 98 | –6 | 0.94 |
| Czech | 3 | 161 (80–204) | 149 (57–197) | –12 (–23 to –2) | 0.93 (0.71–0.99) |
| Degema | 1 | 142 | 130 | –12 | 0.92 |
| Dogri | 1 | 71 | 54 | –17 | 0.76 |
| Dutch | 5 | 156 (108–189) | 133 (87–160) | –23 (–29 to –11) | 0.85 (0.81–0.93) |
| English | 11 | 108 (62–152) | 96 (52–148) | –12 (–25 to 1) | 0.89 (0.81–1.01) |
| Farsi | 2 | 134 (120–149) | 117 (111–122) | –17 (–27 to –9) | 0.87 (0.82–0.93) |
| French | 3 | 125 (98–140) | 102 (88–117) | –23 (–36 to –10) | 0.82 (0.74–0.9) |
| Friulian | 3 | 199 (161–233) | 156 (128–176) | –43 (–68 to –26) | 0.78 (0.71–0.87) |
| Georgian | 1 | 96 | 82 | –14 | 0.85 |
| German | 8 | 115 (70–162) | 88 (56–128) | –27 (–64 to –12) | 0.77 (0.6–0.9) |
| Greek | 1 | 79 | 70 | –9 | 0.89 |
| Gujarati | 1 | 151 | 125 | –26 | 0.83 |
| Hausa | 1 | 83 | 83 | 0 | 1.0 |
| Hindi | 2 | 90 (80–100) | 80 (70–89) | –10 (–11 to –10) | 0.89 (0.88–0.89) |
| Hungarian | 2 | 84 (82–87) | 83 (80–85) | –1 (–2 to –1) | 0.99 (0.97–0.99) |
| Ibibio | 1 | 106 | 96 | –10 | 0.91 |
| Igbo | 1 | 143 | 137 | –6 | 0.96 |
| Italian | 6 | 133 (91–192) | 117 (81–155) | –16 (–36 to –3) | 0.88 (0.81–0.97) |
| Japanese | 7 | 97 (66–139) | 82 (47–116) | –15 (–23 to –6) | 0.85 (0.71–0.93) |
| Javanese | 1 | 320 | 256 | –64 | 0.8 |
| Kabyle | 2 | 76 (53–98) | 70 (53–87) | –6 (–12 to 0) | 0.92 (0.88–1.0) |
| Kanashi | 1 | 71 | 65 | –6 | 0.92 |
| Kannada | 1 | 88 | 73 | –15 | 0.83 |
| Kapampangan | 1 | 70 | 57 | –13 | 0.81 |
| Kazakh | 1 | 81 | 75 | –6 | 0.93 |
| Korean | 3 | 121 (87–156) | 85 (68–94) | –36 (–65 to –19) | 0.7 (0.59–0.79) |
| Lezgian | 1 | 79 | 72 | –7 | 0.91 |
| Lio | 1 | 63 | 54 | –9 | 0.86 |
| Lithuanian | 2 | 152 (136–167) | 130 (116–144) | –22 (–23 to –21) | 0.86 (0.85–0.86) |
| Marathi | 1 | 74 | 61 | –13 | 0.82 |
| Nepali | 1 | 166 | 122 | –44 | 0.73 |
| Oriya | 1 | 104 | 72 | –32 | 0.69 |
| Oromo | 1 | 76 | 68 | –8 | 0.89 |
| Polish | 6 | 108 (83–170) | 100 (76–167) | –8 (–19 to –2) | 0.93 (0.86–0.99) |
| Portuguese | 5 | 143 (123–167) | 115 (98–146) | –28 (–50 to –16) | 0.8 (0.66–0.87) |
| Q’eqchi’ | 1 | 106 | 85 | –21 | 0.8 |
| Romanian | 1 | 148 | 133 | –15 | 0.9 |
| Russian | 3 | 111 (100–127) | 96 (74–115) | –15 (–26 to –8) | 0.86 (0.74–0.92) |
| Serbian | 3 | 161 (145–169) | 136 (115–147) | –25 (–30 to –21) | 0.84 (0.79–0.87) |
| Shanghai Wu | 1 | 134 | 113 | –21 | 0.84 |
| Swedish | 2 | 142 (131–152) | 124 (117–130) | –18 (–23 to –13) | 0.87 (0.85–0.9) |
| Tagalog | 1 | 115 | 113 | –2 | 0.98 |
| Tarifit Berber | 1 | 76 | 72 | –4 | 0.95 |
| Tashlhiyt Berber | 1 | 81 | 79 | –2 | 0.98 |
| Telugu | 1 | 203 | 190 | –13 | 0.94 |
| Turkish | 3 | 86 (71–115) | 68 (54–97) | –18 (–18 to –17) | 0.79 (0.76–0.84) |
| Washo | 1 | 184 | 151 | –33 | 0.82 |
| Wolaytta Doonaa | 1 | 115 | 118 | 3 | 1.03 |
| Yalálag Zapotec | 1 | 170 | 128 | –42 | 0.75 |
| Yoruba | 1 | 142 | 119 | –23 | 0.84 |
Table 7: Vowel duration in milliseconds by language and voicing of the following word-medial fricative.
| Language | # stu-dies | before voiced mean (range) | before vcls mean (range) | V duration diff mean (range) | V duration ratio mean (range) |
| Afrikaans | 1 | 179 | 143 | –36 | 0.8 |
| Arabic | 2 | 102 (97–107) | 85 (78–92) | –17 (–29 to –5) | 0.83 (0.73–0.95) |
| Breton | 1 | 142 | 82 | –60 | 0.58 |
| Bulgarian | 1 | 94 | 84 | –10 | 0.89 |
| Catalan | 1 | 104 | 93 | –11 | 0.89 |
| Dawoodi | 1 | 134 | 109 | –25 | 0.81 |
| Dutch | 2 | 216 (207–225) | 172 (163–181) | –44 | 0.8 (0.79–0.8) |
| English | 3 | 102 (74–124) | 78 (45–103) | –24 (–38 to –6) | 0.76 (0.61–0.94) |
| Farsi | 1 | 117 | 112 | –5 | 0.96 |
| French | 4 | 161 (106–310) | 129 (82–246) | –32 (–64 to –17) | 0.8 (0.77–0.85) |
| Friulian | 2 | 251 (233–269) | 201 (186–217) | –50 (–83 to –16) | 0.8 (0.69–0.93) |
| German | 3 | 121 (99–163) | 102 (66–160) | –19 (–34 to –4) | 0.84 (0.66–0.98) |
| Greek | 1 | 127 | 110 | –17 | 0.87 |
| Hindko | 1 | 204 | 183 | –21 | 0.9 |
| Hungarian | 1 | 91 | 83 | –8 | 0.91 |
| Italian | 3 | 160 (115–203) | 125 (85–148) | –35 (–55 to –19) | 0.78 (0.73–0.88) |
| Japanese | 2 | 112 (98–125) | 109 (91–127) | –3 (–7 to 2) | 0.97 (0.93–1.02) |
| Kabyle | 1 | 125 | 94 | –31 | 0.75 |
| Kazakh | 1 | 93 | 79 | –14 | 0.85 |
| Polish | 3 | 123 (98–164) | 105 (90–131) | –18 (–33 to –8) | 0.85 (0.8–0.91) |
| Portuguese | 2 | 154 (145–162) | 116 (100–133) | –38 (–46 to –29) | 0.75 (0.69–0.82) |
| Romanian | 1 | 152 | 139 | –13 | 0.91 |
| Russian | 2 | 179 (155–203) | 144 (136–152) | –35 (–51 to –19) | 0.8 (0.75–0.88) |
| Serbian | 1 | 157 | 137 | –20 | 0.87 |
| Spanish | 1 | 68 | 57 | –11 | 0.84 |
| Tashlhiyt Berber | 1 | 84 | 65 | –19 | 0.77 |
| Turkish | 1 | 136 | 101 | –35 | 0.74 |
Table 8: Vowel duration in milliseconds by language and voicing of the following word-medial affricate.
| Language | # stu-dies | before voiced mean (range) | before vcls mean (range) | V duration diff mean (range) | V duration ratio mean (range) |
| Catalan | 1 | 97 | 90 | –7 | 0.93 |
| Hungarian | 1 | 88 | 76 | –12 | 0.86 |
| Italian | 2 | 125 (109–140) | 108 (96–121) | –17 (–19 to –13) | 0.86 (0.86–0.88) |
| Japanese | 1 | 127 | 113 | –14 | 0.89 |
| Polish | 1 | 139 | 118 | –21 | 0.85 |
| Serbian | 1 | 135 | 110 | –25 | 0.81 |
Korean differs from the other languages because the lenis stops are only realized as voiced when between voiced sounds (e.g., Han, 2000), whereas the lenis series in the other languages is generally treated as being voiced in the underlying form. However, it has been included here both because voicing is a robust correlate of the fortis/lenis contrast in the environment investigated in these studies and also because it is one of the languages presented in Chen’s (1970) seminal work.
Two languages could not be included due to studies not separating measurements by both position in the word and manner of articulation. Aasmäe et al. (2016) pool results for word-medial Moksha stops and fricatives: Vowels are longer before voiced consonants than before voiceless consonants (134 ms vs 114 ms). Fintoft (1961) provides results for vowel duration preceding labiodental fricatives in Norwegian (175 ms before voiced, 148 ms before voiceless), but does not separate the results based on whether the consonant was word-medial or final. In addition, it is unclear whether the Norwegian voiced labiodental is best analyzed as a fricative or as an approximant.
In some languages, voiceless stops are preaspirated. It is debated whether preaspiration should be included as part of the duration of the vowel. Decisions about how to segment preaspiration have a large effect on vowel duration measurements. The treatment of preaspiration makes the difference between whether vowels are longer before voiced stops or before voiceless stops in Scottish Gaelic (Ní Chasaide, 1985, p. 188) and Norwegian (Ringen & Van Dommelen, 2013). Because of the uncertainty about how vowel duration should be measured before preaspirated consonants and how effects of preaspiration on preceding vowel duration might differ from effects of voicelessness, studies on preaspiration have been excluded here. However, it is important to keep in mind that some languages may have preaspiration which has not been described. Many languages have some form of preaspiration (Craioveanu, 2023).
One salient omission from this list of languages is worth mentioning. Previous work often cites Danish (Fischer-Jørgensen, 1964) among languages exhibiting voicing-conditioned vowel duration. However, the only voiceless consonants used in the study are in word-initial position, and the effect reported is based on preceding consonant voicing. This confusion is likely the result of a labelling error: The legend in Figure 5, which appears in Fischer-Jørgensen’s work, says “Vowels before bdgv, ptkfh,” though the comparison is correctly labelled with “after” in the caption and the following table.
Table 9 looks at vowel duration as conditioned by voicing of word-final stops separated by phonological vowel length in studies that tested both long and short vowels. Languages with word-final devoicing are excluded. Only a few studies provide measurements split by phonological vowel length. Because most of the data here comes from an individual study for each language, the results are separated by study, with a column to indicate the source.
Table 9: Vowel duration in milliseconds by language, voicing of the following word-final stop, and phonological vowel length.
| Language | reference | V length category | before voiced | before vcls | V dur diff | V dur ratio |
| Hungarian | Magdics, 1969 | long | 189 | 163 | -26 | 0.86 |
| short | 95 | 97 | 2 | 1.02 | ||
| Serbian | Sokolović-Perović, 2012 | long | 186 | 172 | -14 | 0.92 |
| short | 129 | 104 | -25 | 0.81 | ||
| Swedish | Elert, 1964 | long | 162 | 153 | -9 | 0.94 |
| short | 118 | 100 | -18 | 0.85 | ||
| Swedish | Helgason et al., 2013 | long | 246 | 228 | -18 | 0.93 |
| short | 143 | 108 | -35 | 0.76 | ||
| Toda | new from UCLA Phonetics Archive | long | 183 | 157 | -26 | 0.86 |
| short | 114 | 92 | -22 | 0.81 |
Table 10 looks at vowel duration as conditioned by voicing of word-medial stops separated by phonological vowel length in studies that tested both long and short vowels. For most of these studies, the consonants of interest were intervocalic. However, the Lithuanian data (Campos-Astorkiza, 2012) comes from vowels before stop+fricative sequences, and the German data (Braunschweiler, 1997) includes one short vowel token before /tr/.
Table 10: Vowel duration in milliseconds by language, voicing of the following word-medial stop, and phonological vowel length.
| Language | reference | V length category | before voiced | before vcls | V dur diff | V dur ratio |
| Arabic | Port et al., 1980 | long | 191 | 168 | –23 | 0.88 |
| short | 82 | 62 | –20 | 0.76 | ||
| Czech | Keating, 1985 | long | 267 | 258 | –9 | 0.97 |
| short | 130 | 136 | 6 | 1.05 | ||
| Dutch | Warner et al., 2004 | long | 184 | 162 | –22 | 0.88 |
| short | 103 | 83 | –20 | 0.81 | ||
| German | Braunschweiler, 1997 | long | 194 | 161 | –33 | 0.83 |
| short | 97 | 80 | –17 | 0.82 | ||
| Hausa | Lindau-Webb, 1985 | long | 127 | 118 | –9 | 0.93 |
| short | 70 | 71 | 1 | 1.01 | ||
| Hindi | Maddieson, 1977 | long | 155 | 148 | –7 | 0.95 |
| short | 82 | 70 | –12 | 0.85 | ||
| Japanese | Ishibashi & Koya, 2023 | long | 202 | 174 | –28 | 0.86 |
| short | 76 | 58 | –18 | 0.76 | ||
| Lithuanian | Campos-Astorkiza, 2012 | long | 166 | 143 | –23 | 0.86 |
| short | 101 | 84 | –17 | 0.83 | ||
| Oromo | Feda Negesse & Tujube Amansa, 2021 | long | 160 | 160 | 0 | 1.0 |
| short | 77 | 68 | –9 | 0.88 | ||
| Serbian | Sokolović-Perović, 2012 | long | 204 | 179 | –25 | 0.88 |
| short | 134 | 111 | –23 | 0.83 | ||
| Swedish | Elert, 1964 | long | 152 | 139 | –13 | 0.91 |
| short | 108 | 94 | –14 | 0.87 | ||
| Swedish | Helgason et al., 2013 | long | 232 | 200 | –32 | 0.86 |
| short | 113 | 95 | –18 | 0.84 | ||
| Telugu | Reddy, 1988 | long | 320 | 300 | –20 | 0.94 |
| short | 85 | 80 | –5 | 0.94 | ||
| Wolaytta Doonaa | F. Elias, 2017 | long | 156 | 162 | 6 | 1.04 |
| short | 74 | 74 | 0 | 1.0 |
Another potentially informative way to separate the data is based on whether the consonants were geminates or singletons. Table 11 separates vowel duration measurements by consonant length for word-medial consonants in studies which divide the results by the phonological length of the consonant. In all cases, these consonants were intervocalic. In some languages, such as Swedish, gemination and preceding vowel length are categorically inversely correlated, i.e., long vowels occur only before singletons and short vowels occur only before geminates; Table 10 gives the results for these length divides.
Table 11: Vowel duration in milliseconds by language and manner, length, and voicing of the following word-medial consonant.
| Language | reference | manner | C length | before voiced | before vcls | V dur diff | V dur ratio |
| Arabic | Alamri, 2022 | stops | gem | 79 | 68 | –11 | 0.86 |
| sing | 94 | 79 | –15 | 0.84 | |||
| Arabic | Al-Tamimi & Khattab, 2018 | stops | gem | 90 | 67 | –23 | 0.74 |
| sing | 123 | 98 | –25 | 0.8 | |||
| Arabic | Ferrat & Guerti, 2015 | stops | gem | 34 | 39 | 5 | 1.15 |
| sing | 88 | 56 | –32 | 0.64 | |||
| Arabic | Alamri, 2022 | fricatives | gem | 103 | 76 | –27 | 0.74 |
| sing | 110 | 79 | –31 | 0.72 | |||
| Arabic | Pajak, 2013 | fricatives | gem | 83 | 79 | –4 | 0.95 |
| sing | 112 | 106 | –6 | 0.95 | |||
| Bengali | Ghosh, 2015 | stops | gem | 81 | 65 | –16 | 0.8 |
| sing | 88 | 72 | –16 | 0.82 | |||
| Bulgarian | Maneva, 1997 | stops | gem | 74 | 61 | –13 | 0.82 |
| sing | 85 | 67 | –18 | 0.79 | |||
| Bulgarian | Maneva, 1997 | fricatives | gem | 90 | 75 | –15 | 0.83 |
| sing | 97 | 92 | –5 | 0.95 | |||
| Chamoru | Santos, 2021 | stops | gem | 89 | 90 | 1 | 1.01 |
| sing | 119 | 105 | –14 | 0.88 | |||
| Dogri | Badyal, 2026 | stops | gem | 64 | 47 | –17 | 0.73 |
| sing | 79 | 62 | –17 | 0.78 | |||
| Farsi | Zirak & Skaer, 2013 | stops | gem | 103 | 106 | 3 | 1.03 |
| sing | 136 | 116 | –20 | 0.85 | |||
| Farsi | Zirak & Skaer, 2013 | fricatives | gem | 109 | 106 | –3 | 0.97 |
| sing | 125 | 118 | –7 | 0.94 | |||
| Italian | Dian et al., 2024 | stops | gem | 81 | 73 | –8 | 0.9 |
| sing | 102 | 90 | –12 | 0.88 | |||
| Italian | Dipino & Celata, 2018 | stops | gem | 151 | 133 | –18 | 0.88 |
| sing | 186 | 166 | –20 | 0.89 | |||
| Italian | Shigemori & Vietti, 2019 | stops | gem | 130 | 107 | –23 | 0.82 |
| sing | 130 | 123 | –7 | 0.95 | |||
| Italian | Di Benedetto & De Nardis, 2021 | fricatives | gem | 135 | 120 | –15 | 0.89 |
| sing | 187 | 165 | –22 | 0.88 | |||
| Italian | Di Benedetto & De Nardis, 2021 | affricates | gem | 118 | 105 | –13 | 0.89 |
| sing | 162 | 137 | –25 | 0.85 | |||
| Japanese | Kawahara, 2006 | stops | gem | 71 | 53 | –18 | 0.75 |
| sing | 46 | 37 | –9 | 0.8 | |||
| Kabyle | A. Elias, 2020 | fricatives | gem | 123 | 100 | –23 | 0.81 |
| sing | 128 | 89 | –39 | 0.7 | |||
| Polish | Malisz, 2013 | stops | gem | 94 | 87 | –7 | 0.93 |
| sing | 82 | 75 | –7 | 0.91 | |||
| Polish | Rojczyk & Porzuczek, 2019 | stops | gem | 85 | 77 | –8 | 0.91 |
| sing | 91 | 77 | –14 | 0.85 | |||
| Polish | Malisz, 2013 | fricatives | gem | 114 | 98 | –16 | 0.86 |
| sing | 83 | 82 | –1 | 0.99 | |||
| Tashlhiyt Berber | Ridouane, 2007 | stops | gem | 67 | 71 | 4 | 1.06 |
| sing | 94 | 86 | –8 | 0.91 | |||
| Tashlhiyt Berber | Ridouane, 2007 | fricatives | gem | 77 | 57 | –20 | 0.74 |
| sing | 91 | 72 | –19 | 0.79 |
Table 12 looks at vowel duration as conditioned by following obstruent voicing separated by consonant length for word-final consonants. Data was only available from two studies.
Table 12: Vowel duration in milliseconds by language and manner, length, and voicing of the following word-final consonant.
| Language | reference | manner | C length | before voiced | before vcls | V dur diff | V dur ratio |
| Kabyle | A. Elias, 2020 | fricatives | gem | 149 | 114 | –35 | 0.77 |
| sing | 215 | 181 | –34 | 0.84 | |||
| Tashlhiyt Berber | Ridouane, 2007 | stops | gem | 99 | 75 | –24 | 0.76 |
| sing | 123 | 98 | –25 | 0.8 | |||
| Tashlhiyt Berber | Ridouane, 2007 | fricatives | gem | 94 | 73 | –21 | 0.78 |
| sing | 112 | 85 | –27 | 0.76 |
While the focus here is on effects of voicing, some of the studies also provide data for effects of aspiration and ejectives on preceding vowels. Vowels are slightly longer before breathy voiced (voiced aspirated) stops than before voiced unaspirated stops, though interpretation of this pattern is limited by the fact that all of the evidence comes from Indic languages (Assamese: Dutta & Kenstowicz, 2018; Maddieson, 1977; Bengali: Maddieson, 1977; Hindi: Durvasula & Luo, 2014; Maddieson, 1977; M. Ohala & Ohala, 1992; Marathi: Maddieson, 1977; Maithili: Yadav, 1979; Nepali: Schwarz, 2024). The effect of aspiration among voiceless consonants is less consistent. The Indic languages generally have slightly longer vowels before aspirated voiceless stops than before unaspirated voiceless stops, paralleling the voiced stops. However, vowels have the same duration or shorter duration before aspirated stops versus before voiceless stops in Armenian (Maddieson, 1977) and Balami (Gautam, 2012). Vowels in Korean are longer before aspirated stops than before fortis (tense) stops (Kang, 1999). Vowels are slightly longer before ejectives than before plain voiceless stops (Georgian: Beguš, 2017; Washo: Yu, 2008). Beguš (2017) discusses how to interpret the effect of ejectives on preceding vowel duration.
2.1. Typological patterns
There are some observable patterns within the data, which provide information for the typical size of voicing-conditioned vowel duration and how it interacts with other characteristics.
Figure 1 shows the distribution of how many languages fall into each range of vowel duration ratios by position and by consonant manner. Affricates are excluded, given the very small number of studies which report them. Measurements from word-final position in languages with final devoicing are excluded. Each datapoint is an average for the consonant manner and position in each language.
Figure 1: Histograms and overlaid density curves for number of languages with each vowel duration ratio (before voiceless versus before voiced), separated by position and by consonant manner. These are the by-language vowel duration ratios reported in Table 2 through Table 7. A smaller ratio reflects a larger difference in vowel duration based on the following consonant voicing.
In medial position, Breton is clearly separate from the other languages in having a particularly large effect when measured as a ratio, i.e., a small ratio of durations; this is the bin which is an outlier on the left side of each distribution in Figure 1. It has a ratio of 0.54 with stops and 0.57 with fricatives, which is much smaller than the next smallest ratio with stops (Oriya, at 0.69) or fricatives (Turkish, at 0.74).
Figure 2 also shows the distribution of languages, but presents the voicing effect measured as a difference. When voicing-conditioned vowel duration is measured as a difference, the overall distribution is largely similar to the distribution of the effect when measured as a ratio, but the position of individual languages in the distribution is substantially different. English has the largest difference (the most extreme negative value) for word-final stops (–73) and fricatives (–79). The next largest difference with stops is in Ozolotepec Zapotec (–64) and for fricatives, it is in French (–61). Javanese has the largest difference for word-medial stops (–64), followed by Breton (–50). Breton has the largest difference for word-medial fricatives (–60), followed by Friulian (–50).
Figure 3 shows the relationship between voicing-conditioned vowel duration in word-medial position and word-final position for the 23 languages with data from both positions and without word-final devoicing. Only manners of articulation observed in both positions for a language were included. In the by-language average, each manner of articulation was given equal weight to avoid potential effects based on how many studies examine each manner of articulation in each language: Using duration ratios, r(21) = 0.0925, p = 0.675; using duration differences, r(21) = –0.0956, p = 0.664.
Figure 4 shows the relationship between voicing-conditioned vowel duration with fricatives and stops from the 23 languages with data from both stops and fricatives in the same position. For languages with final devoicing, word-final position was excluded from the by-language average. Otherwise, each position attested with both manners of articulation was given equal weight. There is a strong positive correlation: Using duration ratios, r(21) = 0.761, p < 0.0001; using duration differences, r(21) = 0.808, p < 0.0001.
Very few languages are consistently found to have a large effect of voicing on preceding vowel duration when measured as a ratio (ratios below 0.7, i.e., vowels before voiceless obstruents are less than 70% the duration of vowels before voiced obstruents). English is the most consistent in having a large effect word-finally (average of 0.67 with stops, 0.7 with fricatives), across a large number of studies, though the effect is weaker word-medially (average of 0.89 with stops, 0.76 with fricatives). Kabyle (Taqbaylit Berber), Korean, and Portuguese have a relatively large effect in some studies but a much weaker effect in other studies, sometimes even for the same manner of articulation within the same position in the word; see the range of vowel duration ratios in Table 2 through Table 7. Large effects of voicing-conditioned vowel duration have also been found in Breton, Hebrew, Oriya, and Ozolotepec Zapotec, but each language is only represented by a single study. When measured as a difference, the patterns are similar; the largest differences are not consistent across studies within a language or evidence for the large difference within a language only comes from a single study. Many of the particular languages exhibiting a large duration difference are also languages with the largest average absolute vowel durations: English (word-finally), Farsi, French (word-finally), Friulian, and Javanese (word-medially).
There are also relatively few languages which are consistently found to have very little effect of voicing on preceding vowel duration when measured as a ratio (ratios above 0.9). For several languages with small observed voicing effects, evidence comes from a single study each (see Table 1 and the Appendix): Argobba, Chamoru, Degema, Gujarati, Hausa, Hindko, Ibibio, Igbo, Kapampangan, Kanashi, Kazakh, Lezgian, Oromo, Romanian, Tagalog, Tarifit Berber, Telugu, and Wolaytta Doonaa. Given that results for a language can vary substantially across studies, it is difficult to interpret the results of an individual study. Some languages exhibit a small effect in certain positions or manners of articulation but larger effects in other conditions (e.g., Japanese, Hungarian, Polish, Tashlhiyt Berber) or have a small effect in certain studies but a larger effect in other studies with the same position and manner of articulation (e.g., Arabic, Czech, English, Farsi, French, German, Hungarian, Russian). Armenian is the language with the most consistent evidence for a small voicing effect: There are two studies, both on word-final stops, where each finds a small voicing effect (Maddieson, 1977; Seyfarth & Garellek, 2018). When measured as a difference, the patterns are similar.
Do any characteristics of a language predict the degree of voicing-conditioned vowel duration? Regression models were calculated using the lme4 package (Bates et al., 2015) in the R software program (R Core Team, 2021). The lmerTest package (Kuznetsova et al., 2015) was used to calculate p-values for the factors in these models. The models were built up from an intercept-only model. Adding factors was tested incrementally with model comparison; each factor was added if it significantly improved the model. Due to the relatively small amount of data and the large amount of variation within it, interactions were not considered.
Table 13 presents the summary of a mixed effects linear regression model for the vowel duration ratio (before voiceless versus before voiced).2 The fixed effects were position in the word (final, medial), manner of articulation (stops, affricates, fricatives), the status of gemination in the language (non-contrastive, contrastive), and the average duration of the vowels before voiceless consonants. This last factor was centered. There was a random intercept for language. The model included all of the data for studies which separated measurements by position and manner of articulation, excluding data for word-final consonants in languages with word-final devoicing: 275 datapoints from 71 languages.
Table 13: Model for vowel duration ratio (before voiceless versus before voiced). Recall that a smaller vowel duration ratio indicates more of a difference in vowel duration based on the voicing of the following consonant. Reference Levels: Position = Final, Manner = Stops, GeminationStatus = Non-contrastive.
| β | SE | t-value | p-value | |
| (Intercept) | 0.79 | 0.014 | 56.6 | < 0.0001 |
| Position Medial | 0.056 | 0.013 | 4.5 | < 0.0001 |
| Manner Affricates | –0.033 | 0.028 | –1.2 | 0.24 |
| Manner Fricatives | –0.048 | 0.012 | –3.9 | 0.00012 |
| GeminationStatus Contrastive | 0.048 | 0.015 | 3.2 | 0.0027 |
| V Dur Before Voiceless | 0.00045 | 0.00012 | 3.7 | 0.00027 |
Table 14: Model for vowel duration difference (before voiceless minus before voiced). A more extreme negative value indicates more of a difference in vowel duration based on the voicing of the following consonant. Reference Levels: Position = Final, Manner = Stops, GeminationStatus = Non-contrastive.
| β | SE | t-value | p-value | |
| (Intercept) | –33.6 | 2.8 | –11.9 | < 0.0001 |
| Position Medial | 11.8 | 2.6 | 4.5 | < 0.0001 |
| Manner Affricates | –7.4 | 6.0 | –1.2 | 0.21 |
| Manner Fricatives | –9.3 | 2.6 | –3.6 | 0.00041 |
| GeminationStatus Contrastive | 8.3 | 3.0 | 2.8 | 0.0068 |
| V Dur Before Voiceless | –0.13 | 0.026 | –5.0 | < 0.0001 |
Position in the word is a significant predictor of the extent of voicing-conditioned vowel duration. The overall effect of voicing on preceding vowel duration is smaller word-medially than word-finally, and is also observed individually in English, Farsi, Hungarian, Kabyle, Kanashi, Q’eqchi’, Tagalog, and Tashylhiyt Berber. However, not every language exhibits the same tendency. Among the languages with contrastive voicing both word-medially and word-finally for which there is data in both positions, some languages exhibit similar effects in both positions (Bengali, French, Hindi, Lezgian, Serbian, and Swedish). In Arabic, Assamese, Gujarati, Javanese, Kapampangan, Marathi, Nepali, and Yalálag Zapotec, voicing-conditioned vowel duration is a larger effect word-medially than word-finally. Among languages with word-final devoicing, the effect of voicing-conditioned vowel duration is usually smaller word-finally (based on underlying voicing) than word-medially, though not always, e.g., Friulian exhibits a larger effect word-finally than word-medially.
Manner of articulation is a significant predictor of the extent of voicing-conditioned vowel duration. Overall, the effect of voicing on preceding vowel duration is larger with fricatives than with stops. A larger effect of voicing-conditioned vowel duration with fricatives than with stops is also found in most individual languages in this study (Arabic, Dutch, English, French, Hungarian, Italian, Kazakh, Polish, Portuguese, Russian, Tashlhiyt Berber, Turkish), though sometimes the interaction is small or only in certain environments. On the other hand, Breton, Bulgarian, German, Japanese, and Ozolotepec Zapotec have a difference in the opposite direction; the effect of voicing-conditioned vowel duration is larger with stops than with fricatives. What is being evaluated here is how the effect of voicing on preceding vowel duration varies by manner—recall that the dependent variable is the duration ratio by voicing environment. One of the factors which might contribute to this difference in voicing-conditioned vowel duration is the absolute effect of manner on vowel duration: Vowels tend to be longer before fricatives than before stops, as has been established in previous work (Crystal & House, 1988; Laeufer, 1992; Peterson & Lehiste, 1960; Tanner et al., 2019).
The existence of a gemination contrast in a language is a significant predictor of the extent of voicing-conditioned vowel duration in that language. The effect of voicing on preceding vowel duration is smaller in languages which have gemination contrasts. One potential confound in experimental design might contribute to why the existence of a gemination contrast in a language would predict voicing-conditioned vowel duration. Several of the studies in languages with geminates were aimed at testing effects of gemination on preceding vowel duration and had minimal pairs for gemination, while the comparisons for voicing were less closely matched. The more limited control of additional influences on vowel duration may weaken the evidence for voicing-conditioned vowel duration.
Despite the language-level effect of gemination contrasts predicting weaker voicing-conditioned vowel duration, there was no strong pattern for gemination in a particular stimulus impacting voicing-conditioned vowel duration, i.e., no clear difference between the size of the voicing effect with singletons versus with geminates. As shown in Table 11, the interaction between gemination and the voicing effect varies substantially across studies, and the average size of the voicing effect is only slightly weaker with geminates than with singletons, both when measured as a ratio (0.87 vs 0.84) and when measured as a difference (14 ms vs 18 ms). This factor was not tested in the model because relatively few of the studies separated measurements based on gemination. While some studies find different results for geminates versus singletons, there is little evidence for consistency of that difference across manners of articulation (e.g., Tashlhyt Berber: Ridouane, 2007; Farsi: Zirak & Skaer, 2013; Polish: Malisz, 2013) or across different studies (e.g., Italian stops: Dian et al., 2024, Dipino & Celata, 2018, and Shigemori & Vietti, 2019; Arabic stops: Alamri, 2022, Al-Tamimi & Khattab, 2018, and Ferrat & Guerti, 2015). The one clear tendency with gemination is that vowels are generally longer before singletons than geminates, as has been noted previously (e.g., Al-Tamimi & Khattab, 2018; Di Benedetto & De Nardis, 2021; Ridouane, 2007).
The absolute duration of the vowel before voiceless consonants is a significant predictor of the extent of voicing-conditioned vowel duration. The effect of voicing on preceding vowel duration is smaller with longer vowels. One potential interpretation of this result is that at least some of the driving factors tend to produce consistent absolute differences rather than proportional differences. Within a language, phonological vowel length interacts with voicing-conditioned vowel duration (Tables 9 and 10), though how the interaction behaves depends on how the effect of voicing is measured. If the size of voicing-conditioned vowel duration is measured as a difference, the effect word-medially is larger with long vowels than with short vowels (17 ms vs. 12 ms) and word-finally is slightly smaller with long vowels than with short vowels (19 ms vs. 20 ms). If the size of the effect is measured as a ratio, the effect word-medially is slightly smaller with long vowels than with short vowels (0.91 vs. 0.88) and word-finally is also smaller with long vowels than with short vowels (0.9 vs. 0.85). Most studies did not use minimal quadruplets, so some differences in voicing-conditioned vowel duration for long vowels and short vowels might reflect confounding factors rather than differences inherent to the length contrast. The factor of phonological vowel length was not tested in the model because relatively few of the studies separated measurements based on phonological vowel length.
Some previous work has suggested that the existence of the phonological vowel length contrast in languages like Swedish can explain why the effect of voicing-conditioned vowel duration is weak (Buder & Stoel-Gammon, 2002). However, whether or not the language has contrastive vowel length did not improve the model (χ2 = 0.085, df = 1, p = 0.77). If there are constraints on using the same cue for multiple contrasts, it is not absolute. Vowel duration can be a cue to multiple contrasts within the same language, e.g., stress and vowel length in Arabic (de Jong & Zawaydeh, 2002), coda voicing and stress in English (de Jong, 2004).
One additional factor was tested but not included because it did not improve the model: Whether additional laryngeal contrasts such as aspiration exist in the language (χ2 = 0.056, df = 1, p = 0.81).
Table 14 presents the summary of a mixed effects linear regression model for the vowel duration difference (before voiceless minus before voiced) rather than vowel duration ratios.3 The fixed effects were position in the word (final, medial), manner of articulation (stops, affricates, fricatives), the status of gemination in the language (non-contrastive, contrastive), and the average duration of the vowels before voiceless consonants. This last factor was centered. There was a random intercept for language. The model included all of the data for studies which separated measurements by position and manner of articulation, excluding data for word-final consonants in languages with word-final devoicing: 275 datapoints from 71 languages.
There is one major difference between this model for duration difference and the preceding model for duration ratio. When measuring the voicing effect as a duration ratio, longer overall vowel durations predicted ratios closer to 1 (i.e., less of an effect). When measuring the voicing effect as a duration difference, longer overall vowel durations predicted more negative difference values (i.e., more of an effect).
The overall effect of position in the word is the same in this model as in the model using duration ratios as the dependent variable: The voicing effect is smaller word-medially than word-finally. However, some individual languages pattern differently depending on whether the effect is measured as a ratio or as a difference. When measured as a ratio, Arabic and Kapampangan exhibit a larger effect word-medially than word-finally, but when measured as a difference, they exhibit similar effect sizes in both environments. When measured as a ratio, French and Lezgian exhibit similar effect sizes in both environments, but when measured as a difference, they exhibit a larger effect in word-final position.
The effect of manner on the voicing effect is the same in this model as in the model using duration ratios as the dependent variable: The voicing effect is larger with fricatives than with stops. However, some individual languages pattern differently depending on whether the effect is measured as a ratio or as a difference. When measured as a difference, more languages exhibit a larger voicing effect with fricatives than with stops; some of these languages exhibited little variation by manner of articulation when measured as a ratio (Friulian, Greek), or exhibited a larger effect with stops than with fricatives when measured as a ratio (Breton).
2.2. Limitations for interpretation
A wide range of variables differ across studies, many of which can impact measurements of vowel duration. This means that no individual result can be interpreted as a definitive measurement of voicing-conditioned vowel duration within that language. Some studies also only have a small number of tokens or a small number of speakers (e.g., Chen, 1970; Maddieson, 1977). See the Supplementary Materials for the summary data from each study along with information about their methods.
Many characteristics which varied across these studies have been shown to interact with voicing-conditioned vowel duration. Some of these factors are: vowel height (Mack, 1982; Tanner et al., 2019), vowel tenseness (Luce & Charles-Luce, 1985; Peterson & Lehiste, 1960; Sharf, 1962), phonological vowel length (Elert, 1964; Maddieson, 1977; Sokolović-Perović, 2012), manner of articulation of the consonants (Crystal & House, 1988; Laeufer, 1992; Peterson & Lehiste, 1960; Tanner et al., 2019), stress (Davis & Van Summers, 1989; de Jong & Zawaydeh, 2002), position within the word (Ridouane, 2007), word length4 (Umeda, 1975), position within the phrase (Crystal & House, 1988; Laeufer, 1992; Luce & Charles-Luce, 1985; Umeda, 1975), and speech rate (Port, 1981; Solé, 2007; Tanner et al., 2019).
Other factors which influence vowel duration could also interact with voicing-conditioned vowel duration, even if previous work has not directly tested the interaction. Some of these factors are lexical frequency (Gahl et al., 2012; Munson & Solomon, 2004), use of real words versus nonce words (Conwell & Barta, 2018; Koob & Shea, 2024), the number of repetitions of each word (Bard et al., 2000; Fowler & Housum, 1987), and characteristics of the preceding consonant (Crystal & House, 1988; Fischer-Jørgensen, 1964). In some languages, the ratio of vowel duration before voiced versus voiceless consonants remains relatively consistent when combined with other factors like stress, speech rate, and inherent vowel duration, while other languages exhibit little proportional scaling (de Jong & Zawaydeh, 2002; Solé, 2007). Language-specific patterns like this mean that different phonological conditions could produce very different estimates of how much the size of voicing-conditioned vowel duration differences varies across languages.
Some of the studies reported here use real words (e.g., Mack, 1982 on English and French; Port et al., 1980 on Arabic; Lousada et al., 2010 on Portuguese), while others use a fixed structure which produces mostly or entirely nonce words (e.g., House & Fairbanks, 1953 on English; Beguš, 2017 on Georgian; Coretta, 2019 on Italian and Polish). Some use mostly real words with a few nonce words (e.g., Maddieson, 1977 on Armenian, Bengali, Hindi, and Marathi; Mohaghegh, 2011 on Farsi).
Most studies have vowels immediately preceding the obstruents of interest, but a few use a vowel+sonorant sequence (e.g., Mărdărescu-Teodorescu, 1995 on Romanian; Bárkányi & Kiss, 2009 on Hungarian), or a mix of words with immediately preceding vowels and words with intervening sonorants (e.g., Chen, 1970 on English; Port & Crawford, 1989 on German).
Many studies use words produced in a frame sentence (e.g., Laeufer, 1992 on English and French; Pape & Jesus, 2015 on German, Italian, and Portuguese; Beguš, 2017 on Georgian), but others use words in meaningful sentences (e.g., Davis & Van Summers, 1989 on English; Aguilar et al., 1997 on Catalan) or words in isolation (e.g., Mack, 1982 on English and French; Warner et al., 2004 on Dutch; Yu, 2004 on Lezgian). Some studies compare multiple contexts (e.g., Barry, 2003 on Russian; Sokolović-Perović, 2009 on Serbian).
Depending on the design of the study, any factors which can influence vowel duration might substantially influence the results if they are not consistent across stimuli. For example, Crystal & House (1988) demonstrate that the proportional size of the voicing effect is similar for tense and lax vowels in English, in contrast to their earlier study (Crystal & House, 1982). They note that their earlier study did not control for whether the words were prepausal or not and found little evidence for voicing-conditioned vowel duration in lax vowels.
Some studies use only exact minimal pairs (e.g., Sharf, 1962 on English; Port & Crawford, 1989 on German; Keating, 1980 on Czech and Polish; as well as most studies using nonce words). Others include some non-minimal pairs (e.g., Chen, 1970 on French, Russian, and Korean; Braunschweiler, 1997 on German), or systematically look at vowels with matched preceding and following consonants (e.g., House & Fairbanks, 1953 on English). Some of the studies do not describe the stimuli in enough detail to make clear whether they were exact minimal pairs or not (e.g., Savithri, 1986 on Kannada). Given the range of factors which contribute to vowel duration, non-minimal pairs might create the appearance of a larger or smaller difference in vowel duration than is actually caused by coda voicing. For example, the French items used by Chen (1970) are about half minimal pairs and half pairs which differ in additional ways; the duration ratio among the minimal pairs is 0.83, while the duration ratio in the other items is 0.9. However, using minimal pairs does not guarantee comparability across studies, given the range of factors which influence the size of the voicing effect.
An additional consideration in interpreting measurements of voicing-conditioned vowel duration is keeping in mind what exactly the “voicing” contrast is in different languages. Contrasts labelled as voicing or fortis/lenis often involve a bundle of several acoustic cues, e.g., Lisker (1986) identifies 16 correlates of intervocalic voicing in English stops. Phonetic evidence sometimes suggests that the fortis/lenis contrast is primarily a voicing contrast (e.g., Russian and Hungarian: Petrova et al., 2006), an aspiration contrast (e.g., German: J. Beckman et al., 2013; English: Keating, 1984), or might not neatly fall into either of these categories (e.g., Ozolotepec Zapotec: Leander, 2008). Phonetic details also vary among languages within each broad categorization, e.g., in the duration of the prevoicing (Souganidis et al., 2024) or the duration of aspiration (Cho & Ladefoged, 1999). The phonetic correlates of a fortis/lenis contrast can vary by position; in languages where unaspirated stops are not consistently voiced word-initially, intervocalic unaspirated stops are often partially voiced (e.g., German: J. Beckman et al., 2013) or even completely voiced (e.g., Farsi: Bijankhan & Nourbakhsh, 2009). Languages were only included in this typology if voicing during the consonant constriction seems to be a correlate of the laryngeal contrast in the environments being considered. This meant that some studies with vowel duration measurements before fortis and lenis consonants were not part of the analysis due to consonant information that indicated that voicing during the consonant constriction is not a consistent correlate of the fortis/lenis contrast: Lule Sami (Engstrand, 1987), Mongolian (Rogers, 2000), and Musey (Shryock, 1995).
Classification as a “true voicing” contrast is generally based on how consistently voicing is present in the lenis consonants (e.g., J. Beckman et al., 2013; Gao & Arai, 2019) and whether these lenis consonants trigger voicing assimilation (e.g., Jansen, 2004; Nicolae & Nevins, 2016). Different consonants within a language will not necessarily exhibit the same voicing behavior. The phonological patterns of laryngeal contrasts within a language can vary based on manner of articulation, e.g., word-initial English /z/ triggers voicing assimilation in preceding stops, whereas /d/ does not (Jansen, 2004). It is not clear whether surface voiced consonants would be expected to behave the same way as underlyingly voiced consonants in how they affect preceding vowel duration. The other side of the question of underlying versus surface voicing is what happens with devoiced consonants. In the current analysis, the voicing effect was evaluated separately for the neutralizing environment in languages which are described as having (nearly) categorical final devoicing. However, many languages have some degree of final devoicing. If voicing-conditioned vowel duration depends on phonetic voicing rather than phonological voicing, even gradient devoicing patterns might be expected to decrease the effect of voicing-conditioned vowel duration.
The phonetic realization of voicing contrasts can also vary across speakers of a language (e.g., English: Flege, 1982). Because of this variation, the broad characterization of a language is likely to be less informative than voicing measurements for the same tokens as the vowel duration measurements. Many studies do not report the degree of voicing in phonologically voiced versus voiceless consonants. Among those which do include this information, there is variation in how they report it. Some studies report the average proportion of voicing in the consonant (e.g., Russian: Kulikov, 2012; Spanish: Campos-Astorkiza, 2019; Turkish: Kopkallı, 1993), while others quantify voicing in other ways, such as what percentage of the consonants are voiced through at least half of their duration (e.g., Arabic: Flege & Port, 1981) or have voicing present at all (e.g., Lio: Miatto & Wivell, 2022).
Voicing during the consonant constriction is not the only correlate of voicing that might be relevant for voicing-conditioned vowel duration. For example, constriction duration is likely to be relevant. Some proposed explanations for the effect of voicing on vowel duration are based on compensatory timing (e.g., Coretta, 2019; Fowler, 1981; Kluender et al., 1988; Lehiste, 1970); such accounts would predict that languages with little difference in the duration of voiced versus voiceless consonants would also exhibit little difference in the duration of preceding vowels.
If some aspects of study design or the nature of voicing within a language have systematic effects, they might be apparent in a model. Table 15 presents the summary of a mixed effects linear regression model for the vowel duration ratio (before voiceless versus before voiced), excluding positions in which voicing is neutralized.5 The fixed effects were position in the word (final, medial), manner of articulation (stops, fricatives, affricates), status of gemination in the language (not contrastive, contrastive), the average duration of the vowels before voiceless consonants, and the language’s contrast type (true voicing, not true voicing). There was a random intercept for language. The model was built up from the model in Table 13. Each factor was added if it significantly improved the model, as tested with model comparison between models differing only in the presence of that factor. Studies were omitted if the value for any model factor was unclear. In the final model, there were 250 datapoints from 55 languages.
Table 15: Model for vowel duration ratio (before voiceless versus before voiced). Recall that a smaller vowel duration ratio indicates more of a difference in vowel duration based on the voicing of the following consonant. Reference Levels: Position = Final, Manner = Stops, GeminationStatus = Non-contrastive, ContrastType = NotTrueVoicing.
| β | SE | t-value | p-value | |
| (Intercept) | 0.74 | 0.02 | 37.1 | < 0.0001 |
| Position Medial | 0.047 | 0.013 | 3.6 | 0.00042 |
| Manner Affricates | –0.028 | 0.028 | –1.0 | 0.32 |
| Manner Fricatives | –0.046 | 0.013 | –3.7 | 0.00031 |
| GeminationStatus Contrastive | 0.044 | 0.015 | 3.0 | 0.0066 |
| V Dur Before Voiceless | 0.00047 | 0.00013 | 3.7 | 0.00026 |
| ContrastType TrueVoicing | 0.061 | 0.02 | 3.1 | 0.0073 |
There is limited evidence here for any methodological factors creating systematic effects in measurements of voicing-conditioned vowel duration. The lack of evidence does not necessarily mean that these factors have no effect; most crucially, some of them might introduce variability rather than consistently increasing or decreasing voicing-conditioned vowel duration. Null results might be due to insufficient data or confounding factors, particularly given that this is not a controlled dataset designed to test these effects.
All of the methodological factors listed for each study in the Supplementary Materials were tested as potential factors in the model if they exhibited sufficient variation across studies to be testable. Most of them did not improve the model and were not included in the final model: number of speakers in the study (χ2 = 2.9, df = 1, p = 0.089); number of tokens in the study (χ2 = 0.47, df = 1, p = 0.49); number of distinct words in the study (χ2 = 0.08, df = 1, p = 0.78); whether the vowel immediately preceded the obstruent with the voicing contrast (χ2 = 2.2, df = 2, p = 0.33); whether the vowel was stressed (χ2 = 0.54, df = 2, p = 0.76); number of syllables in the word (χ2 = 4.2, df = 3, p = 0.24); the broader context of the stimuli, e.g., isolation, frame sentence, meaningful sentences (χ2 = 5.3, df = 4, p = 0.26); whether the stimuli were real words or nonce words (χ2 = 4.5, df = 2, p = 0.1); and whether the stimuli were exact minimal pairs (χ2 = 0.066, df = 1, p = 0.8). Whether the stimuli were elicited or occurred in natural conversation was not tested as a factor because the vast majority of studies used elicited items. Whether the segment following the word-medial consonant of interest was a consonant or a vowel was also not tested, given that very few studies tested pre-consonantal environments.
The vowel duration ratio was significantly predicted by the way that the fortis/lenis contrast in the language is typically categorized. When compared to languages known to have a true voicing contrast, the effect of voicing-conditioned vowel duration was likely to be larger in languages which have been described as having a fortis/lenis contrast that is not primarily dependent on voicing. One important caveat here is that the categorization of languages is typically made at the language level, which has been followed here, e.g., English is categorized as not true voicing. However, phonetic evidence indicates that English word-final and word-medial lenis stops exhibit substantially more voicing than word-initial stops (Davidson, 2016). Another important caveat is that only seven of the languages in the dataset are in the “not true voicing” category, which is not enough for a clear generalization.
One possible interpretation is that this result reflects a tradeoff in cue weighting across languages: If voicing is a weaker cue to a fortis/lenis contrast, preceding vowel duration is likely to be a stronger cue. Friulian is described as having fully shifted from a voicing contrast in word-final obstruents to a length contrast in the preceding vowels (Baroni & Vanelli, 2000). A similar process has been proposed for varieties of English in which word-final stops are devoiced (Holt et al., 2016). This interpretation depends on all of these contrasts having historically been voicing contrasts, i.e., the initial development of vowel duration differences was triggered by voicing. Note that Friulian only exhibits a moderate effect of voicing on preceding vowel duration word-internally, where voicing of the consonants is maintained (for stops, an average ratio 0.78 and average difference 43 ms, vs. word-final ratio 0.47 and difference 136 ms).
A second possibility is that voicing-conditioned vowel duration is actually more directly influenced by characteristics other than voicing itself, e.g., duration of the consonant. Under this interpretation, some of the contrasts might have never consistently involved voicing, and the factor(s) driving the vowel duration differences might still be actively influencing it in languages without true voicing contrasts.
About half of the studies have consonant duration measurements, so it was added as a factor. This model is presented separately because including this factor requires excluding the studies which did not provide consonant duration, resulting in 151 datapoints from 46 languages. The much smaller dataset reduces or eliminates some of the effects which were clear in the models which did not require this subsetting. Table 16 presents the summary of a mixed effects linear regression model for the vowel duration ratio (before voiceless versus before voiced), excluding positions in which voicing is neutralized.6 The fixed effects were position in the word (final, medial), manner of articulation (stops, fricatives, affricates), status of gemination in the language (not contrastive, contrastive), the average duration of the vowels before voiceless consonants, the language’s contrast type (true voicing, not true voicing), and the consonant duration ratio. The continuous factors were centered. There was a random intercept for language.
Table 16: Model for vowel duration ratio (before voiceless versus before voiced). Reference Levels: Position = Final, Manner = Stops, GeminationStatus = Non-contrastive, ContrastType = NotTrueVoicing.
| β | SE | t-value | p-value | |
| (Intercept) | 0.74 | 0.026 | 28.6 | < 0.0001 |
| Position Medial | 0.059 | 0.017 | 3.5 | 0.00075 |
| Manner Affricates | –0.035 | 0.035 | –1.0 | 0.32 |
| Manner Fricatives | –0.035 | 0.015 | –2.3 | 0.022 |
| GeminationStatus Contrastive | 0.029 | 0.019 | 1.5 | 0.14 |
| V Dur Before Voiceless | 0.00077 | 0.00022 | 3.6 | 0.00051 |
| ContrastType TrueVoicing | 0.076 | 0.025 | 3.0 | 0.008 |
| ConsonantDurationRatio | –0.12 | 0.027 | –4.4 | < 0.0001 |
The ratio of constriction duration in voiceless versus voiced consonants is a significant predictor of the ratio of preceding vowel duration. Languages with a larger consonant duration ratio have a smaller vowel duration ratio. Note that because both ratios have the voiceless consonant (or consonant environment) as the numerator and voiceless consonants are typically longer than voiced consonants, a larger consonant duration ratio and smaller vowel duration ratio both reflect a larger difference in duration.
Previous studies have found a similar relationship between vowel duration and following consonant duration within an individual language (e.g., Beguš, 2017; de Jong, 1991). Some authors propose that this relationship reflects a tendency for consistency in word duration (Port et al., 1980, 1987).
There is no evidence that voicing-conditioned vowel duration is mediated by the extent of actual voicing during the consonant. Relatively few studies include measurements for proportion of the consonant that is voiced. Using the subset of studies which include these measurements, adding the ratio of proportion voicing in voiceless versus voiced consonants as a factor does not improve the model (χ2 = 0.33, df = 1, p = 0.57).
Table 17 presents the summary of a mixed effects linear regression model for the vowel duration difference (before voiceless minus before voiced), testing the same factors considered above for duration ratio.7 The fixed effects were position in the word (final, medial), manner of articulation (stops, fricatives, affricates), status of gemination in the language (not contrastive, contrastive), the average duration of the vowels before voiceless consonants, the language’s contrast type (true voicing, not true voicing), and the consonant duration difference. There was a random intercept for language. The continuous variables were centered. In the final model, there were 151 datapoints from 46 languages.
Table 17: Model for vowel duration difference (before voiceless minus before voiced). A more extreme negative value indicates more of a difference in vowel duration based on the voicing of the following consonant. Reference Levels: Position = Final, Manner = Stops, GeminationStatus = Non-contrastive, ContrastType = NotTrueVoicing.
| β | SE | t-value | p-value | |
| (Intercept) | –44.2 | 4.5 | –9.9 | < 0.0001 |
| Position Medial | 8.7 | 3.2 | 2.7 | 0.0083 |
| Manner Affricates | –5.9 | 6.9 | –0.85 | 0.39 |
| Manner Fricatives | –4.0 | 3.0 | –1.3 | 0.19 |
| GeminationStatus Contrastive | 8.4 | 3.4 | 2.4 | 0.022 |
| V Dur Before Voiceless | –0.11 | 0.042 | –2.6 | 0.011 |
| ContrastType TrueVoicing | 14.1 | 4.4 | 3.2 | 0.0056 |
| ConsonantDurationDiff | –0.24 | 0.068 | –3.5 | 0.0006 |
The effects of contrast type and consonant duration in this model are the same as observed above when measuring the voicing effect as a ratio.
The other factors which were tested were not included because they did not improve the model: number of speakers in the study (χ2 = 1.7, df = 1, p = 0.2); number of tokens in the study (χ2 = 0.12, df = 1, p = 0.73); number of distinct words in the study (χ2 = 1.8, df = 1, p = 0.18); whether the vowel immediately preceded the obstruent with the voicing contrast (χ2 = 1.8, df = 2, p = 0.42); whether the vowel was stressed (χ2 = 0.32, df = 2, p = 0.85); number of syllables in the word (χ2 = 3.7 df = 3, p = 0.29); the broader context of the stimuli, e.g., isolation, frame sentence, meaningful sentences (χ2 = 4.1, df = 4, p = 0.39); whether the stimuli were real words or nonce words (χ2 = 0.58, df = 2, p = 0.75); whether the stimuli were exact minimal pairs (χ2 = 0.45, df = 1, p = 0.5); and difference of proportion voicing in voiceless versus voiced consonants (χ2 = 2.3, df = 1, p = 0.13).
There remains an imbalance in how much of the data on voicing-conditioned vowel duration comes from Indo-European languages, which means that patterns common in Indo-European languages (particularly in the Balto-Slavic, Germanic, Indo-Iranian, and Italic branches) might have a disproportionate effect on the cross-linguistic patterns. Of the 76 languages described here, 39 are Indo-European and 37 come from other language families. Most of the Indo-European languages are represented by multiple studies, while most of the non-Indo-European languages are represented by a single study each.
3. Discussion
Although previous work often treats voicing-conditioned vowel duration as a categorical phenomenon, with languages either exhibiting evidence for it or not, the results here provide clear evidence that the size of the effect is a continuum. Across languages, the average ratio of vowel duration before voiceless stops to the vowel duration before voiced stops is 0.86 for word-medial stops, 0.83 for word-final stops, 0.83 for word-medial fricatives, and 0.77 for word-final fricatives. Recall that smaller ratios reflect larger effects. As a difference, the mean size of the effect is 17 ms for word-medial stops, 25 ms for word-final stops, 24 ms for word-medial fricatives, and 41 ms for word-final fricatives.
It has sometimes been proposed that some languages exhibit a particularly large effect because vowel duration is encoded as a controlled target in association with voicing, while others exhibit vowel duration patterns as they arise automatically (e.g., Chen, 1970; Fromkin, 1976; Solé, 2007). Some proposals also add a third category of languages which encode a lack of difference and exhibit little or no effect of voicing on preceding vowel duration (Cho, 2015). Such analyses predict a binary or ternary distribution: one cluster of languages with the effect as it would arise automatically, another cluster of languages with a categorical lengthening (or shortening) process, and potentially a third cluster of languages with representational specifications for vowel duration to be the same in both environments. However, the results here do not provide clear evidence for a bimodal or trimodal distribution. One potential interpretation of the distribution would be that there are two clusters when measuring the voicing effect as a ratio: one which contains only Breton and one which contains all of the other languages. While Breton might just be part of a long tail in this distribution, it is possible that Breton reflects a different representation of voicing-conditioned vowel duration than the other languages. For word-medial stops and fricatives, Breton has a much larger effect than any other language in the dataset when measured as a ratio (0.54 with stops, 0.57 with fricatives). Breton has been described as having a categorical vowel lengthening process, which is also reflected in the preservation of robust vowel duration differences based on underlying voicing of word-final voiceless versus devoiced obstruents (Le Dû, 1986).
The distribution of degree of voicing-conditioned vowel duration across languages seems to suggest that vowel duration is specified in the representation for particular languages in a gradient way. The distribution cannot be captured with lengthening or shortening rules involving a categorical length divide. The distribution is also broader than what would be expected for a purely automatic effect and more consistent within a language. Based on the distribution, we might imagine that the automatic factors driving voicing-conditioned vowel duration would best explain a duration ratio which is common across languages (i.e., a ratio around 0.84 or a difference around 23 ms), while larger effects are likely to be due to speakers exaggerating this cue to voicing and smaller effects might be due to speakers suppressing duration variation. If some languages exhibit an automatic effect while others exhibit a representationally specified effect, there might be a divide in how voicing-conditioned vowel duration interacts with other characteristics. Languages without specified targets for the voicing effect, if they exist, would be most informative for evaluating what drives the effect.
Voicing-conditioned vowel duration is one of the factors that is likely to contribute to characterizations of speech rhythm. Some work characterizes different languages as “stress-timed,” “syllable-timed” (e.g., Abercrombie, 1967, p.96–97; Pike, 1945), or “mora-timed” (Port et al., 1987). Later work demonstrates that rhythm patterns are continuous rather than falling into discrete categories (e.g., Dauer, 1983). Several phonetic measures are often used to compare the rhythm behavior of different languages, including the variability in duration of vocalic intervals and the proportion of vocalic intervals within each sentence (Ramus et al., 1999). Measures of rhythm which quantify variability in vowel duration will be impacted by voicing-conditioned vowel duration. Some work specifically connects rhythm classes to voicing-conditioned vowel duration (e.g., Port et al., 1987), noting that the combination of longer duration of voiceless consonants and shorter vowels before voiceless consonants produces largely consistent word durations.
3.1. Implications for underlying causes
Several different explanations have been proposed to explain voicing-conditioned vowel duration; the lack of consensus about what drives the effect is in large part due to the evidence not converging towards a single explanation. The patterns described above also provide mixed evidence. These varied results seem most consistent with the existence of multiple contributing factors (as proposed by e.g., Beguš, 2017; Coretta, 2020), as well as the duration differences becoming specified as controlled targets in some languages and increasing in size because of their usage as major cues to the voicing contrast.
One of the patterns in support of voicing-conditioned vowel duration being specified as controlled targets is a strong correlation in the size of the effect between manners of articulation. This relationship is consistent with preceding vowel duration being associated with the voicing feature in each language. If the differences between languages are mechanically driven, other language-specific differences would need to explain why the size of the effect differs by language, e.g., in how gestures are coordinated.
The effect of voicing on preceding vowel duration tends to be larger with fricatives than with stops. Voiced fricatives are more aerodynamically challenging than voiced stops; fewer languages have voiced fricatives than voiced stops (J. Ohala, 1983). Frication requires high airflow at the constriction, but the cyclic glottal closure during voicing decreases airflow. When airflow at the beginning of a voiced fricative is not sufficient to achieve frication, it could result in an approximant which may be indistinguishable from long formant transitions in the preceding vowel.
A few languages exhibit a larger effect with stops than with fricatives. There are five languages in which the voicing effect is larger before stops both when measured as a ratio and as a difference: Bulgarian, German, Japanese, Romanian, and Ozolotepec Zapotec. This result might reflect confounds in the stimuli, because many studies focused on effects of voicing do not have exactly parallel stimuli for stops and fricatives. For three of these languages, measurements come from a single study each. For German, the results vary in the three studies which include both stops and fricatives; two find a slightly larger voicing effect with fricatives (Malécot, 1971; Pape & Jesus, 2015) and one finds a much larger voicing effect with stops (Piroth & Janker, 2004). In Breton and Friulian, this direction of relative effect size is observed when measuring the voicing effect as a ratio but not when it is measured as a difference. In these languages, the larger voicing effect ratio with stops than with fricatives is likely to reflect manner of articulation also having a direct effect on vowel duration. Vowels tend to be longer before fricatives than before stops (Crystal & House, 1988; Laeufer, 1992; Peterson & Lehiste, 1960), and the proportional size of the voicing effect tends to be smaller with longer vowels, as shown in Table 13.
There are a few potential explanations for why the size of the voicing effect when measured as a ratio tends to be smaller with longer vowels (i.e., why vowel durations in voiced and voiceless environments are more similar when the absolute duration of the vowel is longer). One possibility is that durational effects of local articulatory pressures (e.g., Chen, 1970; Chomsky & Halle, 1968; Halle & Stevens, 1967; Öhman, 1967) are relatively consistent in absolute size and thus have a larger impact on vowel duration ratios when the vowels are shorter. For example, a 10 ms duration difference due to articulatory demands of transitioning to a voiceless versus voiced obstruent would be a much larger proportional difference in short vowels than in long vowels. Another possibility is that compensatory timing effects (e.g., Fowler, 1981; Lehiste, 1970) are based on approximate isochrony in syllable duration, e.g., a 10 ms longer consonant will result in a 10 ms shorter preceding vowel. This absolute trade-off in duration would result in smaller proportional effects in longer vowels.
On the other hand, the opposite patten is observed when the voicing effect is measured as a difference, i.e., vowel duration differences are smaller when the absolute duration of the vowel is shorter. If compensatory timing is based on expectations of a consistent ratio of vowel duration to consonant duration (cf. Port & Dalby, 1982), then long vowels and short vowels would be expected to exhibit voicing effects of equivalent proportional size but with a larger absolute size with long vowels than with short vowels.
The existence of these opposite patterns in whether long vowels or short vowels exhibit more of a voicing effect is consistent with multiple mechanisms driving voicing-conditioned vowel duration. The opposite patterns are also apparent across languages in the varied ways that the voicing effect interacts with other factors which influence vowel duration. Some languages are relatively consistent in the absolute size of voicing-conditioned vowel duration, while others are more consistent in the relative size (i.e., scaling proportionally when other factors influence vowel duration). Some analyses propose that proportional scaling exists when voicing-conditioned vowel duration is specified as a controlled representational target (Cho, 2015; de Jong & Zawaydeh, 2002; Solé, 2007). However, based on comparisons between phonologically long and short vowels, a larger effect does not correlate with scaling proportionally. Some languages have a small effect that is nonetheless proportionally consistent (e.g., Telugu, with a ratio of 0.94 both with short vowels and long vowels, Reddy, 1988), and other languages have a moderately sized effect that is nearly consistent in absolute size rather than scaling proportionally (e.g., Dutch, with a difference of 22 ms in long vowels and 20 ms in short vowels, Warner et al., 2004). The patterns within this data are potentially consistent with a moderate effect being automatic, while small effects are targets specified by the representation. However, there are also languages with voicing-conditioned vowel duration patterns which are neither consistent in absolute size nor in size relative to the duration of the vowel (see Tables 9 and 10).
Vowels are longer before singleton consonants than before geminates. Proposed explanations for this effect are sometimes similar to proposed explanations for voicing-conditioned vowel duration. Some accounts are based on compensatory timing, particularly in perception. If C/V ratio is a major cue to the contrast between geminates and singletons, then changing the vowel duration helps to cue the length of the consonant (Pickett et al., 1999). Another account is that speakers balance amount of effort across syllables, with longer durations involving more effort (Ridouane, 2007). An alternative account is that the relationship between gemination and vowel duration is the result of closed syllable shortening, with geminates being ambisyllabic (Maddieson, 1985). This explanation would not have implications for voicing-conditioned vowel duration. However, the existence of vowel duration differences conditioned by word-final geminates versus singletons poses a problem for closed-syllable shortening as an explanation (Ridouane, 2007).
Potential explanations for what drives voicing-conditioned vowel duration make different predictions about how voicing and gemination should interact in their effects on preceding vowel duration. However, none of these accounts entirely align with the observed lack of evidence for a consistent interaction. Slowing caused by the precise articulatory adjustments required for voiced obstruents (Chomsky & Halle, 1968; Halle & Stevens, 1967) might predict that the effect of voicing on vowel duration should be greater with geminates than with singletons, because needing to maintain the voiced obstruent for a longer period of time adds an additional challenge which might be expected to increase the time needed for making the articulatory adjustments. Because vowels tend to be shorter before geminates, many other explanations would also predict a larger effect with geminates than with singletons. This includes explanations based on local changes as the vocal tract transitions to voicelessness (e.g., Chen, 1970; Hansson, 2003; Öhman, 1967), as well as secondary developments via other characteristics which influence perceived duration (Javkin, 1976; Sanker, 2020). Compensatory timing (e.g., Fowler, 1981; Lehiste, 1970) might predict a smaller effect with geminates than with singletons. The relative consonant duration differences between voiced versus voiceless geminates is smaller than the relative differences between voiced versus voiceless singletons (Al-Tamimi & Khattab, 2018; Ghosh, 2015; Kawahara, 2006; Rojczyk & Porzuczek, 2019; Shigemori & Vietti, 2019), so vowel duration differences based on those consonant durations should also be proportionally smaller with geminates.
The lack of clear interaction between voicing and gemination further supports the existence of multiple driving factors. Compensatory timing would need to be one of the mechanisms, given that it is the only one which predicts larger proportional voicing effects with singletons than with geminates. Some studies find a larger voicing effect with singletons than with geminates (e.g., Chamoru: Santos, 2021; Kabyle: A. Elias, 2020), but many find little difference (e.g., Arabic: Alamri, 2022; Pajak, 2013; Bengali: Ghosh, 2015; Italian: Dian et al., 2024; Di Benedetto & De Nardis, 2021) or different behavior based on manner of articulation or position (e.g., Bulgarian: Maneva, 1997; Farsi: Zirak & Skaer, 2013; Polish: Malisz, 2013; Tashlhiyt Berber: Ridouane, 2007). Given that most of these languages are represented by a single study each, some of the variation might be due to characteristics of the particular stimulus items rather than a general pattern within the language.
In contrast to the correlation across manners of articulation, there is no evidence for a correlation in the size of the voicing effect across position in the word. The size of voicing-conditioned vowel duration differences word-finally is not predictive of the size of the effect word-medially in the same language. This distinction is important for making comparisons across studies both within a language and across languages; the pattern in one position in the word is not informative as a point of comparison for a different position, and pooling across positions can obscure what patterns exist in a language.
Differing by environment can be explained easily if vowel duration is specified in the representation; position in the word can be part of the environment conditioning vowel duration. In most studies, the word-medial environment was intervocalic, so position in the word generally also aligns with position in the syllable: word-final coda consonants versus word-medial onset consonants. In many languages, there is some degree of intervocalic lenition (e.g., Katz, 2021; Kirchner, 1998), which might also contribute to differences based on position. Some languages exhibit more of an effect word-finally, though other languages exhibit more of an effect word-medially (see Figure 3). The existence of multiple contributing factors might explain how different word-medial and word-final patterns could arise.
Word-final consonants are codas, while word-medial consonants are likely to be onsets, which will produce different gestural coordination patterns. A coda consonant is more directly coordinated with a preceding vowel than is an onset consonant. Because of these differences, accounts of voicing-conditioned vowel duration based on gesture timing (e.g., Lehiste, 1970; Lindblom, 1967; Maddieson, 1999) predict larger effects of consonant voicing on preceding vowel duration for codas and subsequently larger effects of word-final consonants than intervocalic consonants. Languages with larger voicing effects word-finally are consistent with an origin in gesture timing.
In general, word-medial vowels are shorter than vowels in word-final syllables due to final lengthening (e.g., Chen, 1970; Dutta & Kenstowicz, 2018; Elert, 1964; Gráczi, 2012; Maddieson, 1977; Ridouane, 2007; Sharf, 1962). A mechanism which would produce an effect with a consistent absolute size would predict larger proportional differences in shorter vowels, and thus would predict a larger effect word medially when measured as a ratio. The languages with larger effects word-medially when measured as a ratio are consistent with explanations based on local changes as the vocal tract transitions to voicelessness (e.g., Chen, 1970; Hansson, 2003; Javkin, 1976; Öhman, 1967; Sanker, 2020). A compensatory timing mechanism based on relative duration of consonant duration to vowel duration (e.g., Port & Dalby, 1982) could predict a larger voicing effect word-finally than word-medially when measured as a difference, at least in languages where word-final lengthening impacts vowels more than consonants.
The duration of consonants in different positions might also contribute to differences in voicing-conditioned vowel duration by position. In English, vowels in word-final syllables are much longer than word-medial vowels, regardless of stress, while consonants exhibit smaller effects of position (Oller, 1973) and exhibit different patterns depending on whether or not they are in a stressed syllable (Stathopoulos & Weismer, 1983). Across languages with longer word-final consonants than word-medial consonants, the size of the difference varies (e.g., Bengali: Dutta & Kenstowicz, 2018; Kanashi: Saxena et al., 2022; Lezgian: Yu, 2004; Tashlhiyt Berber: Ridouane, 2007). Some languages do not exhibit consistent differences in consonant duration by position (e.g., Swedish: Helgason et al., 2013). Some studies do not control for stress position or number of syllables, which could create a confound.
The difference in duration of voiced versus voiceless consonants does not align with the overall duration patterns by position, though some languages do have a proportionally similar effect in both positions and a larger absolute difference in the position where consonants and vowels are longer (e.g., Tashlhiyt Berber: Ridouane, 2007). It is most common for languages to exhibit a larger effect of voicing on consonant duration word-finally, both as a difference and as a ratio (e.g., Bengali: Dutta & Kenstowicz, 2018; Lezgian: Yu, 2004; English: Stathopoulos & Weismer, 1983; Swedish: Helgason et al., 2013). However, some languages exhibit a larger effect of voicing on consonant duration word-medially both as a difference and as a ratio (e.g., Kanashi: Saxena et al., 2022). Following a compensatory timing account, a larger difference in consonant duration in one position could predict larger voicing-conditioned vowel duration effects in that position. However, the by-position consonant duration differences associated with voicing do not seem to explain the by-position variation in voicing-conditioned vowel duration. As shown in Tables 16 and 17, including consonant duration ratio or difference as a predictor of the voicing effect does not eliminate the effect of position.
The transparency of intervening sonorant consonants as observed in English and several other languages is a challenge for many potential accounts of voicing-conditioned vowel duration, because most mechanisms would not extend for such a long distance. Compensatory timing is the explanation which best accounts for this longer-distance effect. However, if the voicing effect exists purely automatically in some languages, we might find that intervening sonorants block the effect in these languages and that transparency of sonorants only exists when voicing-conditioned vowel duration is specified as a representational target.
3.2. Implication for study design
Studies on the same language sometimes obtain substantially different results. Unusually large or small effects for a language in the summary data here might be due to methodological design more than characteristics of the language, particularly for languages which are represented by a single study. The variable results within a language highlight the importance of considering a range of factors which can influence vowel duration.
Some studies pool results across different positions in a word or different manners of articulation. Given that both factors have a substantial impact on the degree of voicing-conditioned vowel duration, the results here demonstrate the importance of controlling for these factors. Some factors are so frequently pooled that they could not be examined as potential variables here, e.g., vowel height.
The size of voicing-conditioned vowel duration differences in each language might be biased by tendencies in what methods are used to investigate those languages, which might focus on conditions in which the voicing effect is largest. For example, there is a large amount of work on effects of word-final stops in English but much less on word-medial consonants. As noted above, the effect in English is large with word-final stops (a ratio below 0.7 in most studies) but smaller with word-medial stops (average ratio 0.89). In other languages, most work investigating voicing-conditioned vowel duration looks at word-medial consonants, usually because of restrictions on word-final consonants (e.g., Italian). In the collection of data presented here, some of the studies were not investigating voicing-conditioned vowel duration and only incidentally included the necessary measurements; such studies are particularly useful for establishing a more representative distribution of how voicing-conditioned vowel duration behaves.
The results here do not provide any clear evidence for preferring to measure voicing-conditioned vowel duration as a ratio between vowel durations in each voicing environment or as a difference. The main consideration in this choice is whether a particular method is better at capturing how the voicing effect interacts with other influences on vowel duration. As demonstrated here and in previous work, languages vary in whether voicing-conditioned duration differences or duration ratios are more consistent across vowels which also differ in duration based on other factors (de Jong & Zawaydeh, 2002; Solé, 2007), and neither method produces completely consistent measurements within a language. In the models tested here, the one factor which behaves differently based on how the voicing effect is measured is how overall vowel duration predicts the size of the voicing effect. Longer vowel durations predicted ratios closer to 1 (i.e., less of an effect with longer vowels) but larger differences (i.e., more of an effect with longer vowels). This means that methodological differences in factors like vowel height and speech rate will have different effects on estimates of the voicing effect depending on how it is measured.
Ideally, studies should always give the raw duration measurements for each voicing category, to facilitate comparisons across studies. Because the choice of how to measure the size of the voicing effect produces different results depending on the absolute vowel durations within a particular study, this choice can substantially influence how a particular language appears to behave. For example, Lio has a small voicing effect for word-medial stops when measured as a difference (9 ms), but is at the cross-linguistic average when measured as a ratio (0.86). Javanese has the largest voicing effect for word-medial stops when measured as a difference (64 ms), but is much more similar to the cross-linguistic average when measured as a ratio (0.8).
4. Conclusions
This paper compiles data on voicing-conditioned vowel duration from a large number of languages and adds new data from languages for which this effect has not previously been addressed. This typological data helps establish what aspects of the effect are consistent or varied across languages and within a language, which provides a foundation for future work on this topic as well as providing additional lines of evidence for evaluating what might underlie the effect.
Appendix
This section gives the details of the new data from the UCLA Phonetics Lab Archive (UCLA Department of Linguistics, 2007). Voicing minimal pairs were identified based on the transcriptions accompanying the recordings in the UCLA Phonetics Lab Archive database. In Q’eqchi’ and Gujarati, some of the word-final voiced stops are realized as partially or completely voiceless. If a word appeared more than once within the word list, only the first one was included, in order to maintain the same number of tokens for each item within each minimal pair. For all of the languages, the voicing minimal pairs described here are a subset of a larger list of items.
Degema
Data comes from 10 speakers. Eight produced 12 words (6 minimal pairs) in isolation: /àdá, àtá, míìdà, míìtà, mʊ̀dá, mʊ̀tá, mɔdá, mɔtá, mɛ̀dá, mɛ̀tá, màdá, màtá/. Two produced 10 of those words: /míìdà, míìtà, mʊ̀dá, mʊ̀tá, mɔdá, mɔtá, mɛ̀dá, mɛ̀tá, màdá, màtá/. This produced a total of 116 tokens.
The speakers participated as a group; each speaker said each item in turn. Most of these words are part of the same two verb paradigms, and were elicited adjacent to each other.
Gujarati
Data comes from 4 speakers. One produced 6 words (3 minimal pairs) in isolation twice: /əndər, əntər, adro, atro, kaɡədri, kakədri/. The minimal pairs were elicited adjacent to each other. Two produced 72 words (36 minimal pairs) in isolation, once at fast speech rate and once at slow speech rate. The items were C1VC2 with each combination of: C1 as voiced, voiceless unaspirated, and voiced aspirated labial, alveolar, and velar stops and postalveolar affricates; V as /a, i, u/; and C2 as /p/ and /b/. One produced 4 words (2 minimal pairs) in isolation: /vad, vat, naɡ, nak/. The minimal pairs were elicited adjacent to each other. This produced a total of 304 tokens.
Ibibio
Data comes from 3 speakers. One produced 2 words (1 minimal pair), /àdá, àtá/, 6 times in isolation and 2 words (1 minimal pair), /díbé, dípé/, twice in isolation. One produced 2 words (1 minimal pair), /àdá, àtá/, 4 times in isolation and 2 words (1 minimal pair), /díbé, dípé/, twice in isolation. One produced 8 words (4 minimal pairs) twice in isolation: /ádú, átú, àdû, àtû, ádá, átá, ádâ, átâ/. This produced a total of 44 tokens.
Igbo
Data comes from 5 speakers. One produced 6 words (3 minimal pairs) in isolation: /ébà, épà, édà, étà, égà, ékà/. Two produced 2 words (1 minimal pair), /íba, ípa/, twice in isolation and 2 words (1 minimal pair), /íga, íka/, once in isolation. Two produced 2 words (1 minimal pair) twice in isolation: /ideꜜ, iteꜜ/. This produced a total of 26 tokens.
Kapampangan
Data comes from 3 speakers. One produced 4 words (2 minimal pairs) in isolation twice, once alongside the translation and once as a list: /atab, atap, lumod, lumot/. One produced 4 words (2 minimal pairs) in isolation once: /imbis, impis, abugadu, abukadu/. One produced 2 words (1 minimal pair) in isolation twice: /atab, atap/. The minimal pairs were elicited adjacent to each other. This produced a total of 16 tokens.
Q’eqchi’
Data comes from 11 speakers. All of them produced 6 words (3 minimal pairs) in a frame sentence: /tobok, topok, tob, top, hob, hop/. However, one of the speakers did not produce both words of the medial stop minimal pair, so the paired item was also omitted for that speaker. This produced a total of 64 tokens.
Tagalog
Data comes from 4 speakers. One produced 10 words (5 minimal pairs) in isolation: /ibon, ipon, tabon, tapon, baga, baka, bagal, bakal, sukoʔ, sugoʔ/. One produced 8 words (4 minimal pairs) in isolation: /bəɡah, bəkah, talob, talop, bɯkod, bɯkot, usoɡ, usok/. One produced 4 words (2 minimal pairs) twice in isolation: /ibon, ipon, sugat, sukat/. One produced 2 words (1 minimal pair) in isolation: /abo, apo/. This produced a total of 28 tokens.
Toda
Data comes from 13 speakers. Six produced 20 words (10 minimal pairs) both in isolation and in a frame sentence: /poːb, poːp, ed, et, ud, ut, kab, kap, pæːd, pæːt, teg, tek, tog, tok, toːg, toːk, kjeb, kjep, nib, nip/. Seven produced 12 words (6 minimal pairs) in isolation: /poːb, poːp, ed, et, kab, kap, teg, tek, toːg, toːk, kjeb, kjep/. This produced a total of 324 tokens.
The speakers participated as a group; each speaker said each item in turn.
Supplementary Materials
The full dataset used for analysis is included as a CSV file. It includes the acoustic measurements as well as methodological information about each study. There are also CSV files for vowel duration separated by gemination and phonological vowel length for the studies which presented data with each phonological length divide. https://doi.org/10.16995/labphon.19725.s1.
Competing Interests
The author has no competing interests to declare.
Notes
- Learned language-specific characteristics under speaker control are sometimes considered to be phonological. However, it is not the goal of this paper to argue for where a terminological divide between ‘phonetics’ and ‘phonology’ should be drawn. [^]
- lmer(VDurRatio∼Position + Manner + ContrastiveGemination + VDur.before.voiceless.POOLEDcentered + (1|Language), data = CrossLingDuration[which(CrossLingDuration$Manner !=ʺstopsAndFricativesʺ & CrossLingDuration$Manner != ʺStopsAndFricativesAndAffricatesʺ & CrossLingDuration$Position != ʺpooledʺ & CrossLingDuration$NeutralizedHere.== ʺNoʺ),]). [^]
- lmer(VDurDiff∼Position + Manner + ContrastiveGemination + VDur.before.voiceless.POOLEDcentered + (1|Language), data = CrossLingDuration[which(CrossLingDuration$Manner != ʺstopsAndFricativesʺ & CrossLingDuration$Manner != ʺStopsAndFricativesAndAffricatesʺ & CrossLingDuration$Position != ʺpooledʺ & CrossLingDuration$NeutralizedHere.== ʺNoʺ),]). [^]
- In many studies, word length and position within the word are confounded, with comparisons between word-final consonants in monosyllabic words and word-medial consonants in disyllabic words (e.g., Maddieson, 1977; Port, 1981; Sharf, 1962), leaving it somewhat unclear which of these factors is the main contributor to differences in voicing-conditioned vowel duration. [^]
- lmer(VDurRatio∼Position + Manner + ContrastiveGemination + VDur.before.voiceless.POOLEDcentered + TrueVoicing + (1|Language), data = CrossLingDuration[which(CrossLingDuration$Manner != ʺstopsAndFricativesʺ & CrossLingDuration$Manner != ʺStopsAndFricativesAndAffricatesʺ & CrossLingDuration$Position != ʺpooledʺ & CrossLingDuration$NeutralizedHere.== ʺNoʺ & CrossLingDuration$TrueVoicing != ʺunclearʺ),]). [^]
- lmer(VDurRatio∼Position + Manner + ContrastiveGemination + VDur.before.voiceless.POOLEDcentered + TrueVoicing + CDurRatiocentered + (1|Language), data = CrossLingDuration[which(CrossLingDuration$Manner != ʺstopsAndFricativesʺ & CrossLingDuration$Manner != ʺStopsAndFricativesAndAffricatesʺ & CrossLingDuration$Position != ʺpooledʺ & CrossLingDuration$NeutralizedHere.== ʺNoʺ & CrossLingDuration$TrueVoicing != ʺunclearʺ & CrossLingDuration$CDurRatio !=ʺ ʺ),]). [^]
- lmer(VDurDiff∼Position + Manner + ContrastiveGemination + VDur.before.voiceless.POOLEDcentered + TrueVoicing + CDurDiffcentered + (1|Language), data = CrossLingDuration[which(CrossLingDuration$Manner != ʺstopsAndFricativesʺ & CrossLingDuration$Manner != ʺStopsAndFricativesAndAffricatesʺ & CrossLingDuration$Position != ʺpooledʺ & CrossLingDuration$NeutralizedHere.== ʺNoʺ & CrossLingDuration$TrueVoicing != ʺunclearʺ & CrossLingDuration$CDurDiff != ʺ ʺ),]). [^]
References
Aasmäe, N., Pajusalu, K., & Kabajeva, N. (2016). Gemination in the Mordvin languages. Linguistica Uralica, 52(2), 81–92. http://doi.org/10.3176/lu.2016.2.01
Abdelli-Beruh, N. B. (2004). The stop voicing contrast in French sentences: Contextual sensitivity of vowel duration, closure duration, voice onset time, stop release and closure voicing. Phonetica, 61(4), 201–219. http://doi.org/10.1159/000084158
Abercrombie, D. (1967). Elements of general phonetics. Edinburgh University Press.
Abolhasanizadeh, V., Bijankhan, M., & Gussenhoven, C. (2012). The Persian pitch accent and its retention after the focus. Lingua, 122(13), 1380–1394. http://doi.org/10.1016/j.lingua.2012.06.002
Abunima, S., Raihan Syed Jaafar, S., & Hamid, S. A. (2021). The influence of bilingualism on the production of plosive sounds in L1 (Arabic) and L2 (English): Acoustic analysis of the duration of preceding and following vowels. Academic Journal of Modern Philology, 14, 25–44. http://doi.org/10.34616/ajmp.2021.14
Adhikari, S. (2018). Effect of aspiration on vowel duration for voice and voiceless unaspirated (Garhwali Hindi) consonants. BMC Journal of Scientific Research, 2(1), 1–6. http://doi.org/10.3126/bmcjsr.v2i1.42725
Aguilar, L., Giménez, J. A., Machuca, M., Marı́n, R., & Riera, M. (1997). Catalan vowel duration. Proceedings of the 5th European Conference on Speech Communication and Technology (Eurospeech), 771–774. http://doi.org/10.21437/Eurospeech.1997-261
Ahmed, Z. O. (2019). The application of English theories to Sorani phonology. [Doctoral dissertation, Durham University].
Al-Gamdi, N. A. (2022). Voicing contrast in Najdi Arabic stops: Implications for laryngeal realism. [Doctoral dissertation, Newcastle University].
Al-Tamimi, J., & Khattab, G. (2018). Acoustic correlates of the voicing contrast in Lebanese Arabic singleton and geminate stops. Journal of Phonetics, 71, 306–325. http://doi.org/10.1016/j.wocn.2018.09.010
Alamri, S. (2022). Articulatory properties of Saudi Arabic coronals and cross-language assimilation: The case of Saudi Arabic and Seoul Korean listeners. [Doctoral dissertation, George Mason University].
Alhabshan, M., & Alsager, H. N. (2022). The effect of coda voicing contrast on vowel duration in American English and Najdi Arabic: A comparison study. Journal of Language Teaching and Research, 13(5), 1048–1057. http://doi.org/10.17507/jltr.1305.18
Alves, A., de Lucena, R. M., & Alves, U. K. (2023). Duração de vogais antecedentes a consoantes oclusivas na variedade paraibana do português brasileiro. Letrônica, 16(1), Article e42103. http://doi.org/10.15448/1984-4301.2023.1.42103
Avelino, H. (2001). The phonetic correlates of fortis-lenis in Yalálag Zapotec consonants. [Master’s thesis, University of California, Los Angeles].
Badyal, P. (2026). Timing the difference: A study of gemination in Dogri consonants. Acta Universitatis Carolinae Philologica, 2025(3), 177–194. http://doi.org/10.14712/24646830.2025.27
Bard, E. G., Anderson, A. H., Sotillo, C., Aylett, M., Soherty-Sneddon, G., & Newlands, A. (2000). Controlling the intelligibility of referring expressions in dialogue. Journal of Memory and Language, 42, 1–22. http://doi.org/10.1006/jmla.1999.2667
Bárkányi, Z., & Beňuš, Š. (2015). Prosodic conditioning of pre-sonorant voicing. Proceedings of the 18th International Congress of Phonetic Sciences, Glasgow, Scotland, Article 0336.
Bárkányi, Z., & Kiss, Z. (2009). Hungarian v: Is it voiced? In M. den Dikken & R. M. Vago (Eds.), Approaches to Hungarian, Volume 11: Papers from the New York conference (pp. 1–28). John Benjamins. http://doi.org/10.1075/atoh.11.02bar
Bárkányi, Z., & Kiss, Z. G. (2015). Why do sonorants not voice in Hungarian? And why do they voice in Slovak? Approaches to Hungarian: Volume 14: Papers from the Piliscsaba Conference, 65–94. http://doi.org/10.1075/atoh.14.03bar
Baroni, M., & Vanelli, L. (2000). The relationship between vowel length and consonantal voicing in Friulian. In L. Repetti (Ed.), Phonological theory and the dialects of Italy (pp. 13–44). John Benjamins. http://doi.org/10.1075/cilt.212.04bar
Barry, S. M. E. (2003). Temporal correlates of the voicing contrast in Russian. [Doctoral dissertation, University College London].
Bates, D., Mächler, M., Bolker, B., & Walker, S. (2015). Fitting linear mixed-effects models using lme4. Journal of Statistical Software, 67(1), 1–48. http://doi.org/10.18637/jss.v067.i01
Beckman, J., Jessen, M., & Ringen, C. (2013). Empirical evidence for laryngeal features: Aspirating vs. True voice languages. Journal of Linguistics, 49(2), 259–284. http://doi.org/10.1017/S0022226712000424
Beckman, M. (1982). Segment duration and the “mora” in Japanese. Phonetica, 39(2–3), 113–135. http://doi.org/10.1159/000261655
Beguš, G. (2017). Effects of ejective stops on preceding vowel duration. Journal of the Acoustical Society of America, 142(4), 2168–2184. http://doi.org/10.1121/1.5007728
Berkovits, R. (1993). Progressive utterance-final lengthening in syllables with final fricatives. Language and Speech, 67(1), 89–98. http://doi.org/10.1177/002383099303600105
Bhaskararao, P., & Prasad, M. (1986). Length of Marathi /ə/–an instrumental study. Bulletin of the Deccan College Research Institute, 45, 1–6.
Bijankhan, M., & Nourbakhsh, M. (2009). Voice onset time in Persian initial and intervocalic stop production. Journal of the International Phonetic Association, 39(3), 335–364. http://doi.org/10.1017/S0025100309990168
Bouarourou, F., Koya, T., Bouzidi, S., Vaxelaire, B., & Sock, R. (2018). Cross-language and language-specific acoustic correlates of gemination in Berber and Japanese. Proceedings of the 2nd International Conference on Natural Language and Speech Processing (ICNLSP), 1–6. http://doi.org/10.1109/ICNLSP.2018.8374394
Braunschweiler, N. (1997). Integrated cues of voicing and vowel length in German: A production study. Language and Speech, 40(4), 353–376. http://doi.org/10.1177/002383099704000403
Buder, E. H., & Stoel-Gammon, C. (2002). American and Swedish children’s acquisition of vowel duration: Effects of vowel identity and final stop voicing. Journal of the Acoustical Society of America, 111(4), 1854–1864. http://doi.org/10.1121/1.1463448
Campos-Astorkiza, R. (2008). Two sources of voicing neutralization in Lithuanian. Proceedings of the 2nd ISCA Workshop on Experimental Linguistics, 49–52. http://doi.org/10.36505/ExLing-2008/02/0013/000072
Campos-Astorkiza, R. (2012). Length contrast and contextual modifications of duration in the Lithuanian vowel system. Baltic Linguistics, 3, 9–41. http://doi.org/10.32798/bl.418
Campos-Astorkiza, R. (2019). Modeling assimilation: The case of sibilant voicing in Spanish. In M. Gibson & J. Gil (Eds.), Romance phonetics and phonology (pp. 241–275). Oxford University Press. http://doi.org/10.1093/oso/9780198739401.003.0014
Chalise, K. P. (2021). Duration of [a] and [i] as preceding vowels with plosives in Nepali. Interdisciplinary Journal of Linguistics, 14, 50–63.
Chan, L. X. (2021). Acoustic correlates of final stop voicing in Pashto. ICU Working Papers in Linguistics, 17, 1–14.
Chen, M. (1970). Vowel length variation as a function of the voicing of the consonant environment. Phonetica, 22, 129–159. http://doi.org/10.1159/000259312
Cho, T. (2015). Language effects on timing at the segmental and suprasegmental levels. In M. Redford (Ed.), The handbook of speech production (pp. 505–529). John Wiley. http://doi.org/10.1002/9781118584156.ch22
Cho, T., & Ladefoged, P. (1999). Variation and universals in VOT: Evidence from 18 languages. Journal of Phonetics, 27, 207–229. http://doi.org/10.1006/jpho.1999.0094
Chomsky, N., & Halle, M. (1968). The sound pattern of English. Harper & Row.
Clayards, M. (2008). The ideal listener: Making optimal use of acoustic-phonetic cues for word recognition. [Doctoral dissertation, University of Rochester].
Conwell, E., & Barta, K. (2018). Phrase position, but not lexical status, affects the prosody of noun/verb homophones. Frontiers in Psychology, 9, Article 1785. http://doi.org/10.3389/fpsyg.2018.01785
Coretta, S. (2019). An exploratory study of voicing-related differences in vowel duration as compensatory temporal adjustment in Italian and Polish. Glossa, 4(1), Article 125. http://doi.org/10.5334/gjgl.869
Coretta, S. (2020). Vowel duration and consonant voicing: A production study. [Doctoral dissertation, University of Manchester].
Costa, F. (2004). Intrinsic prosodic properties of stressed vowels in European Portuguese. Proceedings of Speech Prosody, Nara, Japan, 53–56. http://doi.org/10.21437/SpeechProsody.2004-12
Craioveanu, R. (2023). Weighing preaspiration: Phonetics, phonology, & typology of a laryngeal phenomenon. [Doctoral dissertation, University of Toronto].
Crystal, T. H., & House, A. S. (1982). Segmental durations in connected speech signals: Preliminary results. Journal of the Acoustical Society of America, 72(3), 705–716. http://doi.org/10.1121/1.388251
Crystal, T. H., & House, A. S. (1988). Segmental durations in connected-speech signals: Current results. Journal of the Acoustical Society of America, 83(4), 1553–1573. http://doi.org/10.1121/1.395911
Dauer, R. M. (1983). Stress-timing and syllable-timing reanalyzed. Journal of Phonetics, 11(1), 51–62. http://doi.org/10.1016/S0095-4470(19)30776-4
Davidson, L. (2016). Variability in the implementation of voicing in American English obstruents. Journal of Phonetics, 54, 35–50. http://doi.org/10.1016/j.wocn.2015.09.003
Davis, S., & Van Summers, W. (1989). Vowel length and closure duration in word-medial VC sequences. Journal of Phonetics, 17(4), 339–353. http://doi.org/10.1016/S0095-4470(19)30449-8
De Jong, K. (1991). An articulatory study of consonant-induced vowel duration changes in English. Phonetica, 48, 1–17. http://doi.org/10.1159/000261868
De Jong, K. (2004). Stress, lexical focus, and segmental focus in English: Patterns of variation in vowel duration. Journal of Phonetics, 32, 493–516. http://doi.org/10.1016/j.wocn.2004.05.002
De Jong, K., & Zawaydeh, B. (2002). Comparing stress, lexical focus, and segmental focus: Patterns of variation in Arabic vowel duration. Journal of Phonetics, 30, 53–75. http://doi.org/10.1006/jpho.2001.0151
Delattre, P. (1962). Some factors of vowel duration and their cross-linguistic validity. Journal of the Acoustical Society of America, 34(8), 1141–1143. http://doi.org/10.1121/1.1918268
Derib Ado Jekale. (2011). An acoustic analysis of Amharic vowels, plosives and ejectives. [Doctoral dissertation, Addis Ababa University].
Di Benedetto, M. G., & De Nardis, L. (2021). Consonant gemination in Italian: The affricate and fricative case. Speech Communication, 134, 86–108. http://doi.org/10.1016/j.specom.2021.07.005
Dian, A., Hajek, J., & Fletcher, J. (2024). Cross-regional patterns of obstruent voicing and gemination: The case of Roman and Veneto Italian. Languages, 9(12), Article 383. http://doi.org/10.3390/languages9120383
Dinnsen, D. A., & Charles-Luce, J. (1984). Phonological neutralization, phonetic implementation and individual differences. Journal of Phonetics, 12, 49–60. http://doi.org/10.1016/S0095-4470(19)30850-2
Dipino, D., & Celata, C. (2018). An UTI study of alveolar stops in Italian. In A. Vietti, L. Spreafico, D. Mereu, & V. Galatà (Eds.), Il parlato nel contesto naturale: Speech in the natural context (pp. 41–53). Associazione Italiana Scienze della Voce.
Durvasula, K., & Luo, Q. (2014). Voicing, aspiration, and vowel duration in Hindi. In S. Ohlsson & R. Catrambone (Eds.), Proceedings of meetings on acoustics. Acoustical Society of America. http://doi.org/10.1121/1.4895027
Dutta, H., & Kenstowicz, M. (2018). The phonology and phonetics of laryngeal stop contrasts in Assamese. In R. Petrosino, P. Cerrone, & H. van der Hulst (Eds.), From sounds to structures: Beyond the veil of Maya (pp. 30–64). Mouton de Gruyter. http://doi.org/10.1515/9781501506734-002
Elert, C.-C. (1964). Phonologic studies of quantity in Swedish based on material from Stockholm speakers. Almqvist & Wiksells Boktryckeri AB.
Elias, F. (2017). An acoustic analysis of vowel duration in Wolaytta Doonaa. [Doctoral dissertation, Addis Ababa University].
Elias, A. (2020). Kabyle “double” consonants: Long or strong? McGill Working Papers in Linguistics, 26.
Engstrand, O. (1987). Preaspiration and the voicing contrast in Lule Sami. Phonetica, 44(2), 103–116. http://doi.org/10.1159/000261784
Falc’Hun, F. (1951). Le système consonantique du breton avec une étude comparative de phonétique expérimentale. Annales de Bretagne et des pays de l’Ouest, 57, 5–194. http://doi.org/10.3406/abpo.1950.1893
Feda Negesse, & Tujube Amansa. (2021). Durational variations in Oromo vowels. In D. Ado, A. W. Gelagay, & J. B. Johannessen (Eds.), Grammatical and sociolinguistic aspects of Ethiopian languages (pp. 345–363). John Benjamins. http://doi.org/10.1075/impact.48.14neg
Ferrat, K., & Guerti, M. (2015). Acoustic and articulatory analysis of the gemination in Modern Standard Arabic. Al-Lisaniyyat, 21(2), 71–88. http://doi.org/10.61850/allj.v21i2.306
Finco, F. (2007). La durata delle vocali friulane: Risultati di un’indagine fonetica. In F. Vicario (Ed.), Ladine loqui: IV colloquium retoromanistich. Società Filologica Friulana.
Fintoft, K. (1961). The duration of some Norwegian speech sounds. Phonetica, 7, 19–39. http://doi.org/10.1159/000258096
Fischer-Jørgensen, E. (1964). Sound duration and place of articulation. Zeitschrift für Phonetik, Sprachwissenschaft und Kommunikationsforschung, 17, 175–208. http://doi.org/10.1524/stuf.1964.17.16.175
Flege, J. E. (1982). Laryngeal timing and phonation onset in utterance-initial English stops. Journal of Phonetics, 10(2), 177–192. http://doi.org/10.1016/S0095-4470(19)30956-8
Flege, J. E., & Port, R. (1981). Cross-language phonetic interference: Arabic to English. Language and Speech, 24(2), 124–146. http://doi.org/10.1177/002383098102400202
Fourakis, M., & Iverson, G. K. (1984). On the “incomplete neutralization” of German final obstruents. Phonetica, 41, 140–149. http://doi.org/10.1159/000261720
Fowler, C. A. (1981). A relationship between coarticulation and compensatory shortening. Phonetica, 38, 35–50. http://doi.org/10.1159/000260013
Fowler, C. A. (1992). Vowel duration and closure duration in voiced and unvoiced stops: There are no contrast effects here. Journal of Phonetics, 20, 143–165. http://doi.org/10.1016/S0095-4470(19)30244-X
Fowler, C. A., & Housum, J. (1987). Talkers’ signaling of “new” and “old” words in speech and listeners’ perception and use of the distinction. Journal of Memory and Language, 26, 489–504. http://doi.org/10.1016/0749-596X(87)90136-7
Frąckowiak-Richter, L. (1973). The duration of Polish vowels. Speech Analysis and Synthesis, 3, 87–115.
Fromkin, V. (1976). Some questions regarding universal phonetics and phonetic representations. In A. Juilland, A. M. Devine, & L. D. Stephens (Eds.), Linguistic studies offered to Joseph Greenberg on the occasion of his sixtieth birthday (pp. 365–380). Anma Libri.
Gahl, S., Yao, Y., & Johnson, K. (2012). Why reduce? Phonological neighborhood density and phonetic reduction in spontaneous speech. Journal of Memory and Language, 66, 789–806. http://doi.org/10.1016/j.jml.2011.11.006
Gao, J., & Arai, T. (2019). Plosive (de-)voicing and f0 perturbations in Tokyo Japanese: Positional variation, cue enhancement, and contrast recovery. Journal of Phonetics, 77, Article 100932. http://doi.org/10.1016/j.wocn.2019.100932
Gautam, B. R. (2012). Phonation types of Balami consonants: An acoustic analysis. Nepalese Linguistics, 27, 42–53.
Ghadessy, E. (1988). Some problems of American students in mastering Persian phonology. [Doctoral dissertation, University of Texas, Austin].
Ghosh, A. (2015). Acoustic correlates of voiced and voiceless geminates in Bangla. [Master’s thesis, The English and Foreign Languages University, Hyderabad].
Ghowail, T. I. (1987). The acoustic phonetic study of the two pharyngeals /ḥ,ʕ/ and the two laryngeals /ʔ,h/ in Arabic. [Doctoral dissertation, Indiana University].
Gráczi, T. E. (2011). Voicing contrast of intervocalic plosives in Hungarian. Proceedings of the 17th International Congress of Phonetic Sciences, Hong Kong, 759–762.
Gráczi, T. E. (2012). Zörejhangok akusztikai fonetikai vizsgálata a zöngésségi oppozı́ció függvényében. [Doctoral dissertation, Eötvös Loránd University].
Gráczi, T. E., & Bárkányi, Z. (2013). Voiced fricatives in Hungarian. In Grammatika és kontextus: Új szempontok az uráli nyelvek kutatásában III (pp. 81–96).
Hajek, J., & Cummins, T. (2006). A preliminary investigation of vowel lengthening in non-final position in Friulian. In P. Warren & C. I. Watson (Eds.), Proceedings of the 11th Australasian International Conference on Speech Science and Technology (pp. 239–242).
Halle, M., & Stevens, K. (1967). On the mechanism of glottal vibration for vowels and consonants. MIT Research Laboratory of Electronics, Quarterly Progress Reports, 85, 267–271. http://doi.org/10.1121/1.2143736
Hallé, P. A., Segui, J., & Androjna, K. (2015). Primary and secondary cues to voice assimilation in French and Slovenian. In The Scottish Consortium for ICPhS 2015 (Ed.), Proceedings of the International Conference on Phonetic Sciences (ICPhS) XVIII.
Hamel, P. (1983). Brazilian Portuguese stressed vowels: A durational study. Kansas Working Papers in Linguistics, 8(1), 31–46. http://doi.org/10.17161/KWPL.1808.471
Han, J.-I. (2000). Intervocalic stop voicing revisited. Speech Sciences, 7(1), 203–216.
Hansson, G. Ó. (2003). Laryngeal licensing and laryngeal neutralization in Faroese and Icelandic. Nordic Journal of Linguistics, 26(1), 45–79. http://doi.org/10.1017/S0332586503001008
Helgason, P., Ringen, C., & Suomi, K. (2013). Swedish quantity: Central Standard Swedish and Fenno-Swedish. Journal of Phonetics, 41(6), 534–545. http://doi.org/10.1016/j.wocn.2013.09.005
Herbert, R. K. (1986). Language universals, markedness theory, and natural phonetic processes. Mouton de Gruyter. http://doi.org/10.1515/9783110865936
Holt, Y. F., Jacewicz, E., & Fox, R. A. (2016). Temporal variation in African American English: The distinctive use of vowel duration. Journal of Phonetics and Audiology, 2(2), Article 1000121. http://doi.org/10.4172/2471-9455.1000121
Homma, Y. (1981). Durational relationship between Japanese stops and vowels. Journal of Phonetics, 9(3), 273–281. http://doi.org/10.1016/S0095-4470(19)30971-4
Hosseinpoor Damirchian, R., & Nourbakhsh, M. (2022). The effect of voicing on constriction duration, voice duration and vowel duration in stop and fricative consonants of Turkish language in Tabrizi dialect. Language Related Research, 13(2), 105–136.
House, A. S. (1961). On vowel duration in English. Journal of the Acoustical Society of America, 33(9), 1174–1178. http://doi.org/10.1121/1.1908941
House, A. S., & Fairbanks, G. (1953). The influence of consonant environment upon the secondary acoustical characteristics of vowels. Journal of the Acoustical Society of America, 25(1), 105–113. http://doi.org/10.1121/1.1906982
Hussein, L. (1994). Voicing-dependent vowel duration in standard Arabic and its acquisition by adult American students. [Doctoral dissertation, The Ohio State University].
Ishibashi, S., & Koya, S. (2023). The effects of voiced stops on adjacent vowel duration in Japanese. In R. Skarnitzl & J. Volı́n (Eds.), Proceedings of the International Conference on Phonetic Sciences (ICPhS) XX (pp. 733–737).
Jacques, B. (1990). Étude de trois indices acoustiques du voisement des consonnes fricatives en français de Montréal. Revue Québécoise de Linguistique, 19(2), 59–71. http://doi.org/10.7202/602676ar
Jansen, W. (2004). Laryngeal contrast and phonetic voicing: A laboratory phonology approach to English, Hungarian, and Dutch. [Doctoral dissertation, University of Groningen].
Jassem, W., & Richter, L. (1989). Neutralization of voicing in Polish obstruents. Journal of Phonetics, 17(4), 317–325. http://doi.org/10.1016/S0095-4470(19)30447-4
Jatteau, A., Audibert, N., Vasilescu, I., Lamel, L., & Adda-Decker, M. (2025). Final devoicing before it happens: A large-scale study of word-final obstruents in French. Laboratory Phonology, 16. http://doi.org/10.16995/labphon.10855
Javkin, H. R. (1976). The perceptual basis of vowel duration differences associated with the voiced/voiceless distinction. Report of the Berkeley Phonology Laboratory, 1, 78–92. http://doi.org/10.1121/1.2002209
Jones, G. E. (1972). Hyd llafariaid yn y Gymraeg. Studia Celtica, 7, 120–129.
Jongmans, P. (2008). The intelligibility of tracheoesophageal speech: An analytic and rehabilitation study. [Doctoral dissertation, University of Amsterdam].
Kang, K.-S. (1999). The phonetics and phonology of Korean and Thai obstruents. [Doctoral dissertation, State University of New York, Buffalo].
Katsanidi, G. (2013). The acquisition of Greek vowels by pre-school and primary school children. [Doctoral dissertation, Aristotle University of Thessaloniki].
Katz, J. (2021). Intervocalic lenition is not phonological: Evidence from Campidanese Sardinian. Phonology, 38(4), 651–692. http://doi.org/10.1017/S095267572100035X
Kawahara, S. (2006). A faithfulness ranking projected from a perceptibility scale: The case of [+voice] in Japanese. Language, 82(3), 536–574. http://doi.org/10.1353/lan.2006.0146
Keating, P. (1980). A phonetic study of a voicing contrast in Polish. [Doctoral dissertation, Brown University].
Keating, P. (1984). Phonetic and phonological representation of stop consonant voicing. Language, 60(2), 286–319. http://doi.org/10.2307/413642
Keating, P. (1985). Universal phonetics and the organization of grammars. In V. A. Fromkin (Ed.), Phonetic linguistics: Essays in honor of Peter Ladefoged (pp. 115–131). Academic Press.
Kenstowicz, M. J. (2021). Phonetic correlates of the Javanese voicing contrast in stop consonants. NUSA Jurnal Ilmu Bahasa Dan Sastra, 70, 1–37.
Khalid, N., & Kiani, Z. H. (2022). Acoustic analysis of Dawoodi consonants: A severely endangered language of Northern Pakistan. Pakistan Languages and Humanities Review, 6(4), 474–492. http://doi.org/10.47205/plhr.2022(6-IV)43
Kharlamov, V. (2014). Incomplete neutralization of the voicing contrast in word-final obstruents in Russian: Phonological, lexical, and methodological influences. Journal of Phonetics, 43, 47–56. http://doi.org/10.1016/j.wocn.2014.02.002
Kirchner, R. M. (1998). An effort-based approach to consonant lenition. [Doctoral dissertation, University of California, Los Angeles].
Kluender, K. R., Diehl, R. L., & Wright, B. A. (1988). Vowel-length differences before voiced and voiceless consonants: An auditory explanation. Journal of Phonetics, 16, 153–169. http://doi.org/10.1016/S0095-4470(19)30480-2
Kohler, K. J. (1979). Dimensions in the perception of fortis and lenis plosives. Phonetica, 36(4–5), 332–343. http://doi.org/10.1159/000259970
Koob, R., & Shea, C. (2024). The influence of abstract phonological processes on the acquisition of a foreign language – an example of German, Spanish, and English. In D. Olson, J. Sturm, O. Dmitrieva, & J. Levis (Eds.), Proceedings of 14th Pronunciation in Second Language Learning and Teaching Conference. Iowa State University Digital Press. http://doi.org/10.31274/psllt.17563
Kopkallı, H. (1993). A phonetic and phonological analysis of final devoicing in Turkish. [Doctoral dissertation, University of Michigan].
Kopkallı-Yavuz, H. (2000). Acoustic analysis of voicing contrast in Turkish stops. In A. Göksel & C. Kerslake (Eds.), Studies on Turkish and Turkic languages (pp. 19–25). Harrassowitz Verlag.
Kulikov, V. (2012). Voicing and voice assimilation in Russian stops. [Doctoral dissertation, University of Iowa].
Kuznetsova, A., Bruun Brockhoff, P., & Haubo Bojesen Christensen, R. (2015). lmerTest: Tests in linear mixed effects models [Computer software]. https://CRAN.R-project.org/package=lmerTest
Laeufer, C. (1992). Patterns of voicing-conditioned vowel duration in French and English. Journal of Phonetics, 20(4), 411–440. http://doi.org/10.1016/S0095-4470(19)30648-5
Landschultz, K. (1967). Duration of French vowels before fricatives. Annual Report of the Institute of Phonetics University of Copenhagen, 2, 109–118. http://doi.org/10.7146/aripuc.v2i.130677
Le Dû, J. (1986). A sandhi survey of the Breton language. In Sandhi phenomena in the languages of Europe (pp. 435–450). Mouton de Gruyter. http://doi.org/10.1515/9783110858532.435
Leander, A. J. (2008). Acoustic correlates of fortis/lenis in San Francisco Ozolotepec Zapotec. [Master’s thesis, University of North Dakota].
Lehiste, I. (1970). Temporal organization of spoken language. Working Papers in Linguistics, 12, 95–114.
Lindau-Webb, M. (1985). Hausa vowels and diphthongs. Studies in African Linguistics, 16(2), 161–182. http://doi.org/10.32473/sal.v16i2.107504
Lindblom, B. (1967). Vowel duration and a model of lip mandible coordination. Speech Transmission Laboratory Quarterly Progress Status Report, 8(4), 1–29.
Lisker, L. (1974). On “explaining” vowel duration variation. Glossa: An International Journal of Linguistics, 8, 233–246.
Lisker, L. (1986). “Voicing” in English: A catalogue of acoustic features signaling /b/ versus /p/ in trochees. Language and Speech, 29(1), 2–11. http://doi.org/10.1177/002383098602900102
Loporcaro, M., Delucchi, R., Nocchi, N., Paciaroni, T., Schmid, S., Savy, R., & Crocco, C. (2006). La durata consonantica nel dialetto di Lizzano in Belvedere (Bologna). In R. Savy & C. Crocco (Eds.), Analisi prosodica. Teorie, modelli e sistemi di annotazione (pp. 491–517). EDK Editore. http://doi.org/10.5167/uzh-33336
Lousada, M., Jesus, L. M. T., & Hall, A. (2010). Temporal acoustic correlates of the voicing contrast in European Portuguese stops. Journal of the International Phonetic Association, 40(3), 261–275. http://doi.org/10.1017/S0025100310000186
Luce, P. A., & Charles-Luce, J. (1985). Contextual effects on vowel duration, closure duration, and the consonant/vowel ratio in speech production. Journal of the Acoustical Society of America, 78(6), 1949–1957. http://doi.org/10.1121/1.392651
Mack, M. (1982). Voicing-dependent vowel duration in English and French: Monolingual and bilingual production. Journal of the Acoustical Society of America, 71(1), 173–178. http://doi.org/10.1121/1.387344
Maddieson, I. (1977). Further studies on vowel length before aspirated consonants. UCLA Working Papers in Phonetics, 38, 82–90.
Maddieson, I. (1985). Phonetic cues to syllabification. In V. Fromkin (Ed.), Phonetic linguistics (pp. 203–221). Academic Press.
Maddieson, I. (1999). Phonetic universals. In W. J. Hardcastle & J. Laver (Eds.), The handbook of phonetic sciences (pp. 619–639). Blackwell. http://doi.org/10.1111/b.9780631214786.1999.00020.x
Magdics, K. (1969). Studies in the acoustic characteristics of Hungarian speech sounds. Indiana University Press.
Malécot, A. (1971). The general phonetic characteristics of languages, final report (DHEW-OEC-0-70-23). U.S. Department of Health, Education, and Welfare. Office of Education Institute of International Studies.
Malisz, Z. (2013). Speech rhythm variability in Polish and English: A study of interaction between rhythmic levels. [Doctoral dissertation, Adam Mickiewicz University].
Malisz, Z., & Klessa, K. (2008). A preliminary study of temporal adaptation in Polish VC groups. Proceedings of Speech Prosody 2008, 383–386. http://doi.org/10.21437/SpeechProsody.2008-87
Maneva, B. (1997). Modèle temporel de la gémination consonantique en bulgare. [Doctoral dissertation, Université de Montréal].
Mărdărescu-Teodorescu, M. (1995). Analiza acustică a vocalelor româneşti ı̂n context nazal. Fonetică Şi Dialectologie, 14, 29–46.
Marković, M. (2007). Kontrastivna analiza akustičkih i artikulacionih karakteristika vokalskog sistema engleskog i srpskog jezika. [Doctoral dissertation, University of Novi Sad].
Maslowski, M., Meyer, A. S., & Bosker, H. R. (2019). Listeners normalize speech for contextual speech rate even without an explicit recognition task. Journal of the Acoustical Society of America, 146(1), 179–188. http://doi.org/10.1121/1.5116004
Meynadier, Y., & Gaydina, Y. (2012). Contraste de voisement en parole chuchotée. In L. Besacier, B. Lecouteux, & G. Sérasset (Eds.), Proceedings of the Joint Conference JEP-TALN-RECITAL 2012 (pp. 361–368).
Miatto, V., & Wivell, G. B. (2022). Stop voicing contrast in Lio. Proceedings of Meetings on Acoustics, 45, Article 060006. http://doi.org/10.1121/2.0001547
Mikuteit, S. (2007). A cross-language approach to voice, quantity and aspiration: An East-Bengali and German production study. [Doctoral dissertation, University of Konstanz].
Mikuteit, S., & Reetz, H. (2007). Caught in the ACT: The timing of aspiration and voicing in East Bengali. Language and Speech, 50(2), 247–277. http://doi.org/10.1177/00238309070500020401
Mitleb, F. (1984). Voicing effect on vowel duration is not an absolute universal. Journal of Phonetics, 12, 23–27. http://doi.org/10.1016/S0095-4470(19)30847-2
Mitterer, H. (2018). The singleton-geminate distinction can be rate dependent: Evidence from Maltese. Laboratory Phonology, 9(6). http://doi.org/10.5334/labphon.66
Mohaghegh, M. (2011). Voicing in Persian word-final obstruents. Canadian Acoustics, 39(3), 158–159.
Moon, S.-J. (2017). Vowel length difference before voiced/voiceless consonants in English and Korean. Phonetics and Speech Sciences, 9(4), 35–41. http://doi.org/10.13064/KSSS.2017.9.4.035
Moosmüller, S., & Ringen, C. (2004). Voice and aspiration in Austrian German plosives. Folia Linguistica, 38(1–2), 43–62. http://doi.org/10.1515/flin.2004.38.1-2.43
Munson, B., & Solomon, N. P. (2004). The effect of phonological neighborhood density on vowel articulation. Journal of Speech, Language, and Hearing Research, 47, 1048–1058. http://doi.org/10.1044/1092-4388(2004/078)
Myers, S. (2010). Regressive voicing assimilation: Production and perception studies. Journal of the International Phonetic Association, 40(2), 163–179. http://doi.org/10.1017/S0025100309990284
Nagano-Madsen, Y. (1992). Mora and prosodic coordination: A phonetic study of Japanese, Eskimo and Yoruba. Lund University Press.
Nı́ Chasaide, A. (1985). Preaspiration in phonological stop contrasts: An instrumental phonetic study. [Doctoral dissertation, University College of North Wales].
Nicolae, A. C., & Nevins, A. (2016). Fricative patterning in aspirating versus true voice languages. Journal of Linguistics, 52(1), 151–174. http://doi.org/10.1017/S0022226715000067
Nirgianaki, E., Kontostavlaki, A., Nikolaenkova, O., & Papanagiotou, M. (2017). Syllable phonology and cross-syllable temporal production in Greek. In A. Botinis (Ed.), Proceedings of the 8th ExLing (pp. 81–84). http://doi.org/10.36505/ExLing-2017/08/0021/000323
Ohala, J. (1983). The origin of sound patterns in vocal tract constraints. In P. F. MacNeilage (Ed.), The production of speech (pp. 189–216). Springer Verlag. http://doi.org/10.1007/978-1-4613-8202-7_9
Ohala, M., & Ohala, J. J. (1992). Phonetic universals and Hindi segment duration. In J. Ohala, T. Nearey, B. Derwing, M. Hodge, & G. Wiebe (Eds.), Proceedings of the International Conference on Spoken Language Processing, Banff, 12–16 October 1992 (pp. 831–834). http://doi.org/10.21437/ICSLP.1992-272
Öhman, S. (1967). Peripheral motor commands in labial articulation. Speech Transmission Laboratory Quarterly Progress Status Report, 30–63.
Oller, D. K. (1973). The effect of position in utterance on speech segment duration in English. Journal of the Acoustical Society of America, 54(5), 1235–1247. http://doi.org/10.1121/1.1914393
Pajak, B. (2013). Non-intervocalic geminates: Typology, acoustics, perceptibility. San Diego Linguistics Papers, 4, 2–27.
Pape, D., & Jesus, L. M. T. (2015). Stop and fricative devoicing in European Portuguese, Italian and German. Language and Speech, 58(2), 224–246. http://doi.org/10.1177/0023830914530604
Peterson, G. E., & Lehiste, I. (1960). Duration of syllable nuclei in English. Journal of the Acoustical Society of America, 32(6), 693–703. http://doi.org/10.1121/1.1908183
Petrova, O., Plapp, R., Ringen, C., & Szentgyörgyi, S. (2006). Voice and aspiration: Evidence from Russian, Hungarian, German, Swedish, and Turkish. Linguistic Review, 23(1), 1–35. http://doi.org/10.1515/TLR.2006.001
Pfiffner, A. M. (2021). Cue-based features: Modeling change and variation in the voicing contrasts of Minnesotan English, Afrikaans, and Dutch. [Doctoral dissertation, Georgetown University].
Pickett, E. R., Blumstein, S. E., & Burton, M. W. (1999). Effects of speaking rate on the singleton/geminate consonant contrast in Italian. Phonetica, 56(3–4), 135–157. http://doi.org/10.1159/000028448
Pike, K. L. (1945). The intonation of American English. University of Michigan Press.
Piroth, H. G., & Janker, P. M. (2004). Speaker-dependent differences in voicing and devoicing of German obstruents. Journal of Phonetics, 32(1), 81–109. http://doi.org/10.1016/S0095-4470(03)00008-1
Podlipský, V. J., & Chládková, K. (2007). Vowel duration variation induced by coda consonant voicing and perceptual short/long vowel categorization in Czech. In R. Vı́ch (Ed.), Speech processing: 17th Czech-German workshop (pp. 68–74).
Port, R. (1979). The influence of tempo on stop closure duration as a cue for voicing and place. Journal of Phonetics, 7(1), 45–56. http://doi.org/10.1016/S0095-4470(19)31032-0
Port, R. (1981). Linguistic timing factors in combination. Journal of the Acoustical Society of America, 69(1), 262–274. http://doi.org/10.1121/1.385347
Port, R., Al-Ani, S., & Maeda, S. (1980). Temporal compensation and universal phonetics. Phonetica, 37, 235–252. http://doi.org/10.1159/000259994
Port, R., & Crawford, P. (1989). Incomplete neutralization and pragmatics in German. Journal of Phonetics, 17(4), 257–282. http://doi.org/10.1016/S0095-4470(19)30444-9
Port, R., & Dalby, J. (1982). Consonant/vowel ratio as a cue for voicing in English. Perception & Psychophysics, 32(2), 141–152. http://doi.org/10.3758/BF03204273
Port, R., Dalby, J., & O’Dell, M. (1987). Evidence for mora timing in Japanese. Journal of the Acoustical Society of America, 81(5), 1574–1585. http://doi.org/10.1121/1.394510
Prusty, V. R., & Venkatesh, L. (2012). Duration of vowels in Oriya: A developmental perspective. Journal of the All India Institute of Speech & Hearing, 31, 10–22.
R Core Team. (2021). R: A language and environment for statistical computing [Computer software]. R Foundation for Statistical Computing. https://www.R-project.org/
Ramus, F., Nespor, M., & Mehler, J. (1999). Correlates of linguistic rhythm in the speech signal. Cognition, 73(3), 265–292. http://doi.org/10.1016/S0010-0277(99)00058-X
Raphael, L. J. (1975). The physiological control of durational differences between vowels preceding voiced and voiceless consonants in English. Journal of Phonetics, 3(1), 25–33. http://doi.org/10.1016/S0095-4470(19)31284-7
Rashid, H. U., & Khan, A. Q. (2014). A phonemic and acoustic analysis of Hindko fricatives. Acta Linguistica Asiatica, 4(3), 71–81. http://doi.org/10.4312/ala.4.3.71-81
Recasens, D., & Espinosa, A. (2007). An electropalatographic and acoustic study of affricates and fricatives in two Catalan dialects. Journal of the International Phonetic Association, 37(2), 143–172. http://doi.org/10.1017/S0025100306002829
Reddy, K. N. (1988). The duration of Telugu speech sounds: An acoustic study. IETE Journal of Research, 34(1), 57–63. http://doi.org/10.1080/03772063.1988.11436705
Ridouane, R. (2007). Gemination in Tashlhiyt Berber: An acoustic and articulatory study. Journal of the International Phonetic Association, 37(2), 119–142. http://doi.org/10.1017/S0025100307002903
Ringen, C., & Van Dommelen, W. A. (2013). Quantity and laryngeal contrasts in Norwegian. Journal of Phonetics, 41(6), 479–490. http://doi.org/10.1016/j.wocn.2013.09.001
Rogers, H. (2000). The effect of rate of speech on laryngeal timing in medial stops in Mongolian. In D. G. Lockwood, P. H. Fries, & J. E. Copeland (Eds.), Functional approaches to language, culture and cognition (pp. 391–401). John Benjamins. http://doi.org/10.1075/cilt.163.29rog
Rojczyk, A., & Porzuczek, A. (2019). Durational properties of Polish geminate consonants. Journal of the Acoustical Society of America, 146(6), 4171–4182. http://doi.org/10.1121/1.5134782
Sanker, C. (2020). A perceptual pathway for voicing-conditioned vowel duration. Laboratory Phonology, 11(1), Article 18. http://doi.org/10.5334/labphon.268
Santos, T. J. M. (2021). I anakko’na sonada siha gi fino’Chamoru: A phonetic investigation of geminate production in Chamoru. [Doctoral dissertation, University of Guam].
Savithri, S. R. (1986). Durational analysis of Kannada vowels. Journal of Acoustical Society of India, 14(2), 34–31.
Saxena, A., Sjöberg, A., & Sagar, P. (2022). The sound system of Kanashi. In A. Saxena & L. Borin (Eds.), Synchronic and diachronic aspects of Kanashi (pp. 13–51). de Gruyter Mouton. http://doi.org/10.1515/9783110703245-002
Schwartz, G., Kaźmierski, K., & Wojtkowiak, E. (2021). Perspectives on final laryngeal neutralisation: New evidence from Polish. Phonology, 38, 693–727. http://doi.org/10.1017/S0952675721000373
Schwarz, M. R. (2024). Realization and representation of Nepali laryngeal contrasts. [Doctoral dissertation, University of California, Berkeley].
Searl, J. P., & Carpenter, M. A. (2002). Acoustic cues to the voicing feature in tracheoesophageal speech. Journal of Speech, Language, and Hearing Research, 45(2), 282–294. http://doi.org/10.1044/1092-4388(2002/022)
Seyfarth, S., & Garellek, M. (2018). Plosive voicing acoustics and voice quality in Yerevan Armenian. Journal of Phonetics, 71, 425–450. http://doi.org/10.1016/j.wocn.2018.09.001
Sharf, D. J. (1962). Duration of post-stress intervocalic stops and preceding vowels. Language and Speech, 5(1), 26–30. http://doi.org/10.1177/002383096200500103
Shigemori, L. S. B., & Vietti, A. (2019). Acoustic analysis of Italian singleton/geminate stop production in two ambient temperature conditions. In S. Calhoun, P. Escudero, M. Tabain, & P. Warren (Eds.), Proceedings of the International Congress of Phonetic Sciences (ICPhS) XIX (pp. 3676–3680). Australasian Speech Science; Technology Association Inc.
Shrager, M. (2012). Neutralization of word-final voicing in Russian. Journal of Slavic Linguistics, 20(1), 71–99. http://doi.org/10.1353/jsl.2012.0000
Shryock, A. M. (1995). Investigating laryngeal contrasts: An acoustic study of the consonants of Musey. [Doctoral dissertation, University of California, Los Angeles].
Slis, I. H., & Cohen, A. (1969). On the complex regulating the voiced-voiceless distinction II. Language and Speech, 12(3), 137–155. http://doi.org/10.1177/002383096901200301
Slowiaczek, L. M., & Dinnsen, D. A. (1985). On the neutralizing status of Polish word-final devoicing. Journal of Phonetics, 13, 325–341. http://doi.org/10.1016/S0095-4470(19)30763-6
Sokolović-Perović, M. (2009). Voicing-conditioned vowel duration in southern Serbian. Newcastle Working Papers in Linguistics, 15, 126–137.
Sokolović-Perović, M. (2012). The voicing contrast in Serbian stops. [Doctoral dissertation, Newcastle University].
Solé, M.-J. (2007). Controlled and mechanical properties in speech: A review of the literature. In M.-J. Solé, P. S. Beddor, & M. Ohala (Eds.), Experimental approaches to phonology (pp. 302–321). Oxford University Press. http://doi.org/10.1093/oso/9780199296675.003.0018
Sóskuthy, M. (2013). Phonetic biases and systemic effects in the actuation of sound change. [Doctoral dissertation, University of Edinburgh].
Souganidis, C., Molinaro, N., & Stoehr, A. (2024). Bilinguals produce language-specific voice onset time in two true-voicing languages: The case of Basque-Spanish early bilinguals. Linguistic Approaches to Bilingualism, 14(3), 370–399. http://doi.org/10.1075/lab.21081.sou
Stathopoulos, E. T., & Weismer, G. (1983). Closure duration of stop consonants. Journal of Phonetics, 11(4), 395–400. http://doi.org/10.1016/S0095-4470(19)30838-1
Strange, W., Weber, A., Levy, E. S., Shafiro, V., Hisagi, M., & Nishi, K. (2007). Acoustic variability within and across German, French, and American English vowels: Phonetic context effects. Journal of the Acoustical Society of America, 122(2), 1111–1129. http://doi.org/10.1121/1.2749716
Suomi, K. (1976). English voiceless and voiced stops as produced by native and Finnish speakers. Jyväskylä Contrastive Studies, 32, 1–91.
Száraz, B. (2019). Explozı́vák és az őket megelőző magánhangzók időtartama modális fonációval létrehozott beszédben és suttogásban. Beszédkutatás, 27(1), 75–86. http://doi.org/10.15775/Beszkut.2019.75-86
Tanner, J., Sonderegger, M., Stuart-Smith, J., & The SPADE Data Consortium. (2019). Vowel duration and the voicing effect across dialects of English. Toronto Working Papers in Linguistics, 41. http://doi.org/10.33137/twpl.v41i1.32769
Tieszen, B. J. (1997). Final stop devoicing in Polish: An acoustic and historical account for incomplete neutralization. [Doctoral dissertation, University of Wisconsin, Madison].
Tofigh, Z., & Abolhasanizadeh, V. (2015). Phonetic neutralization: The case of Persian final devoicing. Athens Journal of Philology, 2(2), 123–134. http://doi.org/10.30958/ajp.2-2-4
Tsuchida, A. (1997). Phonetics and phonology of Japanese vowel devoicing. [Doctoral dissertation, Cornell University].
UCLA Department of Linguistics. (2007). The UCLA Phonetics Lab Archive. http://archive.phonetics.ucla.edu/.
Umeda, N. (1975). Vowel duration in American English. Journal of the Acoustical Society of America, 58(2), 434–445. http://doi.org/10.1121/1.380688
van de Velde, D. J., & van Heuven, V. J. J. P. (2011). Compensatory strategies for voicing of initial and medial plosives and fricatives in whispered speech in Dutch. Proceedings of the International Conference on Phonetic Sciences (ICPhS) XVII, 2058–2061.
Van Dommelen, W. (1981). Kontextbedingte Vokaldehnung im Französischen. In W. J. Barry & K. J. Kohler (Eds.), Universität Kiel Arbeitsberichte 16: Beiträge zur experimentellen und angewandten Phonetik. (pp. 95–107).
Van Dommelen, W. (1982). A contrastive investigation of vowel duration in German and Dutch. Phonetica, 39, 23–35. http://doi.org/10.1159/000261648
Van Rooy, B., Wissing, D., & Paschall, D. D. (2003). Demystifying incomplete neutralisation during final devoicing. Southern African Linguistics and Applied Language Studies, 21(1–2), 49–66. http://doi.org/10.2989/16073610309486328
Van Santen, J. P. H. (1992). Contextual effects on vowel duration. Speech Communication, 11(6), 513–546. http://doi.org/10.1016/0167-6393(92)90027-5
Warner, N., Jongman, A., Sereno, J., & Kemps, R. (2004). Incomplete neutralization and other sub-phonemic duration differences in production and perception: Evidence from Dutch. Journal of Phonetics, 32, 251–276. http://doi.org/10.1016/S0095-4470(03)00032-9
Wissing, D. (1992). Vowel duration in Afrikaans: The influence of postvocalic consonant voicing and syllable structure. Journal of the Acoustical Society of America, 92(1), 589–592. http://doi.org/10.1121/1.404270
Yadav, R. (1979). The influence of aspiration on vowel duration in Maithili. South Asian Languages Analysis, 1, 157–165.
Yawney, H. L. (2025). Phonological and morphological conditions on dorsals in Kazakh: Productivity in production and acceptability. [Doctoral dissertation, University of Toronto].
Yimam Mohammed, & Derib Ado. (2024). A durational comparison of the vowels of Argobba dialects: Shonke and Gachine. Abyssinia Journal of Business and Social Sciences, 9(1), 102–114. http://doi.org/10.20372/ajbs.2024.9.1.1042
Yu, A. C. L. (2004). Explaining final obstruent voicing in Lezgian: Phonetics and history. Language, 80(1), 73–97. http://doi.org/10.1353/lan.2004.0049
Yu, A. C. L. (2008). The phonetics of quantity alternation in Washo. Journal of Phonetics, 36(3), 508–520. http://doi.org/10.1016/j.wocn.2007.10.004
Zhu, X. (1999). Shanghai tonetics. [Doctoral dissertation, Australian National University].
Zirak, M., & Skaer, P. M. (2013). Evidence of gemination in Persian: Phonetic and phonological study of lexical and post-lexical geminates. Hiroshima University Institutional Repository: Studies in Human Sciences, 8, 17–41. http://doi.org/10.15027/35640



