<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20120330//EN" "http://jats.nlm.nih.gov/publishing/1.2/JATS-journalpublishing1.dtd">
<!--<?xml-stylesheet type="text/xsl" href="article.xsl"?>-->
<article article-type="research-article" dtd-version="1.2" xml:lang="en" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">
<front>
<journal-meta>
<journal-id journal-id-type="issn">1868-6354</journal-id>
<journal-title-group>
<journal-title>Laboratory Phonology: Journal of the Association for Laboratory Phonology</journal-title>
</journal-title-group>
<issn pub-type="epub">1868-6354</issn>
<publisher>
<publisher-name>Open Library of Humanities</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.16995/labphon.9019</article-id>
<article-categories>
<subj-group>
<subject>Journal article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Development of a new vowel feature from coarticulation: Biomechanical modeling of rhotic vowels in Kalasha</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Mielke</surname>
<given-names>Jeff</given-names>
</name>
<email>jimielke@ncsu.edu</email>
<xref ref-type="aff" rid="aff-1">1</xref>
<xref ref-type="corresp" rid="cor-1">*</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Hussain</surname>
<given-names>Qandeel</given-names>
</name>
<email>qandeel.hussain@students.mq.edu.au</email>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Moisik</surname>
<given-names>Scott R.</given-names>
</name>
<email>scott.moisik@ntu.edu.sg</email>
<xref ref-type="aff" rid="aff-2">2</xref>
</contrib>
</contrib-group>
<aff id="aff-1"><label>1</label>Department of English, North Carolina State University, Raleigh, NC, USA</aff>
<aff id="aff-2"><label>2</label>Nanyang Technological University, Singapore</aff>
<author-notes>
<corresp id="cor-1"><label>*</label>Corresponding author.</corresp>
</author-notes>
<pub-date publication-format="electronic" date-type="pub" iso-8601-date="2023-08-11">
<day>11</day>
<month>08</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>14</volume>
<issue>1</issue>
<fpage>1</fpage>
<lpage>52</lpage>
<permissions>
<copyright-statement>Copyright: &#x00A9; 2023 The Author(s)</copyright-statement>
<copyright-year>2023</copyright-year>
<license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/4.0/">
<license-p>This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International License (CC-BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. See <uri xlink:href="http://creativecommons.org/licenses/by/4.0/">http://creativecommons.org/licenses/by/4.0/</uri>.</license-p>
</license>
</permissions>
<self-uri xlink:href="http://www.journal-labphon.org/articles/10.16995/labphon.9019/"/>
<abstract>
<p>Coarticulation is an important source of new phonological contrasts. When speakers interpret effects such as nasalization, glottalization, and rhoticization as an inherent property of a vowel, a new phonological contrast is born. Studying this process directly is challenging because most vowel systems are stable and phonological change likely follows a long transitional period in which coarticulation is conventionalized beyond its mechanical basis. We examine the development of a new vowel feature by focusing on the emergence of rhotic vowels in Kalasha, an endangered Dardic (Indo-Aryan) language, using biomechanical and acoustic modeling to provide a baseline of pure rhotic coarticulation.</p>
<p>Several features of the Kalasha rhotic vowel system are not predicted from combining muscle activation for non-rhotic vowels and bunched and retroflex approximants, including that rhotic back vowels are produced with tongue body fronting (shifting the backness contrast to principally a rounding contrast). We find that synthesized vowels that are about 30% plain vowel and 70% rhotic are optimal (i.e., they best approximate observed rhotic vowels and also balance the acoustic separation among rhotic vowels with the separation from their non-rhotic counterparts). Otherwise, dispersion is not generally observed, but the vowel that is most vulnerable to merger differs most from what would be expected from coarticulation alone.</p>
</abstract>
</article-meta>
</front>
<body>
<sec>
<title>1. Introduction</title>
<p>This is an investigation of what happens when a language develops an entirely new vowel feature. It is difficult to observe how various articulatory parameters such as height, backness, or rounding are added to a vowel system, because the vowel systems of most languages already employ them. Kalasha, an endangered Dardic (Indo-Aryan) language has five contrastive pairs of plain (/i e a o u/) and rhotic vowels (/i&#734; e&#734; a&#734; o&#734; u&#734;/), as well as nasalized counterparts of all of these (/&#297; &#7869; &#227; &#245; &#361; &#297;&#734; &#7869;&#734; &#227;&#734; &#245;&#734; &#361;&#734;/) (<xref ref-type="bibr" rid="B14">Cooper, 2005</xref>; <xref ref-type="bibr" rid="B32">Hussain &amp; Mielke, 2020</xref>;. <xref ref-type="bibr" rid="B38">Kochetov, Arsenault, Petersen, Kalas, &amp; Kalash, 2021</xref>). The rhotic vowels are thought to have developed recently from the combination of plain vowels and a source of retroflexion (e.g., /&#637; &#635; &#627;/, <xref ref-type="bibr" rid="B27">Heeg&#229;rd &amp; M&#248;rch, 2004</xref>), as schematized in <xref ref-type="fig" rid="F1">Figure 1</xref>.</p>
<fig id="F1">
<label>Figure 1</label>
<caption>
<p>Schematic representation of the development of rhotic vowels in Kalasha.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-9019-g1.png"/>
</fig>
<p>This recent appearance of rhotic vowels in Kalasha provides an opportunity to explore the development of a new vowel feature (i.e., how does an articulatory gesture get combined with all the vowel qualities in a vowel system?) The genesis of rhotic vowels strongly points to retroflexion, given that it was the retroflex class of consonants that triggered the change (<xref ref-type="bibr" rid="B27">Heeg&#229;rd &amp; M&#248;rch, 2004</xref>). Apart from the approximant, retroflex consonants are very likely to have a tip-up or true retroflex configuration. However, Hussain and Mielke (<xref ref-type="bibr" rid="B33">2021</xref>) found that present-day Kalasha rhotic vowels are produced with tongue bunching rather than retroflexion. Clearly the vowel system has undergone reorganization from its articulatory basis.</p>
<p>Height, backness, and rounding are ubiquitous vowel quality features. We are interested in how other features interact with vowel quality when they are introduced into a vowel system, as in the case of rhoticity being added to Kalasha&#8217;s vowel system. While rhotic vowels in Kalasha and other languages have lower F3 than the corresponding plain vowels (<xref ref-type="bibr" rid="B33">Hussain &amp; Mielke, 2021</xref>; <xref ref-type="bibr" rid="B38">Kochetov et al., 2021</xref>, see section 2.3), making a plain vowel rhotic is not simple, because the gestures involved in lowering F3 are complex and interact with the other aspects of vowel quality such as F1 and F2. To combine plain vowel qualities with rhoticity, it is necessary to consider the gestures involved in producing vowels and rhotics as well as the acoustic consequences of combining these gestures. In this paper we use biomechanical modeling to compare rhotic vowels with the result of applying retroflex and bunched tongue gestures to plain vowels. More generally, this is an attempt to isolate the phonetic bases of phonological patterns. While phonetic explanations are widely invoked to account for phonological observations, the arguments often take the form of a typological observation concerning the distribution of a particular pattern and a plausible phonetic explanation for it. It is often difficult to study the phonetic motivation directly, because its phonologized result and/or its conventionalized phonetic precursor may already be present in a language under investigation. Here we attempt to produce a realistic vowel+rhotic coarticulation baseline to compare with Kalasha&#8217;s rhotic vowels.</p>
<p>Previous studies have investigated the phonetics of rhotic vowels and vowel+rhotic sequences in the Dardic languages Kalasha and Dameli, and the Nuristani languages Eastern Kataviri and Kamviri (<xref ref-type="bibr" rid="B34">Hussain &amp; Mielke, 2022</xref>). The pairs of rhotic and non-rhotic vowels differ considerably in their lingual articulation (<xref ref-type="bibr" rid="B33">Hussain &amp; Mielke, 2021</xref>), but it is not known how much of this difference is directly attributable to retroflex coarticulation and how much is due to subsequent changes as the new vowels have been incorporated into the sound system of Kalasha. In this paper we seek to explore the apparent coarticulatory basis for a new vowel feature through biomechanical modeling of the effects of combining plain vowels with a coarticulatory source derived from a different language. We are interested in whether the development of rhotic vowels from plain vowel+retroflex is basically additive from an articulatory standpoint. We use the rhotic approximant /&#635;/ produced by contemporary speakers of Kamviri as representative of the articulation that conditioned the rhotic vowels of Kalasha. Kamviri speakers have retroflex and bunched versions of rhotic approximant /&#635;/, which provides us the opportunity to investigate how coarticulation of vowels and retroflex/bunched approximants results in the development of rhotic vowels in modern Kalasha.</p>
<p>Kalasha rhotic vowels are predominantly bunched and they retain the lip rounding gestures of the corresponding non-rhotic vowels, but rhotic vowels differ considerably from the corresponding plain vowels in tongue posture and acoustic vowel quality (<xref ref-type="bibr" rid="B33">Hussain &amp; Mielke, 2021</xref>, see section 2.3). These differences are difficult to interpret without a realistic model of coarticulation between plain vowels and rhotic consonants. Thus, we use biomechanical and acoustic modeling to address the following questions:</p>
<p>(a) <bold>Do the rhotic vowels retain the lingual articulation of the corresponding non-rhotic vowels?</bold> This will be assessed by combining the corresponding plain vowels with both retroflex and bunched approximants and comparing them with the rhotic vowels and determining whether the differences in tongue height and backness observed within each plain-rhotic pair are accounted for by adding retroflexion or bunching to the plain vowel.</p>
<p>(b) <bold>Do the rhotic vowels retain any signs of retroflexion of the historically present retroflex consonants?</bold> It is already known that Kalasha rhotic vowels are predominantly bunched, but it is unknown how much other aspects of their articulation might be more similar to the retroflex consonants that provided the original coarticulation leading to their development.</p>
<p>(c) <bold>Do the formant frequencies of rhotic vowels differ from what would be expected from adding retroflexion or bunching to their non-rhotic counterparts?</bold> Although the modern Kalasha vowels are bunched, they could have acoustic properties that are attributable to their previous existence as plain vowels coarticulated with retroflexion in particular. Furthermore, since five-way rhotic vowel quality contrasts are unusual, we wonder if additional quality adjustments are required to keep them perceptually distinct.</p>
</sec>
<sec>
<title>2. Background</title>
<sec>
<title>2.1 Typology of vowel distinctions</title>
<p>Retroflex or rhotic vowels are found in fewer than 1% of the world&#8217;s languages (<xref ref-type="bibr" rid="B52">Maddieson, 1984</xref>; <xref ref-type="bibr" rid="B58">Moran, McCloy, &amp; Wright, 2014</xref>). While vowel rhoticity may be considered marginal from a broad crosslinguistic perspective, it is a basic vowel feature in Kalasha. To help contextualize the interaction of rhoticity and vowel quality in Kalasha, we begin by surveying various phonetic vowel distinctions and their interaction with the lingual articulation of vowel quality. <xref ref-type="table" rid="T1">Table 1</xref> shows the proportion of languages (defined as distinct ISO 693-3 codes) in the PHOIBLE database (<xref ref-type="bibr" rid="B58">Moran et al., 2014</xref>) that employ various phonetic distinctions in their vowel systems, as well as vowel quality dimensions that they are particularly likely to interact with.<xref ref-type="fn" rid="n1">1</xref> We did not require these distinctions to be minimal, only to be present (i.e., /i/ vs. /u/ counts as both backness and rounding).</p>
<table-wrap id="T1">
<label>Table 1</label>
<caption>
<p>Occurrence of vowel distinctions in the PHOIBLE database (<xref ref-type="bibr" rid="B58">Moran et al., 2014</xref>).</p>
</caption>
<table>
<thead>
<tr>
<td align="left" valign="top"><bold>Distinction</bold></td>
<td align="left" valign="top"><bold>Example</bold></td>
<td align="left" valign="top"><bold>Count</bold></td>
<td align="left" valign="top"><bold>Percentage</bold></td>
<td align="left" valign="top"><bold>Interacts with</bold></td>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">height</td>
<td align="left" valign="top">i and a</td>
<td align="right" valign="top">2101</td>
<td align="right" valign="top">100.00%</td>
<td align="left" valign="top">[see below]</td>
</tr>
<tr>
<td align="left" valign="top">backness</td>
<td align="left" valign="top">i and u</td>
<td align="right" valign="top">2095</td>
<td align="right" valign="top">99.71%</td>
<td align="left" valign="top">[see below]</td>
</tr>
<tr>
<td align="left" valign="top">lip rounding</td>
<td align="left" valign="top">i and u</td>
<td align="right" valign="top">2088</td>
<td align="right" valign="top">99.38%</td>
<td align="left" valign="top">backness</td>
</tr>
<tr>
<td align="left" valign="top">length</td>
<td align="left" valign="top">a&#720;</td>
<td align="right" valign="top">883</td>
<td align="right" valign="top">42.03%</td>
<td align="left" valign="top">height</td>
</tr>
<tr>
<td align="left" valign="top">nasalization</td>
<td align="left" valign="top">&#227;</td>
<td align="right" valign="top">500</td>
<td align="right" valign="top">23.80%</td>
<td align="left" valign="top">height</td>
</tr>
<tr>
<td align="left" valign="top">creakiness (or glottalization, ejective)</td>
<td align="left" valign="top">a&#816; a&#8217; or a<sup>&#660;</sup></td>
<td align="right" valign="top">30</td>
<td align="right" valign="top">1.43%</td>
<td align="left" valign="top">height</td>
</tr>
<tr>
<td align="left" valign="top">breathiness</td>
<td align="left" valign="top">a&#804;</td>
<td align="right" valign="top">28</td>
<td align="right" valign="top">1.33%</td>
<td align="left" valign="top">height</td>
</tr>
<tr>
<td align="left" valign="top">pharyngealization</td>
<td align="left" valign="top">a<sup>&#661;</sup></td>
<td align="right" valign="top">10</td>
<td align="right" valign="top">0.48%</td>
<td align="left" valign="top">height &amp; backness</td>
</tr>
<tr>
<td align="left" valign="top">rhoticity</td>
<td align="left" valign="top">a&#734;</td>
<td align="right" valign="top">9</td>
<td align="right" valign="top">0.43%</td>
<td align="left" valign="top">height &amp; backness</td>
</tr>
<tr>
<td align="left" valign="top">voicing</td>
<td align="left" valign="top">a&#805;</td>
<td align="right" valign="top">7</td>
<td align="right" valign="top">0.33%</td>
<td align="left" valign="top">height</td>
</tr>
<tr>
<td align="left" valign="top">tongue root advancement or retraction</td>
<td align="left" valign="top">a&#799; or a&#800;</td>
<td align="right" valign="top">6</td>
<td align="right" valign="top">0.29%</td>
<td align="left" valign="top">height</td>
</tr>
<tr>
<td align="left" valign="top">velarization</td>
<td align="left" valign="top">a<sup>&#611;</sup></td>
<td align="right" valign="top">2</td>
<td align="right" valign="top">0.10%</td>
<td align="left" valign="top">height &amp; backness</td>
</tr>
<tr>
<td align="left" valign="top">epilaryngeal source</td>
<td align="left" valign="top">a<sup>E</sup></td>
<td align="right" valign="top">1</td>
<td align="right" valign="top">0.05%</td>
<td align="left" valign="top">backness</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>While lip rounding is utilized in nearly all vowel systems, it is worth considering how it interacts with tongue posture. Although the lips and tongue are articulatorily independent, they may interact through the trading relations involving their effects on formant frequencies. Vowel pairs such as /i y/ that ostensibly differ only in lip rounding often differ in tongue position as well (<xref ref-type="bibr" rid="B35">Jackson &amp; McGowan, 2012</xref>). Length and nasalization are the most frequent additional features that are at least partially independent of vowel quality, followed by various forms of laryngealization (creakiness/glottalization/ejective), breathiness, pharyngealization, and rhoticity. Each of these semi-independent vowel features has opportunities to interact with vowel quality.</p>
<p>A prototypical example of length-quality interaction is the development of Persian vowels: Classical Persian has been analyzed as having a three-way quality contrast combined with a two-way length contrast /a a&#720; i i&#720; u u&#720;/ (<xref ref-type="bibr" rid="B40">Kr&#225;msk&#7923;, 1939</xref>), but in Modern Persian this is a six-way quality contrast /a &#593; e i o u/ (<xref ref-type="bibr" rid="B62">Nye, 1955</xref>) or /a &#593;&#720; e i&#720; o u&#720;/ (<xref ref-type="bibr" rid="B77">Toosarvandani, 2004</xref>). In this case, the Classical Persian long vowels have developed higher and/or more peripheral qualities in Modern Persian than the corresponding short vowels. This is consistent with the idea that short vowels are more vulnerable to undershoot (<xref ref-type="bibr" rid="B45">Lindblom, 1963</xref>), which can be reinterpreted as a basic property of a vowel. The quality differences among the non-low vowels are also consistent with the idea that higher vowels sound longer than lower vowels and speakers may use lowering to signal short duration and raising to signal long duration (<xref ref-type="bibr" rid="B25">Gussenhoven, 2007</xref>).</p>
<p>Vowel nasalization is achieved by velum lowering, which is also independent of tongue position but has overlapping acoustic consequences that may lead to lingual differences among ostensibly similar oral-nasal pairs. Nasalization leads to F1 raising in high vowels and F1 lowering in low vowels due to an additional vocal tract resonance in the vicinity of an F1 frequency that is typical for a low-mid vowel (<xref ref-type="bibr" rid="B18">Diehl, Kluender, Walsh, &amp; Parker, 1991</xref>; <xref ref-type="bibr" rid="B21">Feng &amp; Castelli, 1996</xref>; <xref ref-type="bibr" rid="B23">Fujimura &amp; Lindqvist, 1971</xref>; <xref ref-type="bibr" rid="B68">Serrurier &amp; Badin, 2008</xref>). This acoustic effect of nasalization may lead to enhancement or compensation (<xref ref-type="bibr" rid="B4">Beddor, 1982</xref>; <xref ref-type="bibr" rid="B39">Krakow, Beddor, Goldstein, &amp; Fowler, 1988</xref>). The acoustic effects of nasalization are enhanced with a more neutral tongue height in Northern Metropolitan French (<xref ref-type="bibr" rid="B10">Carignan, 2014</xref>) and Brazilian Portuguese (<xref ref-type="bibr" rid="B3">Barlaz et al., 2015</xref>), and also in Kalasha (<xref ref-type="bibr" rid="B33">Hussain &amp; Mielke, 2021</xref>).</p>
<p>Creaky voice quality and glottalization are associated with larynx raising, which shortens the vocal tract and raises F1 and other formants (<xref ref-type="bibr" rid="B41">Laver, 1980</xref>), making vowels sound lower, whereas breathy vowels are often produced with a lowered larynx, which has the opposite effect (<xref ref-type="bibr" rid="B20">Esposito, Sleeper, &amp; Sch&#228;fer, 2021</xref>), making vowels sound higher. Indeed, listeners perceive creaky vowels as sounding lower (<xref ref-type="bibr" rid="B9">Brunner &amp; Zygis, 2011</xref>) and breathy vowels as sounding higher (<xref ref-type="bibr" rid="B50">Lotto, Holt, &amp; Kluender, 1997</xref>). It is reasonable to expect phonation differences to be enhanced by vowel height, much like length and nasalization differences (see <xref ref-type="bibr" rid="B20">Esposito et al., 2021</xref> for more discussion of the interaction of voice quality and vowel quality).</p>
<p>Pharyngealized vowels are typically produced with centralization of F1 and F2, and they may or may not show signs of rhoticity such as low F3 (<xref ref-type="bibr" rid="B13">Catford, 1983</xref>). Pharyngealization makes the back cavity smaller and limits the tongue&#8217;s freedom of movement to produce extreme vowel postures. Hussain and Mielke (<xref ref-type="bibr" rid="B33">2021</xref>) concluded that there is probably considerable overlap between vowels described as pharyngealized and vowels described as rhotic, but that these terms are not equivalent.</p>
<p>Most of the vowel types described so far have acoustic or perceptual effects that may lead to changes in vowel quality. By contrast, rhotic vowels directly interact with the tongue movements used to produce vowels, and accordingly rhoticity is expected to be particularly aggressive at rearranging vowel systems it is introduced to. Rhotic vowels such as /&#602;/ and rhotic approximants such as retroflex (tip-up) [&#635;] and bunched (tip-down) [&#633;] are produced with a wide range of tongue shapes often grouped into retroflex and bunched categories (<xref ref-type="bibr" rid="B16">Delattre &amp; Freeman, 1968</xref>; <xref ref-type="bibr" rid="B33">Hussain &amp; Mielke, 2021</xref>; <xref ref-type="bibr" rid="B55">Mielke, 2015</xref>; <xref ref-type="bibr" rid="B84">Zhou et al., 2008</xref>). Retroflexion is achieved by raising and retracting the tongue tip (possibly involving subapical or sublaminal contact with the alveolar ridge). Bunching can be achieved by lowering and retracting the tongue tip while depressing the medial portion of the tongue dorsum, resulting in a concavity that is prominent in the mid-sagittal plane (<xref ref-type="bibr" rid="B33">Hussain &amp; Mielke, 2021</xref>; <xref ref-type="bibr" rid="B57">Moisik, 2013</xref>; <xref ref-type="bibr" rid="B70">Stavness, Gick, Derrick, &amp; Fels, 2012</xref>). Tongue dorsum concavity is a key characteristic of bunched [&#633;] and has also been observed in retroflex tongue shapes in English and Canadian French (<xref ref-type="bibr" rid="B55">Mielke, 2015</xref>; <xref ref-type="bibr" rid="B84">Zhou et al., 2008</xref>). Retroflexion, in addition to raising of the tongue tip towards hard palate, may also include a dip in between the tongue tip and dorsum (<xref ref-type="bibr" rid="B26">Hamann, 2003</xref>).</p>
<p>Low F3 is a hallmark of bunching and retroflexion (<xref ref-type="bibr" rid="B42">Lehiste, 1962</xref>). Tongue tip raising and tongue bunching both give rise to a large sublingual cavity, which results in lowering of F3 (<xref ref-type="bibr" rid="B84">Zhou et al., 2008</xref>). Delattre and Freeman (<xref ref-type="bibr" rid="B16">1968</xref>) reported a wide range of articulatory to acoustic correlations for the American English [&#633;]. (1) A narrow palato-velar constriction lowers the F3 or brings F2 and F3 close to each other. (2) A dip in the tongue dorsum lowers F3. (3) A wider pharyngeal constriction increases the distance between F2 and F3 but a narrow pharyngeal constriction brings the two formants closer. (4) Lip rounding lowers all the formants.</p>
<p><xref ref-type="table" rid="T2">Table 2</xref> organizes the vowel distinctions that are at least as frequent as rhoticity in descending order of their expected effect on the lingual articulation of vowels. Rhoticity directly affects most aspects of tongue posture, and pharyngeal constriction and tongue root movements affect the posterior tongue. The others have effects that are mediated primarily by acoustics, by affecting the end of the vocal tract as in lip rounding and laryngeal gestures, or the effect of velopharyngeal coupling on F1, or primarily by perception, as in the case of length.</p>
<table-wrap id="T2">
<label>Table 2</label>
<caption>
<p>Vowel distinctions by hypothesized hierarchy of impact on lingual articulation of vowels.</p>
</caption>
<table>
<thead>
<tr>
<td align="left" valign="top"><bold>Distinction</bold></td>
<td align="left" valign="top"><bold>Impact</bold></td>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">rhoticity</td>
<td align="left" valign="top">directly affects tongue posture</td>
</tr>
<tr>
<td align="left" valign="top">pharyngealization, tongue root</td>
<td align="left" valign="top">directly affects pharynx volume</td>
</tr>
<tr>
<td align="left" valign="top">lip rounding</td>
<td align="left" valign="top">affects front cavity</td>
</tr>
<tr>
<td align="left" valign="top">creakiness, breathinesss</td>
<td align="left" valign="top">affects pharynx length</td>
</tr>
<tr>
<td align="left" valign="top">nasalization</td>
<td align="left" valign="top">affects acoustics of F1</td>
</tr>
<tr>
<td align="left" valign="top">length</td>
<td align="left" valign="top">affects undershoot and perception of height</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>2.2. Development of Kalasha rhotic vowels</title>
<p>The question of how rhotic vowels emerged in Kalasha is a major issue in the current Dardic literature (<xref ref-type="bibr" rid="B14">Cooper, 2005</xref>; <xref ref-type="bibr" rid="B17">Di Carlo, 2016</xref>; <xref ref-type="bibr" rid="B27">Heeg&#229;rd &amp; M&#248;rch, 2004</xref>; <xref ref-type="bibr" rid="B34">Hussain &amp; Mielke, 2022</xref>; <xref ref-type="bibr" rid="B38">Kochetov et al., 2021</xref>). There are reasons to attribute the development of rhotic vowels to internal sound change as well as to external factors. The internal source of rhoticity is the occurrence of retroflex consonants in the environment of a vowel that developed rhoticity. The oral plain vowels of Northern (Birir, Bumburet, and Rumbur) and Southern (Urtsun and Jinjiret) dialects of Kalasha underwent rhoticization due to the presence of liquids and retroflex consonants in a word. For instance, the rhotic-nasal vowels of modern Kalasha can be reconstructed from the Old Indo-Aryan (Sanskrit) retroflex nasal /&#627;/ (Sanskrit /pa&#627;i/ &#8216;hand&#8217; <italic>&#8594;</italic> Kalasha /p&#7869;&#734;/ &#8216;palm of the hand&#8217;: <xref ref-type="bibr" rid="B27">Heeg&#229;rd &amp; M&#248;rch, 2004</xref>).</p>
<p>Some evidence about earlier stages of Kalasha comes from fieldwork by Leitner in 1866&#8211;72 (<xref ref-type="bibr" rid="B44">Leitner, 1880</xref>) and Morgenstierne in the late 1920s (<xref ref-type="bibr" rid="B60">Morgenstierne, 1973</xref>). In Leitner and Morgenstierne&#8217;s wordlists, Kalasha rhotic vowels were transcribed as sequences of plain vowels+palatal fricative /&#345;/ or rhotics /r &#7771;/ (an underdot is generally used to denote retroflexion in Indo-Aryan and Nuristani literature). For example, the word for &#8216;heart&#8217; was transcribed as /h&#233;ra/ by Leitner and /h&#239;&#720;&#345;a/ by Morgenstierne; in modern Bumburet Kalasha, the same word is produced as /hi&#734;a/, with a rhotic vowel. Morgenstierne (<xref ref-type="bibr" rid="B60">1973</xref>) originally used the term <italic>palatal fricative</italic> to refer to a class of r-colored speech sounds in Dardic and Nuristani languages, which may affect the quality of the neighboring vowels. This resembles modern descriptions of Dameli and Nuristani languages (Eastern Kataviri and Kamviri), which have non-phonemic and r-colored vowels in the vicinity of /&#635;/ (<xref ref-type="bibr" rid="B65">Perder, 2013</xref>; <xref ref-type="bibr" rid="B73">Strand, 2011</xref>).</p>
<p>We infer from these descriptions that the development of rhotic vowels in Kalasha is the result of vowels coalescing with following liquids and retroflex consonants and that researchers 95&#8211;160 years ago heard the consonantal portions as palatal fricative /&#345;/ or rhotics /r &#7771;/. During the time period that Kalasha has been developing rhoticity in vowels that are followed by retroflex consonants, it has been in contact with Nuristani languages that have abundant phonetic vowel rhoticity. This vowel rhoticity takes the form of a contrastive retroflex approximant /&#635;/ and non-phonemic retroflex vowels (<xref ref-type="bibr" rid="B27">Heeg&#229;rd &amp; M&#248;rch, 2004</xref>; <xref ref-type="bibr" rid="B59">Morgenstierne, 1954</xref>), as well as r-coloring of vowels in the vicinity of the retroflex flap /&#637;/ (<xref ref-type="bibr" rid="B73">Strand, 2011</xref>).</p>
<p>In summary, Kalasha rhotic vowels are believed to have emerged via loss of retroflex consonants (e.g., /&#637; &#635; &#627;/). The development of phonemic rhotic vowels may have also been encouraged by intensive contact with the Nuristani languages with a contrastive retroflex approximant /&#635;/ (<xref ref-type="bibr" rid="B17">Di Carlo, 2016</xref>; <xref ref-type="bibr" rid="B27">Heeg&#229;rd &amp; M&#248;rch, 2004</xref>; <xref ref-type="bibr" rid="B33">Hussain &amp; Mielke, 2021</xref>, <xref ref-type="bibr" rid="B34">2022</xref>).</p>
</sec>
<sec>
<title>2.3 Rhotic vowels in modern Kalasha</title>
<p>A handful of studies have investigated the phonetic correlates of rhotic vowels in Kalasha (<xref ref-type="bibr" rid="B33">Hussain &amp; Mielke, 2021</xref>, <xref ref-type="bibr" rid="B34">2022</xref>; <xref ref-type="bibr" rid="B38">Kochetov et al., 2021</xref>). <xref ref-type="fig" rid="F2">Figure 2</xref> shows mean formant frequencies for the ten plain and rhotic oral vowels of four male speakers reported by Hussain and Mielke (<xref ref-type="bibr" rid="B33">2021</xref>) for the tokens included in this study. It can be observed that the rhotic vowels are generally more centralized in F2 relative to their non-rhotic counterparts, their F1 is centralized or raised, and they have much lower F3 (as also shown by <xref ref-type="bibr" rid="B38">Kochetov et al., 2021</xref>).<xref ref-type="fn" rid="n2">2</xref> The rhotic vowels are close to steady-state (i.e., they are not generally more rhotic at the end than at the beginning; <xref ref-type="bibr" rid="B33">Hussain &amp; Mielke, 2021</xref>).</p>
<fig id="F2">
<label>Figure 2</label>
<caption>
<p>F1, F2, and F3 of observed plain and rhotic oral vowels of Kalasha (<xref ref-type="bibr" rid="B33">Hussain &amp; Mielke, 2021</xref>).</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-9019-g2.png"/>
</fig>
<p><xref ref-type="fig" rid="F3">Figure 3</xref> shows Smoothing-Spline ANOVA (SSANOVA) comparisons of tongue shapes used to produce these ten vowels. The five non-rhotic vowels are articulated as expected (e.g., on the basis of Wood (<xref ref-type="bibr" rid="B82">1979</xref>): /i/ and /e/ have constrictions made toward the hard palate, /u/ has tongue raising toward the velum, and /o/ and /a/ have constrictions in the upper and lower pharynx, respectively). The five rhotic vowels are produced with tongue bunching (rather than retroflexion) by all speakers investigated by Hussain and Mielke (<xref ref-type="bibr" rid="B33">2021</xref>). In addition, all five rhotic vowels, including /a&#734; o&#734; u&#734;/, are produced with relatively front tongue body position.</p>
<fig id="F3">
<label>Figure 3</label>
<caption>
<p>SSANOVA figures showing tongue shapes used to produce plain (top) and rhotic (bottom) vowels by one Kalasha speaker (Kal1). Tongue tip is to the right. X and Y axes are in centimeters (<xref ref-type="bibr" rid="B33">Hussain &amp; Mielke, 2021</xref>).</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-9019-g3.png"/>
</fig>
<p>While the rhotic vowel tongue shapes are quite different from the tongue shapes used to produce the corresponding plain vowels, Hussain and Mielke (<xref ref-type="bibr" rid="B33">2021</xref>) found no significant differences between the lip postures for rhotic vowels and the corresponding non-rhotic vowels. /u/, /u&#734;/, /o/, and /o&#734;/ and their nasal counterparts are all rounded and all the other vowels are not. /u/ and /u&#734;/ are produced with a slightly smaller lip opening than /o/ and /o&#734;/.</p>
</sec>
<sec>
<title>2.4. Tongue muscle activations during vowel and rhotic production</title>
<p>In this paper we model the coarticulation between non-rhotic vowels and rhotic consonants. The tongue shapes involved in coarticulated vowel+rhotic sequences are determined by the combination of muscle activations involved in vowels and rhotic consonants. We begin by reviewing the muscle activations involved in producing vowels and rhotic approximants. The human tongue is controlled by intrinsic and extrinsic sets of muscles (see <xref ref-type="bibr" rid="B66">Sanders &amp; Mu, 2013</xref>; chapters 8&#8211;9 of <xref ref-type="bibr" rid="B24">Gick, Wilson, &amp; Derrick, 2013</xref>; and <xref ref-type="fig" rid="F1">Figures 1</xref>, <xref ref-type="fig" rid="F2">2</xref>, <xref ref-type="fig" rid="F3">3</xref>, <xref ref-type="fig" rid="F4">4</xref> in <xref ref-type="bibr" rid="B36">Jang, 2022</xref> and references therein for more details). Minor contractions in intrinsic muscles change the shape of the tongue from the inside. The intrinsic tongue muscles consist of inferior longitudinal, which lowers and retracts the tongue tip; superior longitudinal, which raises and retracts the tongue tip; verticalis, which flattens the tongue body; and transversus, which narrows the tongue laterally and causes a sagittal expansion, enlarging the tongue along both anteroposterior and inferosuperior axes due to the muscular hydrostatic nature of the tongue (<xref ref-type="bibr" rid="B69">Smith &amp; Kier, 1989</xref>).</p>
<fig id="F4">
<label>Figure 4</label>
<caption>
<p>Registration of one participant&#8217;s tongue contours (gray lines) against the ArtiSynth tongue contour in black line (left) and a selection of these tongue contours (representing the all types of plain and rhotic vowels) within the ArtiSynth vocal tract (right). The magenta circles are the inverse target locations that define the inverse simulation trajectory.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-9019-g4.jpg"/>
</fig>
<p>The extrinsic tongue muscles connect the tongue to other parts of the body (see <xref ref-type="bibr" rid="B28">Honda, 1996</xref>; <xref ref-type="bibr" rid="B75">Takano &amp; Honda, 2007</xref>). The genioglossus courses from the superior mental spines inside the mandible to the full length of the tongue. Contracting the anterior fibers of the genioglossus lowers the front of the tongue and contracting middle fibers lowers the tongue dorsum. Contracting posterior fibers of the genioglossus moves the whole tongue forward and raises it toward the palate, as in the production of high and front vowels. The hyoglossus courses from the hyoid bone at the root of the tongue up to the sides of the tongue and pulls the tongue down and back when contracted (as in low and back vowels). The palatoglossus courses from the soft palate to the sides of the tongue and pulls the soft palate down or the tongue up depending on the state of other muscles attached to these structures. The styloglossus courses from the styloid processes in the skull below the ears forward to the sides of the tongue, and has the potential to retract and stabilize the tongue in back vowels. In addition to these intrinsic and extrinsic tongue muscles, the geniohyoid and mylohyoid are muscles that do not directly insert into the tongue but form the floor of the mouth and accordingly support the base of the tongue and facilitate tongue raising.</p>
<p>Stavness, Gick, et al. (<xref ref-type="bibr" rid="B70">2012</xref>) modeled the articulation of retroflex and bunched English /r/ variants using the ArtiSynth biomechanical modeling toolkit (<ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="www.artisynth.org">www.artisynth.org</ext-link>; <xref ref-type="bibr" rid="B49">Lloyd, Stavness, &amp; Fels, 2012</xref>). Their bunched /r/ involved the contraction of the superior longitudinal, inferior longitudinal, and anterior genioglossus (to retract the tongue tip without raising it), middle genioglossus (to lower the tongue dorsum medially), along with some retraction of the transversus and verticalis to bunch the tongue body. Their tip-up (retroflex) /r/ involved greater contraction of the superior longitudinal without contraction of the inferior longitudinal (to retract and raise the tongue tip), plus some contraction of the middle genioglossus (but less than in the bunched variant), and optionally contraction of the hyoglossus to do some of the work of retracting the tongue without depressing the tip. Stavness, Gick, et al. (<xref ref-type="bibr" rid="B70">2012</xref>) showed that their bunched /r/ is very compatible with their simulation of [i], which differs from it by principally having contraction of the posterior genioglossus to front and raise the tongue and less contraction of the middle genioglossus. Their retroflex /r/ is compatible with their simulation of [a], which similarly involves middle genioglossus and hyoglossus to lower the tongue body, and differs from retroflex /r/ in having contraction of the verticalis to further depress the tongue body and lacking contraction of the superior longitudinal (for no retroflexion). These results helped account for the observation that English /r/ is typically bunched in the context of /i/ and that retroflex /r/ is particularly frequent in the environment of /a/ (<xref ref-type="bibr" rid="B56">Mielke, Baker, &amp; Archangeli, 2016</xref>; <xref ref-type="bibr" rid="B63">Ong &amp; Stone, 1998</xref>).</p>
<p>The aim of the current study is to investigate the similarity between Kalasha rhotic vowels and the superposition of plain vowels and rhotic approximants (i.e., additive combination of the articulatory gestures of plain vowels and retroflex /&#635;/ and bunched /&#633;/ approximants). We use the phonetic data presented in Hussain and Mielke (<xref ref-type="bibr" rid="B33">2021</xref>, <xref ref-type="bibr" rid="B34">2022</xref>) as the basis for the current investigation of the biomechanics involved in the production of Kalasha rhotic vowels and combine Kalasha plain vowels with Kamviri-style rhotic approximants /&#635; &#633;/. Moreover, we also examine the role of different tongue muscles in the production of plain and rhotic vowels of Kalasha and compare them with the retroflex and bunched approximants of Kamviri.</p>
</sec>
</sec>
<sec>
<title>3. Methods</title>
<sec>
<title>3.1. Languages and speakers</title>
<p>Acoustic and articulatory data used as the basis for simulations were described in more detail in Hussain and Mielke (<xref ref-type="bibr" rid="B33">2021</xref>, <xref ref-type="bibr" rid="B34">2022</xref>). Recordings were made of four Kalasha (Bumburet dialect) and two Kamviri speakers (all males in their 20s or 30s). All six speakers were from Bumburet valley, Chitral, northern Pakistan. In addition to their native languages, all the speakers could speak Khowar, Pashto, Urdu, and/or English.</p>
</sec>
<sec>
<title>3.2. Speech materials</title>
<p>The modeling described in this paper is based on the rhotic vowels in the Kalasha words /pi&#734;&#720;/ &#8216;press&#8217;, /he&#734;/ &#8216;theft&#8217;, /ba&#734;/ &#8216;lazy&#8217;, /t&#643;o&#734;i/ &#8216;parasite&#8217;, and /k<sup>h</sup>u&#734;/ &#8216;hat&#8217;, the non-rhotic vowels in the Kalasha words /pi/ &#8216;drink (verb); from&#8217;, /pe/ &#8216;if&#8217;, /pa&#720;/ &#8216;go&#8217;, /po/ &#8216;footprint&#8217;, and /tu/ &#8216;you&#8217;, and the word-final rhotic approximant /&#635;/ (retroflex) or /&#633;/ (bunched) in the Kamviri word /parma&#635;/ &#8216;child.&#8217;</p>
</sec>
<sec>
<title>3.3. Recording procedure</title>
<p>The participants were invited into a quiet room at a hotel in Bumburet valley, Chitral, Pakistan. A Terason t3000 ultrasound machine with Ultraspeech 1.3 software (<xref ref-type="bibr" rid="B30">Hueber, Chollet, Denby, &amp; Stone, 2008</xref>) was used for recording the ultrasound data. The tongue ultrasound and lip video recordings were made in direct-to-disk mode, generating 640 <italic>&#215;</italic> 480 pixel bitmap images at 60 frames per second. A Terason 8MC3 3&#8211;8 MHz ultrasound transducer was positioned underneath each participant&#8217;s chin, stabilized with an Articulate Instruments aluminum Probe Stabilisation Headset (<xref ref-type="bibr" rid="B67">Scobbie, Wrench, &amp; van der Linden, 2008</xref>). A frontal view of the lips was captured using a board camera (The Imaging Source DFM 22BUC03-ML with a 12 <italic>&#215;</italic> 0.5 mm lens) mounted on the headset about five centimeters in front of each participant&#8217;s lips using two clip-on LED book lights, which also illuminated the participants&#8217; lips.</p>
<p>Simultaneous audio recordings were made with a Shure Beta 53 head-mounted omnidirectional condenser microphone (44.1 kHz, 16-bit). Before the recordings, all the participants were familiarized with the task and went through the wordlists. After audio and ultrasound recording commenced, the participants were asked to hold a mouthful of water to generate ultrasound images of the palate (not used in this analysis) and then held a tongue depressor between their teeth and pressed their tongue against it in order to generate ultrasound images of the occlusal plane. The target wordlists were presented to the participants on a computer screen or they were described to them in Urdu, which is a lingua franca of Pakistan. All the words were elicited in citation form. Each target word was repeated five times.</p>
</sec>
<sec>
<title>3.4. Acoustic and articulatory analyses</title>
<p>The ultrasound frames were selected from the midpoints of the Kalasha vowel intervals and the Kamviri rhotic approximant intervals. The tongue contours and the lip opening were manually traced in Palatoglossatron (<xref ref-type="bibr" rid="B2">Baker, 2005</xref>). The lip data used to inform the modeling were two points placed mid-sagittally on the edge of the upper and lower lips. The frequencies of the first three formants were extracted at the same time points using Praat (<xref ref-type="bibr" rid="B8">Boersma &amp; Weenink, 2007</xref>).</p>
</sec>
<sec>
<title>3.5. ArtiSynth modeling</title>
<p>Our goal was to test whether Kalasha rhotic vowels are similar articulatorily and acoustically to the superposition of plain vowels and rhotic approximant articulations. Thus, we have taken plain vowels coarticulated with a following rhotic approximant to be the initial state for Kalasha rhotic vowels. To model this coarticulation, we used the ArtiSynth biomechanical modeling toolkit (<ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://www.artisynth.org">www.artisynth.org</ext-link>; <xref ref-type="bibr" rid="B49">Lloyd et al., 2012</xref>). ArtiSynth is a free, open-source computational platform for simulating multibody systems comprising rigid and deformable bodies (the latter implemented as finite-element models or FEMs). These bodies can be made to interact via unilateral (e.g., contact/collision) and bilateral (e.g., joint) constraints and manipulated with force effectors of various kinds, such as musculature based on realistic mathematical muscle models. We performed two types of simulations: (i) Inverse simulations of the Kalasha vowel shapes, both plain and rhotic, and Kamviri rhotic approximants (both retroflex /&#635;/ and bunched /&#633;/ variants); and (ii) forward simulations, which combine the excitations computed during the inverse simulations of the Kalasha plain vowels with excitations of either of the two types of Kamviri rhotic approximant. Along with the biomechanical simulation, we also simulated 1-dimensional acoustics of these vowels, employing frequency-domain acoustic simulation (<xref ref-type="bibr" rid="B6">Birkholz, 2005</xref>; <xref ref-type="bibr" rid="B7">Birkholz &amp; Jackel, 2004</xref>) and making use of the airway skin mesh (<xref ref-type="bibr" rid="B1">Anderson et al., 2017</xref>) to allow for estimation of the vocal tract area function. More details about the biomechanical modeling are included in the Supplementary Materials.</p>
<p>Biomechanical modeling approaches are generally characterized as either forward or inverse. Forward modeling (aka forward-dynamics simulation), is a process of controlling a biomechanical model with given muscle activation signals. Inverse modeling (aka inverse-dynamics simulation), in contrast, is a process which estimates the underlying muscle activations from previously obtained kinematic measurements by using a biomechanical model (<xref ref-type="bibr" rid="B19">Eskes et al., 2017</xref>). The ArtiSynth forward models (retroflex and bunched versions) represent the predicted vocal tract shapes due to superposing a bunched or retroflex approximant on plain vowels. The ArtiSynth inverse model of Kalasha vowels is meant to be a close approximation to the actual observed Kalasha vocal tract shapes. When the inverse simulation tongue shapes are quite different from the forward simulation tongue shapes, it suggests that Kalasha speakers are doing something different (such as a more optimal articulatory strategy to achieve a particular perceptual target). If the acoustic output is also different, this suggests that Kalasha has phonologized new perceptual targets for rhotic vowels (i.e., the goal for the phonological category does not even sound like the vowel+rhotic superposition that it is thought to have originated from).</p>
<p>We next describe the inverse simulations, which form the basis of the forward simulations. Inverse models in ArtiSynth take time-varying target points &#8212; or trajectories &#8212; as input, and they output a set of muscle activations that minimize trajectory-tracking errors while being subject to additional terms for resolving muscle redundancy (<xref ref-type="bibr" rid="B1">Anderson et al., 2017</xref>; <xref ref-type="bibr" rid="B71">Stavness, Lloyd, &amp; Fels, 2012</xref>) and ensuring smooth activations (as rapid changes can lead to model instability). In our simulations, we used the empirical data from the lips and the tongue to define the trajectories. The final set of inverse target points were two targets for the face FEM (on the midline of the upper and lower lips) and eleven target points along the contour of the tongue, running from the tongue tip to the tongue root. We also developed target points for the mandible (one on the central incisors and one on the pogonion) and hyoid bone (one point on the anterosuperior most point of the body) giving optional parameters to use in cases where model stability was an issue. In all cases, the inverse trajectories started at 0.0 s from their initial configuration and ended at their target location at 0.2 s.</p>
<p>We used MATLAB (<xref ref-type="bibr" rid="B53">MATLAB, 2019</xref>) to register the participants&#8217; data (lip points and tongue contours) against corresponding details of the ArtiSynth model (<xref ref-type="fig" rid="F4">Figure 4</xref>). While it was straightforward to identify points on the lips that match the flesh points tracked on the participants&#8217; lips, it was less clear how to map the ultrasound tongue contour data into the ArtiSynth model. This is because we have no information in the ultrasound image as to what part of the tongue exactly is being imaged, and the imaged portion of the tongue also changes from frame to frame. Thus, homology cannot be guaranteed in the registration, and it was therefore necessary to make assumptions about what the typical visible portion of the tongue was and select reference points on the ArtiSynth tongue model (from tongue tip to root) that matched this. With this in mind, we extracted the location of the lower lip, upper lip, mouth corners, and midsagittal contour of the tongue from tip to root from the ArtiSynth model in its neutral configuration and imported these landmarks into MATLAB for further processing alongside the empirical articulatory data. For each participant, the lip and tongue data were independently registered (using Procrustes superimposition) against the ArtiSynth landmarks. In the case of the tongue, we registered the (participant-wise) mean observed tongue contour against the ArtiSynth tongue contour, using resampling to ensure all contours had the same number of sample points. The registered empirical data were then brought into ArtiSynth and visually examined for how well they fit within the ArtiSynth vocal tract, with the possibility of making slight adjustments to the overall scaling and translation of the registered data once it was in the ArtiSynth environment. The entire process was iterated upon using (most notably) slightly different selections for tongue tip and root positions on the ArtiSynth model to try to best fit the tongue contours while also being articulatorily reasonable. The task proved difficult. Ideally, if more participant data were available for other vocal tract structures, the fit could be improved; it would even be possible to register the ArtiSynth model itself to the participant (if, for example, structural MRI data were available for the participant). In practice, the inverse simulation does not always manage to match the target points exactly and thus no severe acoustic issues arose from the articulation occasionally treading slightly outside of the airway skin.<xref ref-type="fn" rid="n3">3</xref> The appearance of the tongue contour data within the ArtiSynth vocal tract setting is illustrated in <xref ref-type="fig" rid="F4">Figure 4</xref>.</p>
<p>While alternative reference points on the ArtiSynth tongue model could have been explored systematically, small differences in reference points would not result in large differences with the current findings (see, e.g., <xref ref-type="bibr" rid="B29">Howson, Moisik, &amp; &#379;ygis, 2022</xref>). This is in part because there is always some amount of error that the inverse simulation makes in hitting the targets (particularly when there are many of them for a given contour and many contours are being used across a large set of simulations). Systematic exploration of this choice is infeasible, especially given the many other assumptions we have made alongside this that could also arguably be deserving of similar attention (such as the number of inverse targets to employ). The best we can do then is to present our findings in the light of the model design choices that we have made with the knowledge that small deviations (such as shifting the reference nodes by one node forward or backward or by using more or less inverse target nodes) would lead to similarly small changes to the results.</p>
<p>For the inverse simulation, we simulated selected productions from four Kalasha participants. We ran simulations covering five basic vowel qualities /i e a o u/ within two vowel types (plain and rhotic), each with approximately five tokens, as shown in <xref ref-type="table" rid="T3">Table 3</xref>, giving us 208 simulations. We also simulated five tokens each of the Kamviri rhotic approximant variants (retroflex and bunched), using data from two different participants who produced the sound differently. The process of bringing these data into ArtiSynth followed the same procedure as that used for the Kalasha simulations outlined above. <xref ref-type="fig" rid="F5">Figure 5</xref> illustrates inverse simulations of /o/ and /o&#734;/.</p>
<table-wrap id="T3">
<label>Table 3</label>
<caption>
<p>Number of simulated tokens by participant and vowel. The number of failed simulation runs included in the count are indicated with a corresponding number of asterisks.</p>
</caption>
<table>
<thead>
<tr>
<td align="left" valign="top"><bold>Participant</bold></td>
<td align="left" valign="top"><bold>i</bold></td>
<td align="left" valign="top"><bold>e</bold></td>
<td align="left" valign="top"><bold>a</bold></td>
<td align="left" valign="top"><bold>o</bold></td>
<td align="left" valign="top"><bold>u</bold></td>
<td align="left" valign="top"><bold>i&#734;</bold></td>
<td align="left" valign="top"><bold>e&#734;</bold></td>
<td align="left" valign="top"><bold>a&#734;</bold></td>
<td align="left" valign="top"><bold>o&#734;</bold></td>
<td align="left" valign="top"><bold>u&#734;</bold></td>
<td align="left" valign="top"><bold>Total</bold></td>
<td align="left" valign="top"></td>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Kal1</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">4</td>
<td align="left" valign="top">4</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">48</td>
<td align="left" valign="top"></td>
</tr>
<tr>
<td align="left" valign="top">Kal4</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">6</td>
<td align="left" valign="top">5*</td>
<td align="left" valign="top">6</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">6*</td>
<td align="left" valign="top">6**</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">54</td>
<td align="left" valign="top">****</td>
</tr>
<tr>
<td align="left" valign="top">Kal5</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">3</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">6</td>
<td align="left" valign="top">6</td>
<td align="left" valign="top">6*</td>
<td align="left" valign="top">6</td>
<td align="left" valign="top">9</td>
<td align="left" valign="top">56</td>
<td align="left" valign="top">*</td>
</tr>
<tr>
<td align="left" valign="top">Kal8</td>
<td align="left" valign="top">4</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">5*</td>
<td align="left" valign="top">6</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">5</td>
<td align="left" valign="top">50</td>
<td align="left" valign="top">*</td>
</tr>
<tr>
<td align="left" valign="top">All</td>
<td align="left" valign="top">19</td>
<td align="left" valign="top">21</td>
<td align="left" valign="top">20*</td>
<td align="left" valign="top">19*</td>
<td align="left" valign="top">20</td>
<td align="left" valign="top">20</td>
<td align="left" valign="top">22*</td>
<td align="left" valign="top">22***</td>
<td align="left" valign="top">21</td>
<td align="left" valign="top">24</td>
<td align="left" valign="top">208</td>
<td align="left" valign="top">******</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F5">
<label>Figure 5</label>
<caption>
<p>Sample tongue shapes of inverse simulations of plain /o/ in /po/ &#8216;footprint&#8217; and rhotic /o&#734;/ in /kro&#734;/ &#8216;chest&#8217; (orange points = detailed ultrasound samples; purple points = inverse targets).</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-9019-g5.jpg"/>
</fig>
<p>Once the inverse simulations were complete, we proceeded with the second type of modeling by creating forward simulations of the Kalasha rhotic vowels using the muscle activations estimated in the inverse simulations. Specifically, for both the retroflex and bunched versions of the Kamviri rhotic approximants, we combined the token-wise means of the muscle activations of these sounds with the token-wise means of the activations that were estimated for the Kalasha plain vowels. This resulted in a further 10 sets of simulations of &#8220;pseudo&#8221; Kalasha rhotic vowels &#8211; five with a retroflex basis and five with a bunched basis &#8211; that were intended to serve as a point of comparison against the inverse simulations of the Kalasha rhotic vowels. To accomplish the superposition of the muscle activations, we used a simple additive combination rule, adding activations from each component articulation (plain vowel and rhotic approximant variant) for each muscle exciter. To model a range of possible coarticulatory blends of plain vowels and rhotic approximants, each rhotic vowel was simulated using 11 different mixtures of plain vowel and rhotic approximant activations, essentially a &#8220;crossfade&#8221; from 100% plain vowel to 100% rhotic approximant at 10% increments. This resulted in 11 steps of rhotic mix proportion for each of the five forward simulation rhotic vowels that could be compared with the inverse simulation of the same rhotic vowel.</p>
<p>We applied several measures to the simulated tongue postures in order to compare them, as illustrated in <xref ref-type="fig" rid="F6">Figure 6</xref>. Each tongue shape is represented by a polygon with 50 vertices. The centroid of the polygon is indicated by a dot in the middle of the polygon. Its <italic>x</italic> and <italic>y</italic> coordinates represent the overall advancement and height of the tongue body, respectively. The most posterior point on each polygon (at a point indicated by another dot) represents tongue root advancement. The tongue tip is defined as the most anterior point on the tongue polygon (indicated by a third dot). The <italic>x</italic> value of this point represents tongue tip advancement. The angle above the horizontal from the tongue centroid to the tongue tip (represented by a line segment) is an indicator of retroflexion. The tongue blade is taken to be the portion of the tongue that is 2&#8211;5 points (out of the 50 points) posterior to the tip. The angle between the two points defining the tongue blade (represented by a thick line between these two points) is another indicator of retroflexion.</p>
<fig id="F6">
<label>Figure 6</label>
<caption>
<p>Illustration of measurements applied to the simulated tongue shapes. Dots indicate tongue root, tongue centroid, and tongue tip. Lines indicate angle measure (tongue blade angle and angle from tongue centroid to tongue tip).</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-9019-g6.png"/>
</fig>
<p>Here we reprise the research questions introduced at the end of &#167;1 in terms of the inverse and forward simulations:</p>
<p>(a) <bold>Do the rhotic vowels retain the lingual articulation of the corresponding non-rhotic vowels?</bold> This is a tongue shape comparison between the inverse simulation rhotic vowels (representing modern Kalasha rhotic vowels) and the retroflex and bunched forward simulations (representing an earlier coarticulatory stage). Articulatory properties of the inverse simulation rhotic vowels that are not found in the forward simulations suggest an articulatory reorganization of the Kalasha vowel system.</p>
<p>(b) <bold>Do the rhotic vowels retain any signs of retroflexion of the historically present retroflex consonants?</bold> This is also a tongue shape comparison between the inverse simulation rhotic vowels and the retroflex and bunched forward simulations. Articulatory properties of the inverse simulation rhotic vowels that resemble the retroflex forward simulations in particular suggest that the modern bunched rhotic vowels retain articulatory signs of their retroflex origins.</p>
<p>(c) <bold>Do the formant frequencies of rhotic vowels differ from what would be expected from adding retroflexion or bunching to their non-rhotic counterparts?</bold> This is a formant comparison (based on the acoustic synthesis derived from the biomechanical models), comparing the inverse simulation rhotic vowels with the retroflex and bunched forward simulations, particularly looking for acoustic similarities between the inverse simulations and retroflex forward similations (suggesting that the modern Kalasha vowels are preserving acoustic details attributable to the original coarticulatory basis) and looking for signs that the inverse simulation vowels are acoustically more distinct than would be expected from the forward simulation vowels (suggesting compensation for the centralizing effects of rhoticity).</p>
<p>To address these questions, we have compared a set of inverse simulations based on actual productions by four Kalasha speakers (representing our best biomechanical models of actual Kalasha vowels) to 22 different forward simulations for each rhotic vowel, representing 11 different degrees of overlap between averaged inverse simulation Kalasha plain vowels and averaged inverse simulation Kamviri retroflex and bunched approximants (the rhotic mix proportion). Any point along the continuum from 100% plain vowel to 100% rhotic could have formed the basis for modern Kalasha rhotic vowels, so we consider these possibilities as a group, asking, for example, does the inverse simulation /o&#734;/ resemble <italic>any</italic> of the mixtures of /o/ and /&#635;/ or /&#633;/ among the forward simulations?</p>
</sec>
</sec>
<sec>
<title>4. Results</title>
<p><xref ref-type="fig" rid="F7">Figure 7</xref> shows the tongue body shape in the inverse simulations of the five non-rhotic and five rhotic vowels, and retroflex/bunched approximants of Kamviri. They are broadly consistent with what is shown for observed Kalasha vowels above in <xref ref-type="fig" rid="F3">Figure 3</xref>. The rhotic vowels involve more tongue front bunching than their non-rhotic counterparts, and for front vowels and /u&#734;/, they involve more tongue root retraction. The Kamviri retroflex approximant is characterized by a slightly raised (tip-up) tongue posture, whereas the bunched approximant exhibits a bunched (tip-down) tongue gesture. It can also be observed that the the Kamviri bunched approximant resembles the rhotic vowels of Kalasha.</p>
<fig id="F7">
<label>Figure 7</label>
<caption>
<p>Tongue shape comparisons in inverse simulation plain and rhotic vowels, and retroflex and bunched approximants. Tongue tip points to the right.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-9019-g7.png"/>
</fig>
<p><xref ref-type="fig" rid="F8">Figure 8</xref> shows the muscle activity used in these inverse simulations. Anterior fibers of genioglossus are used to produce /a/ and /a&#734;/. Medial genioglossus fibers are used to produce all four front vowels and rhotic /a&#734;/. Consistent with how vowels are observed to be produced in Kalasha, the greatest differences in overall tongue shape are observed between the five non-rhotic vowels. Consistent with this, extreme contraction of the posterior genioglossus fibers is seen only in /i/, and contraction of the hyoglossus is seen only in /a/. Inferior longitudinal is more active in rhotic vowels (retracting the tongue tip) than in corresponding non-rhotic vowels for all pairs except for /o o&#734;/, where the plain vowel involves more tongue retraction than the rhotic vowel. Styloglossus and superior longitudinal muscles are also generally more active in rhotic vowels than their non-rhotic counterparts. Unexpectedly, transversus is more active in /a&#734; o&#734;/ than in their non-rhotic counterparts, and verticalis is more active in /o u/ than in their rhotic counterparts. This may be because the inverse simulation is fed only information about the mid-sagittal plane. Geniohyoid is more active in /e&#734; a a&#734; o/. Kamviri&#8217;s bunched approximant is consistently characterized by higher muscle activation in genioglossus anterior, medial, and posterior, inferior longitudinal, styloglossus, and verticalis. Hyoglossus, superior longitudinal, and transversus muscles are actively involved in the production of retroflex approximant.</p>
<fig id="F8">
<label>Figure 8</label>
<caption>
<p>Tongue muscle activations in plain and rhotic vowels, and retroflex and bunched approximants.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-9019-g8.png"/>
</fig>
<p>There are some differences between our inverse simulation Kamviri rhotic approximants and Stavness et al.&#8217;s (<xref ref-type="bibr" rid="B70">2012</xref>) English /&#633;/ simulations. Their bunched /&#633;/ does not involve posterior genioglossus or styloglossus but it does involve a small contraction of transversus and it involves equal contraction of the inferior longitudinal and superior longitudinal. Their retroflex /&#633;/ involves medial genioglossus but not transversus, and it optionally involves hyoglossus.</p>
<p>Recall that the forward simulations were designed to simulate coarticulation of vowels and retroflex or bunched approximants. They were produced at 11 rhotic mix proportion steps ranging from 100% vowel and 0% rhotic (similar to the inverse simulation plain vowels) to 0% vowel and 100% rhotic (where all five rhotic vowels are identical because there is no vowel information included in the simulation). We are interested in whether any of the intermediate points resemble the inverse simulation rhotic vowels, which would support the idea that rhotic vowels originated from coarticulation between plain vowels and retroflex or bunched consonants without much further phonetic development. The forward simulated combinations of plain vowels and rhoticity are considered to be vowel qualities that naturally occur in the event of coarticulation.<xref ref-type="fn" rid="n4">4</xref> We are interested in how they differ from inverse simulation Kalasha rhotic vowels, because these differences point to ways in which the actual Kalasha rhotic vowels have been established as separate vowel categories that are distinct from vowel+rhotic superposition.</p>
<p>In each panel of <xref ref-type="fig" rid="F9">Figure 9</xref>, thin horizontal lines indicate the formant frequencies of the inverse simulation plain and rhotic vowels for one vowel quality. In all cases the rhotic vowel has lower F3 frequency, and in most cases the rhotic vowel has less extreme F1 and F2 frequencies. From left to right within each panel, the thick solid contour shows the formant frequencies of the forward simulation as more and more retroflex approximant (and less and less plain vowel) is mixed in. The thick dashed contour shows the formant frequencies of the forward simulation as more and more bunched approximant is mixed in. The thick vertical lines indicate the step at which the retroflex and bunched forward simulation vowels are acoustically most similar to the inverse simulation rhotic vowels according to the Root Mean Squared Error (RMSE) for the bark-scaled formant frequencies. In some cases, like /a/ and /u/, the 0% rhotic mix steps are fairly close to the inverse simulation non-rhotic vowel formants. In other cases, there are some differences, attributable to the fact that the inverse simulation formant values are based on multiple tokens averaged in acoustic space, whereas the forward simulation formant values are those produced by an average articulatory configuration.</p>
<fig id="F9">
<label>Figure 9</label>
<caption>
<p>Formant frequencies of forward simulation rhotic vowels ranging from 0 to 100% rhotic mix proportion, with reference formant values for inverse simulation plain and rhotic vowels.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-9019-g9.png"/>
</fig>
<p>For the /i i&#734;/ and /e e&#734;/ simulations, as more of either type of rhotic is mixed in, F1 mostly increases and F2 and F3 mostly decrease. The step at which the forward simulation of bunched /i&#734;/ most resembles the inverse simulation rhotic vowel is 50%, an equal mix of plain /i/ and a rhotic approximant. For retroflex /i&#734;/, 60% rhotic most resembles the acoustics of the inverse simulation rhotic vowel. The /e e&#734;/ simulations most closely resemble the inverse simulation [e&#734;] when 70% rhotic is mixed in for both retroflex and bunched versions.</p>
<p>For the back vowels, the picture is a bit different. For /a a&#734;/, the most similar steps for retroflex and bunched /a&#734;/ are 90% rhotic for bunched and 100% rhotic for retroflex. This is consistent with the fact that the model rhotics used for these simulations were produced in an /a/ context, so the rhotic approximant itself is already similar to a blend of /a/ and a rhotic approximant.<xref ref-type="fn" rid="n5">5</xref> Forward simulation /o&#734;/ is acoustically most similar to inverse simulation /o&#734;/ at 70% retroflex rhotic and 90% bunched rhotic. Forward simulation /u&#734;/ is acoustically most similar to inverse simulation /u&#734;/ at 100% retroflex rhotic (similar to the other back vowels) and 40% bunched rhotic (similar to the other high vowel). Forward simulation back vowels tend to resemble the inverse simulation rhotic vowels near the 100% rhotic end of the scale. This is consistent with the fact that the Kalasha back rhotic vowels are articulatorily quite different from the corresponding plain vowels (as shown in <xref ref-type="fig" rid="F2">Figure 2</xref>), and thus the plain vowel portion of the mixture is not as helpful as it is for the front rhotic vowels. The fact that the plain back vowels do not help much is also a clue that some of the Kalasha rhotic vowels are not articulated in a way that can be predicted from combining the corresponding plain vowel with a gesture for rhoticity.</p>
<p><xref ref-type="fig" rid="F10">Figure 10</xref> shows how the forward simulation vowels move through F1&#8211;F2 and F2&#8211;F3 space as they become increasingly rhotic. The forward simulation 100% plain vowels are indicated by IPA symbols /i e a o u/, and the 100% retroflex and bunched rhotic ends of the scales are indicated by &#8220;&#635;&#8221; and &#8220;&#633;&#8221;. The inverse simulation formant frequencies are indicated by small black IPA symbols, and the observed formant frequencies are indicated by small gray IPA symbols. Even though they differ from the observed formant frequencies, the inverse simulation formant frequencies are a more appropriate point of comparison for the forward simulation formant frequencies. Comparing the forward simulations with the observed Kalasha vowels would conflate effects of coarticulation with effects of biomechanical modeling, but comparing the forward simulations with the inverse simulations isolates the effects of coarticulation. The paths from 0% rhotic to 100% rhotic are indicated by curves (solid for retroflex and dashed for bunched), and each step along the way is indicated by a dot. The front vowels pass by the inverse simulation rhotic versions as they become more rhotic, and the back vowels generally do not. This suggests that the present-day /i&#734; e&#734;/ are similar acoustically to coarticulated /i e/ and a rhotic consonant, either bunched or retroflex, while the back vowels generally are not similar. As /a/ becomes more rhotic, it moves away from /a&#734;/ in F1&#8211;F2 space, and becomes more similar to it mainly by dropping F3 as it approaches the rhotic end of the scale. /o/ becomes more similar to /o&#734;/ only in F1. /u/ approaches /u&#734;/, but the inverse simulation version of the rhotic vowel is located beyond the 100% rhotic endpoints in terms of F2. One thing that these back vowel mismatches have in common is that the inverse simulation rhotic vowels (and for /a&#734; o&#734;/ also the observed vowels) have high F2 frequencies that are not accounted for by the superposition of rhoticity. This is likely because the rhotic counterparts of back vowels are produced with tongue fronting that is not accounted for by the gestures used to produce plain back vowels or rhotics.</p>
<fig id="F10">
<label>Figure 10</label>
<caption>
<p>Paths through formant space as vowels become more rhotic: Forward simulation rhotic vowels ranging from 0 to 100% rhotic mix proportion in F1&#8211;F2 space (top) and F2&#8211;F3 space (bottom). Plain vowel IPA symbol endpoint is 0% rhotic and [&#635;] and [&#633;] endpoints are 100% rhotic (retroflex and bunched, respectively). Steps are at 10% increments.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-9019-g10.png"/>
</fig>
<p><xref ref-type="fig" rid="F11">Figures 11</xref>, <xref ref-type="fig" rid="F12">12</xref>, <xref ref-type="fig" rid="F13">13</xref>, <xref ref-type="fig" rid="F14">14</xref>, <xref ref-type="fig" rid="F15">15</xref> show the shape of the simulated tongue for all the steps of the retroflex and bunched forward simulations and the acoustic distance between each of these steps and the corresponding inverse simulation rhotic vowel. Each of the five figures contains four panels showing the same types of information for each of the five rhotic vowels. The top two panels show tongue shapes and the bottom two panels show acoustic distance. The left two panels show retroflex forward simulations and the right two panels show bunched forward simulations. In the top tongue trace panels, all 11 steps of the forward simulation (from 0% rhotic to 100% rhotic) are shown with solid lines. The step that is acoustically most similar to the rhotic vowel is indicated by a heavier contour.</p>
<fig id="F11">
<label>Figure 11</label>
<caption>
<p>Mid-sagittal tongue shapes (top) and acoustic distances (bottom) for all steps of the retroflex (left) and bunched (right) forward simulations of /i&#734;/. Tongue tip is to the right in tongue traces. The step that is acoustically most similar to the inverse simulation rhotic vowel is indicated by a heavy contour in the top panel and a rhotic vowel IPA symbol in the bottom panel. The least rhotic step has the brightest shading and the most rhotic step has the darkest shading, in both the top and bottom panels.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-9019-g11.png"/>
</fig>
<fig id="F12">
<label>Figure 12</label>
<caption>
<p>Mid-sagittal tongue shapes (top) and acoustic distances (bottom) for all steps of the retroflex (left) and bunched (right) forward simulations of /e&#734;/.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-9019-g12.png"/>
</fig>
<fig id="F13">
<label>Figure 13</label>
<caption>
<p>Mid-sagittal tongue shapes (top) and acoustic distances (bottom) for all steps of the retroflex (left) and bunched (right) forward simulations of /a&#734;/.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-9019-g13.png"/>
</fig>
<fig id="F14">
<label>Figure 14</label>
<caption>
<p>Mid-sagittal tongue shapes (top) and acoustic distances (bottom) for all steps of the retroflex (left) and bunched (right) forward simulations of /o&#734;/.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-9019-g14.png"/>
</fig>
<fig id="F15">
<label>Figure 15</label>
<caption>
<p>Mid-sagittal tongue shapes (top) and acoustic distances (bottom) for all steps of the retroflex (left) and bunched (right) forward simulations of /u&#734;/.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-9019-g15.png"/>
</fig>
<p>The bottom acoustic distance panels show the Root Mean Squared Error (RMSE) for the bark-scaled formant frequencies of all the forward simulation rhotic vowels, with the same <italic>x</italic>-axis scale as <xref ref-type="fig" rid="F9">Figure 9</xref>. Non-rhotic IPA symbols are always at the first (0% rhotic) step and rhotic IPA symbols are located at the step of each series (retroflex and bunched) that is most similar acoustically to the inverse simulation rhotic vowel (corresponding to the heavy tongue trace in the panel above it and the vertical lines in <xref ref-type="fig" rid="F9">Figure 9</xref>).</p>
<p>The acoustic distances in the bottom panels of these figures show that the best matches for the front vowels are closer matches than the best matches for the back vowels, which mostly occur close to the 100% rhotic end of the scale. In other words, the inverse simulation rhotic front vowels /i&#734; e&#734;/ are intermediate between the corresponding plain front vowel and a rhotic approximant, while the inverse simulation rhotic non-high back vowels /a&#734; o&#734;/ are not very similar to any of the mixture proportions, but most similar to the rhotic end of the scale. The bunched rhotic high back vowel /u&#734;/ is more similar to the front vowels, and the retroflex rhotic /u&#734;/ is more similar to the back vowels.</p>
<p><xref ref-type="fig" rid="F16">Figures 16</xref>, <xref ref-type="fig" rid="F17">17</xref> show how the various simulations compare according to the articulatory parameters illustrated in <xref ref-type="fig" rid="F6">Figure 6</xref>. The distributions of inverse simulation plain and rhotic vowels are represented by boxes, and the individual steps of the forward simulations are represented by circles in between them, with the 11 retroflex simulations (labeled with &#635;) always to the left of the 11 bunched simulations (labeled with &#633;), with each group getting more rhotic from left to right. The circle for the step that is acoustically most similar to the inverse simulation rhotic is filled. Note that for every vowel, there is no difference between the first step of the retroflex and bunched forward simulation series, because no rhotic is mixed in to the plain vowel, and that the final step for every retroflex (or bunched) simulation is the same as all the other retroflex (or bunched) simulations, because no vowel is mixed in.</p>
<fig id="F16">
<label>Figure 16</label>
<caption>
<p>Tongue body and tongue root measures for all vowel simulations. Distributions of simulated plain and rhotic vowels are represented by boxes. Each series of 11 forward simulations is represented by circles, 0&#8211;100% retroflex first, followed by 0&#8211;100% bunched, with a filled circle for the forward simulation acoustically closest to the inverse simulation rhotic vowel.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-9019-g16.png"/>
</fig>
<fig id="F17">
<label>Figure 17</label>
<caption>
<p>Tongue tip and tongue blade measures for all vowel simulations. Distributions of simulated plain and rhotic vowels are represented by boxes. Each series of 11 forward simulations is represented by circles, 0&#8211;100% retroflex first, followed by 0&#8211;100% bunched, with a filled circle for the forward simulation acoustically closest to the inverse simulation rhotic vowel.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-9019-g17.png"/>
</fig>
<p>In terms of tongue body height, all of the forward simulations are close to the inverse simulation rhotic vowel distribution at the step that is acoustically most similar to it (the filled circle), with the exception of retroflex /o&#734;/. Adding retroflexion to /o/ does not provide any of the tongue body raising that is observed in Kalasha /o&#734;/.</p>
<p>In terms of tongue body advancement, the front and back vowels react differently to the addition of rhoticity. The retroflex /i&#734;/ and /e&#734;/ match the inverse simulation rhotic vowels at the acoustically most similar step, but the corresponding bunched simulations do not match the inverse simulation rhotic vowels until additional bunching is added. Retroflex /a&#734;/ has the right amount of tongue body advancement, but the acoustically most similar bunched step overshoots it by a little. The other back vowels have an appropriate amount of tongue body advancement only in the bunched simulations. The forward simulation retroflex versions of /o&#734;/ and /u&#734;/ have insufficient tongue body advancement.</p>
<p>For tongue root advancement, all /i&#734;/ and /e&#734;/ simulations get close to the corresponding inverse simulations. The /a&#734;/ simulations do not show any tongue root advancement relative to the plain category until the last step of each series. None of the steps of the /o&#734;/ forward simulations have the tongue root advancement that is seen in /o&#734;/. Like the front vowel simulations, the /u&#734;/ simulations span a wide range of tongue root positions that include the degree of retraction of the inverse simulation /u&#734;/, but resemblance to the inverse simulation occurs at the acoustically most similar step only for the bunched simulation.</p>
<p>Moving to the tongue tip and blade (<xref ref-type="fig" rid="F17">Figure 17</xref>), the bunched and retroflex simulations produce too little tongue tip retraction for /i&#734; e&#734; a&#734;/ at the step that is closest acoustically to the inverse model, although most of them reach the appropriate amount of retraction at a more extreme step. The forward simulation bunched /o&#734;/ has a good amount of tip advancement, but the retroflex simulation is too retracted. The forward simulation retroflex /u&#734;/ has a good amount of tip retraction, but the bunched simulation is too advanced. The angle to the tongue tip is too low for all of the bunched simulations and too high for the retroflex /o&#734; u&#734;/ simulations. Tongue blade angle is too high for all retroflex simulations except /i&#734;/, and too low for all bunched rhotic simulations.</p>
<p>Another way that the forward simulation rhotic vowels differ from observed Kalasha rhotic vowels is that their lip posture is interpolated between the plain vowel and the rhotic approximant, but observed Kalasha rhotic vowels have virtually the same lip posture as the corresponding non-rhotic vowels. The rhotic approximants used for the forward simulations are produced with a small lip opening that is closer to /u/ than any of the other non-rhotic vowels, but it is achieved by a relatively closed jaw and relatively relaxed facial muscles rather than by orbicularis oris contraction as in /u/. So most of the forward simulation rhotic vowels generally have a smaller lip opening than the observed Kalasha rhotic vowels, but less lip protrusion in /u&#734;/. This lack of variation in lip posture is likely to be one reason for the lack of acoustic differentiation between the forward simulation rhotic vowels.</p>
<p>We also note that all of the observed Kalasha rhotic vowels have higher F1 than their non-rhotic counterparts, as seen in <xref ref-type="fig" rid="F2">Figure 2</xref>, but in the inverse and forward simulations all the rhotic vowels are closer to the middle of the F1 range than their non-rhotic counterparts. This may be accounted for by an additional phonetic feature of Kalasha rhotic vowels that has not been included in the modeling: Kalasha rhotic vowels appear to be produced with a raised larynx voice quality. We have recognized this as an auditory feature of rhotic vowels and we have noticed large changes in the angle of the hyoid bone shadow in our ultrasound images. Larynx raising shortens the vocal tract, raising the frequency of F1 in particular. If larynx raising were included in the inverse simulation rhotic vowels, their F1 values would be higher, more in line with the observed rhotic vowels. In any case, larynx raising appears to be an additional feature of Kalasha rhotic vowels that is not accounted for by coarticulation to rhotic approximants.</p>
</sec>
<sec>
<title>5. Discussion</title>
<p>This study investigated the development of rhoticity in a vowel system using biomechanical modeling. The likely sources of Kalasha rhotic vowels are deleted postvocalic retroflex consonants (by sound change) and Nuristani rhotic approximants (by contact). It is reasonable to consider the initial state of Kalasha rhotic vowels to be a plain vowel coarticulated with a rhotic approximant. So we have sought to find out how similar present-day Kalasha rhotic vowels are to a rhotic approximant blended with the corresponding plain vowel. Here we revisit the questions from the introduction, where we asked whether the development of rhotic vowels from plain vowel + rhoticity is basically additive from an articulatory standpoint (i.e., whether the rhotic vowels retain the lingual articulation of the corresponding non-rhotic vowels, whether they retain any signs of retroflexion from the historically present retroflex consonants, and whether the acoustic properties of rhotic vowels resemble what would be expected from adding retroflexion or bunching to the corresponding non-rhotic vowels). The genesis of rhotic vowels strongly points to retroflexion (given that it was the retroflex class of consonants that triggered the change).</p>
<sec>
<title>5.1. Answers to research questions</title>
<p><bold>Do the rhotic vowels retain the lingual articulation of the corresponding non-rhotic vowels?</bold> The Kalasha plain vowels /i e a o u/ differ from each other in tongue height and advancement in the expected ways, as seen above in <xref ref-type="fig" rid="F16">Figure 16</xref>: /i e/ have a relatively advanced tongue body, /u o a/ have a relatively retracted tongue body, and within those front and back groups the vowels are distinguished by tongue body height. /i/ has the most advanced tongue root, followed by /e u/ and then /a o/. For most of these vowels, adding either bunching or retroflexion is expected to increase tongue body height and neutralize advancement while also increasing tongue root retraction.</p>
<p>The inverse simulation rhotic vowels /i&#734; e&#734; o&#734; u&#734;/ all have similar tongue body height, and /a&#734;/ is lower. The retroflex and bunched forward simulations all generally capture these tongue height effects, indicating that tongue height in most rhotic vowels is attributable to the results of coarticulation. The exception is retroflex /o&#734;/, which shows none of the tongue body raising observed in /o&#734;/.</p>
<p>Turning to tongue body advancement, the inverse simulation /e&#734; o&#734; u&#734;/ have advancement similar to each other and similar to plain /u/. /i&#734;/ is more advanced and /a&#734;/ is less advanced. The retroflex forward simulations capture the tongue body advancement of /i&#734; e&#734; a&#734;/ well, and the bunched forward simulations have steps that match the degree of tongue body advancement seen in the inverse simulation, but the step that is most similar acoustically has too much tongue body advancement for these three vowels. For /o&#734; u&#734;/, the bunched forward simulations are closer. Retroflex inverse simulation /o&#734;/ has insufficient advancement and retroflex inverse simulation /u&#734;/ introduces unnecessary retraction. As with tongue body height, introducing retroflexion to plain /o/ does not yield appropriate tongue body advancement for /o&#734;/.</p>
<p>Turning to tongue root advancement, the inverse simulation /i&#734; o&#734; u&#734;/ have similar tongue root position, while /e&#734; a&#734;/ show tongue root retraction similar to plain /a o/. The bunched and retroflex forward simulations of /i&#734; e&#734; u&#734;/ all include steps with tongue root retraction similar to the inverse simulation rhotic vowels, although the retroflex /u&#734;/&#8217;s tongue root retraction match occurs at a step that is not similar acoustically to inverse simulation /u&#734;/. None of the forward simulations of /a&#734; o&#734;/ approach the tongue root advancement that is seen in these vowels, which is particularly large for /o&#734;/.</p>
<p>In summary, the advancement and raising of /o&#734;/ is not accounted for by coarticulation to any type of rhotic. It is similar in tongue posture to /e&#734;/, from which it is distinguished by lip rounding. Hussain and Mielke (<xref ref-type="bibr" rid="B33">2021</xref>) concluded that /o&#734;/ might better be classified as a front vowel, and we have shown that this fronting is not accounted for by coarticulation to a rhotic consonant. Beyond this fact about /o&#734;/, bunching does the best job of accounting for the tongue position of back rhotic vowels and retroflexion does the best job of accounting for the tongue position of front rhotic vowels.<xref ref-type="fn" rid="n6">6</xref></p>
<p>The addition of rhoticity to plain vowels does a good job of capturing the tongue body and tongue root position found in the front rhotic vowels /i&#734; e&#734;/, and the retroflex version is more accurate, particularly for tongue body advancement. On the other hand, adding bunching to the back vowels yields a closer match to the rhotic vowels than adding retroflexion. Retroflexion misses the tongue body and tongue root advancement of /o&#734; u&#734;/ and the tongue body raising of /o&#734;/. Bunching provides a much closer approximation of /o&#734; u&#734;/ but misses the tongue root advancement of /o&#734;/ and underestimates its tongue body advancement. The bunched and retroflex forward simulations of /a&#734;/ are similar, and the main miss is the slight tongue root advancement.</p>
<p>A possible interpretation of these facts is that the front vowels straightforwardly reflect the effects of coarticulation to retroflex consonants, and the tongue body fronting observed in back vowels may have been introduced at a stage when the dominant articulatory strategy for rhotic vowels shifted from retroflexion to bunching. This shift seems likely to have occurred once the vowel+retroflex consonant sequences were reinterpreted as vowels. Rhotic vowels have been found to be predominantly bunched (more so than rhotic approximants) in languages where they have been studied articulatorily (<xref ref-type="bibr" rid="B37">Jiang, Chang, &amp; Hsieh, 2019</xref>; <xref ref-type="bibr" rid="B55">Mielke, 2015</xref>; <xref ref-type="bibr" rid="B56">Mielke et al., 2016</xref>).</p>
<p><bold>Do Kalasha rhotic vowels retain any signs of retroflexion of the historically present retroflex consonants?</bold> Although the direct source of Kalasha rhotic vowels is believed to be coarticulation to a following retroflex (not bunched) consonant, Hussain and Mielke (<xref ref-type="bibr" rid="B34">2022</xref>) found only bunched variants of Kalasha rhotic vowels, for all vowel qualities. Here we simulated coarticulation to both retroflex and bunched rhotic approximants. As seen above in <xref ref-type="fig" rid="F17">Figure 17</xref>, the inverse simulation rhotic vowels all have higher tongue blade angle and a higher angle to the tongue tip than their non-rhotic counterparts, and all but /o&#734;/ have a more retracted tongue tip. The bunched forward simulation rhotic vowels achieve a much better match to the inverse simulation rhotic vowels in terms of tongue blade angle, and they are comparable in terms of tongue tip advancement and the angle to the tongue tip. Coarticulation to the Kamviri bunched approximant provides a clearly better match to Kalasha rhotic vowels than coarticulation to the Kamviri retroflex approximant does, and this difference is most apparent in tongue blade angle. Kalasha /o&#734;/ has a more advanced tongue tip than would be expected for a retroflex version of /o/ because it is a more advanced vowel than /o/, but the bunched simulation nevertheless does a good job of approximating the observed tongue tip advancement.</p>
<p><bold>Do the formant frequencies of rhotic vowels differ from what would be expected from adding retroflexion or bunching to their non-rhotic counterparts?</bold> Nearly all of the ways that the observed Kalasha rhotic vowels differ from their plain counterparts are replicated in the inverse simulation vowels, although the magnitudes of the differences vary a lot. This is shown above in <xref ref-type="fig" rid="F10">Figure 10</xref>. The rhotic vowels all have lower F3 than their plain counterparts. F2 is lower in front rhotic vowels /i&#734; e&#734;/ and higher in back rhotic vowels /a&#734; o&#734; u&#734;/ relative to their plain counterparts. F1 is higher in /i&#734; u&#734; e&#734;/ and lower in /a&#734;/, and it is similar between /o&#734;/ and its non-rhotic counterpart. This differs somewhat from the wholesale F1 increase shown in <xref ref-type="fig" rid="F2">Figure 2</xref>, and we have suggested that the F1 increase in Kalasha rhotic vowels could be due to larynx raising, which is not included in any of the simulations.</p>
<p>The bunched and retroflex forward simulations both do a good job of capturing the F1 increase and F2 and F3 reduction observed in /i&#734; e&#734;/. Both forward simulations of /a&#734;/ have F1 and F2 that are too low. The bunched and retroflex forward simulations differ the most for the two back rounded rhotic vowels /o&#734; u&#734;/, with higher F3 for bunched /o&#734;/ and higher F2 for bunched /u&#734;/, relative to the retroflex versions of these vowels. The bunched models are somewhat better at capturing the F3 values (the retroflex models yield too much F3 decrease), but all of the models of these vowels are too low in F2. They do not account for the acoustic centralization or fronting of these vowels. In summary, there does not seem to be any evidence that any acoustic details of the Kalasha rhotic vowels are attributable specifically to the retroflexion in their history rather than rhoticity in general.</p>
<p>An important possibility to consider is that Kalasha rhotic vowels are somehow optimized to maintain perceptual distinctiveness within the entire vowel system. This is explored in the next subsection.</p>
</sec>
<sec>
<title>5.2. Vowel dispersion</title>
<p>Kalasha appears to have once had a system of five oral vowels distinguished primarily by F1 and F2, and now it has a system of ten oral vowels distinguished by F1, F2, and F3.<xref ref-type="fn" rid="n7">7</xref>&#160;<xref ref-type="fig" rid="F18">Figure 18</xref> shows how adding rhoticity affects the dispersion of vowels in F1&#8211;F2&#8211;F3 space. The top panel shows that at a rhotic proportion of zero, all rhotic vowels are identical to their non-rhotic counterparts, and at a rhotic proportion of one, all rhotic vowels are identical to each other. As the rhotic proportion is increased from zero to one, the difference within rhotic-non-rhotic pairs generally increases, and the difference between rhotic vowels generally decreased. The curves for rhotic vowels and rhotic-non-rhotic pairs cross at a rhotic proportion between 0.7 and 0.8, where the average distances between vowels are closely matched by the average distances between inverse simulation vowels. This is also a rhotic mix proportion that results in vowels that most closely match the inverse simulation acoustically. At low rhotic mix proportions, bunched rhotic vowels are somewhat less distinct from each other than similar retroflex rhotic vowels, but these differences disappear before the rhotic mix proportion reaches 0.7. We note that for /i&#734; e&#734;/, the two vowels with a good acoustic match between inverse and forward simulations, the best match was found at a rhotic mix proportion of 0.5&#8211;0.7 (a mixture of 50&#8211;70% bunched or retroflex approximant and 30&#8211;50% plain vowel). For the vowels with a relatively poor acoustic match, the closest step tended to be closer to the 1.0 rhotic approximant end of the scale. Thus, maximizing the acoustic dispersion of the forward simulation vowels and maximizing their acoustic similarity to inverse simulation vowels both point to 70% rhotic approximant and 30% plain vowel as a reasonable mixture, at least for rhotic front vowels, or at least for vowels that resemble rhotic versions of plain vowels.</p>
<fig id="F18">
<label>Figure 18</label>
<caption>
<p>Vowel dispersion by rhotic mix proportion and simulation type. Mean Euclidean distance (based on bark-scaled F1, F2, and F3) among all pairs of vowels, rhotic vowels only, and rhotic-non-rhotic pairs only (top); all rhotic-non-rhotic pairs (middle); all pairs of rhotic vowels (bottom).</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-9019-g18.png"/>
</fig>
<p>The middle and bottom panels show all rhotic-non-rhotic pairs, and all pairs of rhotic vowels, respectively. Liu and Kewley-Port (<xref ref-type="bibr" rid="B48">2004</xref>) identify 0.37 barks as a threshold for listeners to distinguish vowels.<xref ref-type="fn" rid="n8">8</xref> All pairs of vowels at rhotic mix proportions from 0.2 to 0.9 exceed this threshold. The most indistinct rhotic-non-rhotic pair is /o o&#734;/. The most indistinct pairs of rhotic vowels are bunched /a&#734; o&#734;/, retroflex /e&#734; i&#734;/, and bunched /e&#734; i&#734;/. Pairs of vowels differing in both quality and rhoticity are not depicted here, but for all rhotic mix proportions from 0.2 to 0.7, and for both retroflex and bunched simulations, the most indistinct among these pairs is /e i&#734;/.</p>
<p>If Kalasha vowels are more dispersed than would be expected based on coarticulation, we expect the inverse simulation vowels to be more dispersed than the forward simulation vowels. Looking at the whole system, the inverse simulation vowels are no more dispersed than any step of rhotic proportion, and bunching does not make for a more dispersed vowel system than retroflexion. The mean dispersion within the most vulnerable vowels (the rhotic-non-rhotic pairs and the pairs of rhotic vowels) is very similar between the inverse simulation vowels and the forward simulation vowels at a rhotic mix proportion of 0.7, which is also the point where the forward simulation rhotic tended to match the inverse simulation vowels most closely. There does appear to be some exaggeration of the differences in particular vulnerable pairs of vowels. /o o&#734;/ is the least distinct pair of vowels in the forward simulations, but it is the most distinct pair of vowels in the inverse simulations, and in the observed Kalasha tongue shapes. As previously discussed, /o&#734;/ is considerably different from /o/ in articulation, and probably better described as a front vowel with a tongue shape similar to /e&#734;/ combined with lip rounding similar to /o/.</p>
<p>Some Kalasha vowels are more dispersed acoustically than they would be if they were simply plain vowels combined with either type of a rhotic approximant. This is consistent with adaptive dispersion (<xref ref-type="bibr" rid="B15">de Boer, 2000</xref>; <xref ref-type="bibr" rid="B22">Flemming, 2002</xref>; <xref ref-type="bibr" rid="B46">Lindblom, 1986</xref>, <xref ref-type="bibr" rid="B47">1990</xref>; <xref ref-type="bibr" rid="B64">Padgett &amp; Tabain, 2005</xref>) manifesting early in the development of a new vowel sub-system. The observed acoustic dispersion involves articulatory gestures not explained by the rhotic approximants or the plain vowels believed to form the basis for the new rhotic vowels. However, we do not observe dispersion generally across the Kalasha vowel system. If it is an active force here, it appears to be limited to the most vulnerable pairs of vowels that are created by adding rhoticity to the system.</p>
</sec>
<sec>
<title>5.3. Comparison to nasal vowel subsystems</title>
<p>The most likely sequence of events in the development of Kalasha vowels seems to be that coarticulatory retroflexion was exaggerated and extended to more of the vowel interval, introducing new vowel qualities that were achieved through tongue bunching by later generations. When the vowels coalesced with rhoticity, they retained their characteristic lip postures. In the course of developing new vowel categories, the back rhotic vowels, and /o&#734;/ in particular, were established as front rounded vowels. The development of rhotic vowels in Canadian French (e.g., [pn&#248;] <italic>&#8764;</italic> [pn&#602;] &#8216;tire&#8217;) has shown a similar reorganization from a different starting point: rhotic vowels developed from front rounded vowels, which were undergoing backing over time (<xref ref-type="bibr" rid="B54">Mielke, 2013</xref>). The modern rhotic vowels are produced as retroflex by some speakers (<xref ref-type="bibr" rid="B55">Mielke, 2015</xref>). The Kalasha /o&#734;/ is phonetically similar to the Canadian French rhotic /&#248;/, both produced with lip rounding and a bunched tongue in the front of the oral cavity, despite one originating from a back vowel and retroflexion and the other originating from a front vowel and no articulatory source of retroflexion. See Hussain and Mielke (<xref ref-type="bibr" rid="B34">2022</xref>) for further discussion of these two cases.</p>
<p>The reorganization of the rhotic vowel subsystem recalls the reorganization that is seen in nasalized vowel subsystems. Vowel nasalization changes the acoustic output of the vocal tract in ways that interact with the perception of vowel quality, such as by making high vowels sound lower and low vowels sound higher (<xref ref-type="bibr" rid="B18">Diehl et al., 1991</xref>; <xref ref-type="bibr" rid="B21">Feng &amp; Castelli, 1996</xref>; <xref ref-type="bibr" rid="B23">Fujimura &amp; Lindqvist, 1971</xref>; <xref ref-type="bibr" rid="B68">Serrurier &amp; Badin, 2008</xref>). As such, nasalization is articulatorily independent but acoustically integrated with vowel quality. The acoustic effects of vowel nasalization may be enhanced or compensated for through direct changes to tongue or pharynx posture (<xref ref-type="bibr" rid="B3">Barlaz et al., 2015</xref>; <xref ref-type="bibr" rid="B4">Beddor, 1982</xref>; <xref ref-type="bibr" rid="B10">Carignan, 2014</xref>, <xref ref-type="bibr" rid="B11">2018</xref>; <xref ref-type="bibr" rid="B12">Carignan, Shosted, Fu, Liang, &amp; Sutton, 2015</xref>; <xref ref-type="bibr" rid="B39">Krakow et al., 1988</xref>). Articulatory enhancement of vowel nasalization is also seen in Kalasha /&#227; &#245; &#361;/ and /&#227;&#734;/, which appears to be merged with /&#7869;&#734;/ in some speakers (<xref ref-type="bibr" rid="B33">Hussain and Mielke, 2021</xref>).</p>
<p>Since rhoticity is achieved largely with the tongue, it is less independent of vowel quality than nasalization is. The articulatory gestures that help achieve acoustic characteristics of rhoticity such as low F3 also directly affect the quality of vowels as realized through F1 and F2 by changing the shape of the oral cavity and pharynx. We have seen that the present-day Kalasha rhotic vowel subsystem has apparently compensated for some of these effects through enhancement of F1 and F2 differences among the rhotic vowels. It would not have been possible to observe this as enhancement without the reference point provided by the biomechanical and acoustic modeling.</p>
</sec>
<sec>
<title>5.4. Comparison to pre-rhotic vowel subsystems</title>
<p>It seems very likely that Kalasha went through a stage with distinct pre-rhotic vowel allophones prior to developing monophthongal rhotic vowels from coarticulated vowel-rhotic sequences. North American English has a bunched/retroflex rhotic approximant that affects the quality of vowels around it. We can examine the pre-rhotic vowel system of North American English for possible clues about the earlier development of Kalasha rhotic vowels before they became established as monophthongs.</p>
<p>North American English is typically described as having one rhotic vowel quality /&#602;/ (which may be transcribed /&#604;&#734;/ or /&#633;&#809;/), but it has many distinct pre-rhotic vowel variants. Thomas (<xref ref-type="bibr" rid="B76">2001, p. 44</xref>) summarizes the effects of following /&#633;/ on English vowels as follows:</p>
<disp-quote>
<p>Coarticulation with /r/ shrinks the vowel space of pre-/r/ vowels and obliterates certain cues such as gliding that are used to distinguish vowels. These effects, in turn, lead to difficulty by speakers in identifying pre-/r/ vowels with particular vowel phonemes&#8230;. They also lead to mergers&#8230;</p>
</disp-quote>
<p>Kalasha vowels are generally monophthongal (<xref ref-type="bibr" rid="B33">Hussain &amp; Mielke, 2021</xref>), so they are less susceptible to changes in gliding, and Kalasha has fewer non-rhotic oral vowel qualities than English, so mergers may be less likely. The syllabic R of North American English is the result of a merger of Middle English /&#603;&#633; &#618;&#633; &#650;&#633; &#604;&#633; &#601;&#633;/ (<xref ref-type="bibr" rid="B80">Wells, 1982</xref>). None of these vowel qualities are found in Kalasha. The remaining English pre-rhotic vowels are more similar to the Kalasha vowels that merged with rhotics.</p>
<p>Thomas (<xref ref-type="bibr" rid="B76">2001, pp. 44&#8211;48</xref>) describes North American English monophthongs before /&#633;/ as follows. /i&#633;/ has retracted and sporadically merged with /e&#633;/. /e&#633;/ lacks the upgliding found in non-pre-rhotic contexts and it is lowered and retracted (and sporadically merged with /i&#633;/). /&#593;&#633;/ varies considerably in the front-back dimension ([&#593;&#633;] to [&#230;&#633;]) and also may be rounded. /o&#633;/ has mostly merged with /&#596;&#633;/ and is pronounced as [o&#633;] (without the upgliding found in non-pre-rhotic contexts). In words which historically had /u&#633;/, this sequence has merged with /&#604;&#734;/ or /o&#633;/, or it has been reinterpreted as bisyllabic /u&#602;/. These patterns in English are largely consistent with Kalasha rhotic vowels. We have seen that rhoticity causes F2 decrease in Kalasha front vowels, and /e&#734; i&#734;/ are very close together in the simulations. Kalasha /a&#734;/ is acoustically fronted, much like English /&#593;&#633;/. Kalasha /o&#734;/ is quite different from English /o&#633;/, and much more like English /&#604;&#734;/. Kalasha /u&#734;/ is acoustically lowered, which could have caused it to merge with /o&#734;/ if /o&#734;/ had not moved toward the front of the vowel space. Indeed, /u&#734;/ and /o&#734;/ are very close in the forward simulations, which do not take into account the fronting of Kalasha /o&#734;/. The fronting of /o&#633;/ to avoid a merger with /u&#633;/ would have been more problematic in North American English, with /&#602;/ sitting in front of /o&#633;/ in the vowel space.</p>
<p>Thomas (<xref ref-type="bibr" rid="B76">2001, p. 44</xref>) notes that the distinct pre-rhotic variants of English vowels are difficult for speakers to connect to vowel categories occurring in non-pre-rhotic contexts. While the conventional IPA transcription /i&#734; e&#734; a&#734; o&#734; u&#734;/ suggests a close relationship between rhotic vowels and their non-rhotic counterparts, we have seen that some pairs, most notably /o o&#734;/, bear very little resemblance to each other in the way of lingual articulation. The rhotic vowels are spelled &lt;a&#8217; e&#8217; i&#8217; o&#8217; u&#8217;&gt;, transparently relating them their non-rhotic counterparts &lt;a e i o u&gt; in Cooper&#8217;s (<xref ref-type="bibr" rid="B14">2005</xref>) orthography. This may encourage association between historically related vowel pairs among speakers who read and write using this orthography. We have not collected metalinguistic judgments about associations between rhotic and non-rhotic vowels, and we are not aware of any phonological patterns supporting phonological relationships between particular vowels (see <xref ref-type="bibr" rid="B38">Kochetov et al. 2021</xref>), but we note one similarity (lip rounding) which may support the continued connection between rhotic vowels and their historically related non-rhotic vowels. Rhotic vowels are produced with virtually identical lip postures to their non-rhotic counterparts (<xref ref-type="bibr" rid="B33">Hussain &amp; Mielke, 2021</xref>). Even though lip rounding would be an excellent way to help achieve the low F3 targets for these vowels, speakers maintain lip postures that match the vowels&#8217; historical non-rhotic counterparts.</p>
<p>Walker and Proctor (<xref ref-type="bibr" rid="B79">2019</xref>) show that American English /&#593;&#633;/ and /o&#633;/ involve little movement of the posterior tongue compared to /i&#633;/ and /e&#633;/, in which the tongue retracts considerably going from the vowel to the rhotic. They interpret this as a reason why /&#593;&#633;/ and /o&#633;/ appear to only have a single mora each and may be followed by coda consonants as in <italic>dark</italic> and <italic>fork</italic>, whereas other vowel+/&#633;/ sequences appear to be bimoraic and cannot be followed by coda consonants. Walker and Proctor (<xref ref-type="bibr" rid="B79">2019</xref>) suggest that /&#593;&#633;/ and /o&#633;/ might be particularly well-suited to being monophthongal rhotic vowels resembling coarticulated sequences of plain vowels and rhotic approximants. However, we have seen that the anterior tongue position is quite different between rhotic and non-rhotic /a/ and /o/ in Kalasha, and the rhotic vowels are rather different from what is predicted based on coarticulation. In summary, what makes /&#593;&#633;/ and /o&#633;/ particularly good sequences in English does not extend well to monophthongal rhotic vowels with a single tongue-front posture throughout the vowel.</p>
</sec>
</sec>
<sec>
<title>6. Conclusion</title>
<p>Despite the historical source of rhotic vowels being <italic>retroflex</italic>, Kalasha rhotic vowels are (at least predominantly) <italic>bunched</italic>. To explore what kind of vowels would result from coarticulation between Kalasha&#8217;s plain vowels and a source of rhoticity, we used biomechanical modeling to superimpose the gestures and acoustic modeling to examine the resulting formant frequencies. We found that present-day Kalasha rhotic vowels are not simply plain vowels with retroflexion or bunching added. The introduction of rhoticity has caused a reorganization of the vowel space, most notably in the form of the fronting of back vowels, especially /o&#734;/, which is much more acoustically distinct from /o/ than would be predicted based on coarticulation alone. We did not observe wholesale dispersion of Kalasha&#8217;s ten oral vowels, but we observed that /o&#734;/, the rhotic vowel most vulnerable to merger in its apparent original coarticulated form, is the vowel that differs the most from what would be expected based on coarticulation. It has shifted to being a front rounded vowel, potentially avoiding the mergers that affected pre-rhotic /u/ and /o/ in North American English.</p>
<p>We found some converging evidence that ideal rhotic vowels are 70% rhotic and 30% vowel in terms of muscle activation. The best balance of acoustic dispersions within the rhotic vowel subsystem and within the pairs of rhotic and non-rhotic vowels happens at rhotic mix proportions of 0.7 and 0.8. Also, the rhotic vowels whose forward simulations best matched their inverse simulations matched the best at rhotic mix proportions of 0.6 or 0.7.</p>
<p>The present study shows how a language can exploit already-available consonantal features to transform a simple vowel inventory into a crowded vowel space. Compared to more frequent features such as nasality and length, rhoticity directly interacts with the tongue postures used to achieve vowel quality and necessitates reorganization of the articulatory realization of vowel contrasts. The introduction of rhoticity has transformed the lingual articulation of rhotic vowels, especially /o&#734;/. This involves changes in its position in the vowel space relative to what would be expected from coarticulation, and it also involves new lingual means of achieving the low F3 of rhotic vowels (bunching rather than retroflexion). Through all this, lip rounding stands as a conspicuously unexploited means of lowering F3 in rhotic vowels. Despite the fact that lip rounding could aid in producing low F3 in rhotic vowels that are not already rounded, the rhotic vowels maintain the lip rounding specifications of their corresponding non-rhotic vowels, suggesting that they are still related in the minds of Kalasha speakers.</p>
<p>We end with a brief history of technology and our understanding of rhotic vowels. Trail and Cooper (<xref ref-type="bibr" rid="B78">1985</xref>), Heeg&#229;rd and M&#248;rch (<xref ref-type="bibr" rid="B27">2004</xref>), and Cooper (<xref ref-type="bibr" rid="B14">2005</xref>) described the 20 vowels of modern Kalasha in terms of vowel quality, nasality, and retroflexion, and Heeg&#229;rd and M&#248;rch (<xref ref-type="bibr" rid="B27">2004</xref>) and Di Carlo (<xref ref-type="bibr" rid="B17">2016</xref>) described how they likely originated from coarticulation and contact. Hussain and Mielke (<xref ref-type="bibr" rid="B34">2022</xref>) postulated that the emergence of Kalasha rhotic vowels was probably further amplified by retroflex approximant /&#635;/, which is widely found in neighboring Nuristani languages. Kochetov et al. (<xref ref-type="bibr" rid="B38">2021</xref>) provided an acoustic description of the Kalasha vowel system confirming the low F3 of rhotic vowels and the effects of rhoticity on F1 and F2. Hussain and Mielke (<xref ref-type="bibr" rid="B33">2021</xref>) used ultrasound and lip video data to show that the rhotic vowels are bunched and their lip postures closely match their non-rhotic counterparts, but could only speculate about how the observed tongue postures relate to the apparent origins of rhoticity in the Kalasha vowel system. Present-day Kalasha rhotic vowels probably owe their existence to coarticulation but they are many decades removed from it and even the rawest phonetic data incorporates the phonological consequences of communicating with rhotic vowels for multiple generations of speakers. Biomechanical modeling has enabled us to construct a plausible coarticulation-only simulation of the Kalasha rhotic vowel system to compare with our simulation of the actual rhotic vowels, and learn about how they are different, and find out what happens when a language develops a new vowel feature.</p>
</sec>
<sec>
<title>Supplementary materials</title>
<p>The overall construction of the model made use of pre-existing face (<xref ref-type="bibr" rid="B61">Nazari, Perrier, Chabanas, &amp; Payan, 2010</xref>), tongue (<xref ref-type="bibr" rid="B72">Stavness, Lloyd, Payan, &amp; Fels, 2011</xref>), and rigid structure model components available within ArtiSynth (<xref ref-type="bibr" rid="B1">Anderson et al., 2017</xref>). Collisions are used very sparingly since they can greatly reduce the stability of the model and were not deemed essential here. Collision processing was only used between the lips and the teeth. The design closely (although not exactly) follows the material properties and coupling specified in the original model sources. However, modifications were made to several model components to improve the speed of model development and provide greater model flexibility deemed necessary to achieve the widely varying set of empirical tongue shapes observed in the ultrasound data. The most important changes were converting the face and tongue FEMs to tetrahedral meshes, trimming of the face FEM, and elaborating the musculature of the face and tongue models. Concerning the first change, using meshes dominantly comprised of linear tetrahedral elements is less preferred compared to other types (such as quadratic tetrahedral or linear hexahedral topologies) because of issues such as mesh locking that can arise (<xref ref-type="bibr" rid="B5">Benzley, Perry, Merkley, Clark, &amp; Sjaardama, 1995</xref>; <xref ref-type="bibr" rid="B31">Hughes, 1987</xref>). This issue was not found to be a major impediment to simulation here, but rather the gains in design flexibility and speed of model prototyping were deemed to be worthy trade-offs to make.</p>
<p><xref ref-type="fig" rid="F19">Figure 19</xref> shows original face FEM and the trimmed version used in the present simulations. Concerning trimming of the face FEM, the original face model has an inferosuperior extent from below the chin (at roughly the level of the vocal folds) to above the brow ridge and an anteroposterior extent from the tip of the nose to roughly a coronal plane set just in front of the ears. While this expansive amount of facial structure could allow for simulation of details of facial expression, it is unnecessary for speech models and its presence only adds to the computational burden of the simulation. The trimming, which was conducted in Blender (<ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://www.blender.org">www.blender.org</ext-link>), reduced the face to just the lower orofacial area, with a superior, transverse planar-border set immediately below the nose, an oblique plane oriented roughly parallel to the inferior border of the body of the mandible, and a posterior, roughly coronal planar-border aligned to the posterior margin of the mandibular rami.</p>
<fig id="F19">
<label>Figure 19</label>
<caption>
<p>Original face FEM (top) and the trimmed version (bottom) used in the present simulations.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-9019-g19.jpg"/>
</fig>
<p><xref ref-type="fig" rid="F20">Figures 20</xref> and <xref ref-type="fig" rid="F21">21</xref> show the tongue and face muscles. Elaboration of the tongue and face musculature was performed to increase the density of muscle fiber representation in both models, which are, in the original versions, quite sparse. This was done to improve the distribution of the musculature within a material-based modeling approach (and specification of muscle fiber directions at element integration points). Redesign of the muscles was performed in Blender and a flexible system was developed along with an export script to facilitate the process of making small changes to the musculature and facilitate further refinements in future models. Both sets of musculature are based on available anatomical resources (e.g., <xref ref-type="bibr" rid="B83">Zemlin, 1998</xref>), but the tongue muscles are also closely based upon details presented in Sanders and Mu (<xref ref-type="bibr" rid="B66">2013</xref>), with the images therein having been imported into Blender and used to guide the layout of the muscle fibers. Muscles requiring extrinsic attachment to skeletal structure (outside of the bounding FEMs) were supported by axial muscles partially embedded in the FEM. A tendinous origin of the genioglossus muscle fibers was also developed.</p>
<fig id="F20">
<label>Figure 20</label>
<caption>
<p>Tongue muscles for ArtiSynth modeling.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-9019-g20.jpg"/>
</fig>
<fig id="F21">
<label>Figure 21</label>
<caption>
<p>Refined muscles of the face (top) and tongue (bottom).</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-9019-g21.jpg"/>
</fig>
<p>All inverse simulations used a damping term of 0.5 and an l2-normalization term of 1.0. Target weighting across the different articulators was freely adjusted to facilitate finding a balance between accuracy of the tracking and stability of the simulation (some empirical configurations were difficult to match with the model and we needed to relax the target weighting where this occurred).</p>
</sec>
</body>
<back>
<fn-group>
<fn id="n1"><p>This classification does not distinguish tense/lax from vowel height, and many vowel systems that are analyzed as having tongue root contrasts but transcribed using base symbols are also treated as basically vowel height. The tongue root advancement or retraction category here only includes languages that are listed with diacritics indicating tongue root position.</p></fn>
<fn id="n2"><p>There is some ambiguity in how the rhotic vowels differ in F1. Kochetov et al. (<xref ref-type="bibr" rid="B38">2021</xref>) found centralized F1 for all of the rhotic vowels, and Hussain &amp; Mielke (<xref ref-type="bibr" rid="B33">2021</xref>) found similar differences in nonlow vowels but little difference in F1 between /a/ and /a&#734;/, and in the data used for this paper, /a&#734;/ has higher F1 than /a/. We note also that we have observed signs of larynx raising in Kalasha rhotic vowels, which could account for general F1 raising in some or all speakers. This is discussed more in the discussion.</p></fn>
<fn id="n3"><p>As a precaution to prevent collapse of the area function (and thus zero acoustic output), we clamp the minimum cross-sectional area of the airway skin to a value of 1.0 <italic>&#215;</italic> 10<sup>&#8211;5</sup> m<sup>2</sup> (0.1 cm<sup>2</sup>).</p></fn>
<fn id="n4"><p>It should be noted that articulatory-target timing is introduced into the simulation to allow the model to attain the target posture from its resting position at a pace commensurate with what might be observed in speech (200 ms for the inverse simulation) while also providing ample timing for stability purposes. However, the simulations are idealized in all other respects with respect to timing and should be viewed as canonical steady-state productions (i.e., free of context and hence coarticulatory influence on muscle activity patterns). We did not make any attempt to simulate muscle activation dynamics that might arise from effects such as coarticulation, which are empirically documented for speech (<xref ref-type="bibr" rid="B43">Leidner, 1976</xref>; <xref ref-type="bibr" rid="B51">MacNeilage &amp; DeClerk, 1969</xref>; <xref ref-type="bibr" rid="B74">Sussman, MacNeilage, &amp; Hanson, 1973</xref>) and other complex motor-control tasks (e.g., <xref ref-type="bibr" rid="B81">Winges, Furuya, Faber, and Flanders 2013</xref>).</p></fn>
<fn id="n5"><p>In our earlier exploration of the methods for performing the blending simulations, we attempted to use idealized rhotics as targets (based in part on the work of <xref ref-type="bibr" rid="B70">Stavness, Gick, et al., 2012</xref>). We abandoned this approach in favor of one that makes use of exemplars drawn from our ultrasound data of the Kamviri-style rhotic approximants to improve the connection of our simulation materials to our research question. This decision, however, comes with the drawback that the use of natural speech data involving real-word productions makes coarticulatory contamination unavoidable. We opted to use the rhotic low-vowel context because this is the only context where we have articulatory data from the relevant languages (<xref ref-type="bibr" rid="B34">Hussain and Mielke, 2022</xref>). Thus our findings must be interpreted with this in mind.</p></fn>
<fn id="n6"><p>Note that this is unlikely to be accounted for bunched/retroflex allophony in rhotic approximants. Where vowel-conditioned bunched/retroflex allophony has been observed in English, the distribution of bunched and retroflex allophones is the opposite of what would account for this pattern: Retroflexion is compatible with back vowels and bunching is compatible with front vowels (<xref ref-type="bibr" rid="B56">Mielke et al., 2016</xref>; <xref ref-type="bibr" rid="B63">Ong &amp; Stone, 1998</xref>; <xref ref-type="bibr" rid="B70">Stavness, Gick, et al., 2012</xref>)</p></fn>
<fn id="n7"><p>It also has the ten nasalized counterparts of these vowels, which are not addressed here.</p></fn>
<fn id="n8"><p>We note that Liu and Kewley-Port (<xref ref-type="bibr" rid="B48">2004</xref>) compared pairs of vowels differing in one formant, and here we are measuring Euclidean distance based on three formants.</p></fn>
</fn-group>
<ack>
<title>Acknowledgements</title>
<p>We wish to thank the developers of ArtiSynth for their support, including Sid Fels, John Lloyd, Ian Stavness, Bryan Gick, and many more. We thank the attendees of LabPhon 2020 for comments and suggestions, and Erik Thomas for guidance with English prerhotic vowels.</p>
</ack>
<sec>
<title>Funding information</title>
<p>This project was funded by a Documenting Endangered Languages grant (BCS-1562134) from the National Science Foundation and the NCSU Department of English.</p>
</sec>
<sec>
<title>Competing interests</title>
<p>The authors have no competing interests to declare.</p>
</sec>
<sec>
<title>Authors&#8217; contributions</title>
<p><bold>JM:</bold> Conceptualization, Investigation, Methodology, Data analysis and visualization, Writing</p>
<p><bold>QH:</bold> Conceptualization, Investigation, Methodology, Data collection and visualization, Writing</p>
<p><bold>SRM:</bold> Investigation, Methodology, Data analysis and visualization, Writing</p>
</sec>
<ref-list>
<ref id="B1"><label>1</label><mixed-citation publication-type="book"><string-name><surname>Anderson</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Fels</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Harandi</surname>, <given-names>N. M.</given-names></string-name>, <string-name><surname>Ho</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Moisik</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>S&#225;nchez</surname>, <given-names>C. A.</given-names></string-name>, &#8230; <string-name><surname>Tang</surname>, <given-names>K.</given-names></string-name> (<year>2017</year>). <chapter-title>Frank: A hybrid 3d biomechanical model of the head and neck</chapter-title>. In <source>Biomechanics of living organs</source> (pp. <fpage>413</fpage>&#8211;<lpage>447</lpage>). <publisher-name>Academic Press</publisher-name>. DOI: <pub-id pub-id-type="doi">10.1016/B978-0-12-804009-6.00020-1</pub-id></mixed-citation></ref>
<ref id="B2"><label>2</label><mixed-citation publication-type="webpage"><string-name><surname>Baker</surname>, <given-names>A.</given-names></string-name> (<year>2005</year>). <chapter-title>Palatoglossatron 1.0 [Computer software manual]</chapter-title>. <publisher-loc>Tucson, Arizona</publisher-loc>. (<uri>http://dingo.sbs.arizona.edu/&#8764;apilab/pdfs/pgman.pdf</uri>)</mixed-citation></ref>
<ref id="B3"><label>3</label><mixed-citation publication-type="book"><string-name><surname>Barlaz</surname>, <given-names>M. S.</given-names></string-name>, <string-name><surname>Fu</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Dubin</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Liang</surname>, <given-names>Z.-P.</given-names></string-name>, <string-name><surname>Shosted</surname>, <given-names>R.</given-names></string-name>, &amp; <string-name><surname>Sutton</surname>, <given-names>B. P.</given-names></string-name> (<year>2015</year>). <chapter-title>Lingual differences in Brazilian Portuguese oral and nasal vowels: An MRI study</chapter-title>. In <source>Proceedings of the 18th International Congress of Phonetic Sciences</source> (pp. <fpage>1</fpage>&#8211;<lpage>5</lpage>). <publisher-loc>Glasgow, UK</publisher-loc>: <publisher-name>University of Glasgow</publisher-name>. (Paper number 819)</mixed-citation></ref>
<ref id="B4"><label>4</label><mixed-citation publication-type="thesis"><string-name><surname>Beddor</surname>, <given-names>P. S.</given-names></string-name> (<year>1982</year>). <source>Phonological and phonetic effects of nasalization on vowel height</source> (Unpublished doctoral dissertation). <publisher-name>University of Minnesota</publisher-name>. (Bloomington, IN: Indiana University Linguistics Club).</mixed-citation></ref>
<ref id="B5"><label>5</label><mixed-citation publication-type="journal"><string-name><surname>Benzley</surname>, <given-names>S. E.</given-names></string-name>, <string-name><surname>Perry</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Merkley</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Clark</surname>, <given-names>B.</given-names></string-name>, &amp; <string-name><surname>Sjaardama</surname>, <given-names>G.</given-names></string-name> (<year>1995</year>). <article-title>A comparison of all hexagonal and all tetrahedral finite element meshes for elastic and elasto-plastic analysis</article-title>. In <source>Proceedings of the 4th international meshing roundtable</source> (pp. <fpage>179</fpage>&#8211;<lpage>191</lpage>).</mixed-citation></ref>
<ref id="B6"><label>6</label><mixed-citation publication-type="thesis"><string-name><surname>Birkholz</surname>, <given-names>P.</given-names></string-name> (<year>2005</year>). <source>3d-artikulatorische sprachsynthese</source> (Unpublished doctoral dissertation). <publisher-name>Universit&#228;t Rostock</publisher-name>, <publisher-loc>Germany</publisher-loc>.</mixed-citation></ref>
<ref id="B7"><label>7</label><mixed-citation publication-type="webpage"><string-name><surname>Birkholz</surname>, <given-names>P.</given-names></string-name>, &amp; <string-name><surname>Jackel</surname>, <given-names>D.</given-names></string-name> (<year>2004</year>). <article-title>Influence of temporal discretization schemes on formant frequencies and bandwidths in time domain simulations of the vocal tract system</article-title>. In <source>Proceedings of Interspeech 2004</source> (pp. <fpage>1125</fpage>&#8211;<lpage>1128</lpage>). DOI: <pub-id pub-id-type="doi">10.21437/Interspeech.2004-409</pub-id></mixed-citation></ref>
<ref id="B8"><label>8</label><mixed-citation publication-type="webpage"><string-name><surname>Boersma</surname>, <given-names>P.</given-names></string-name>, &amp; <string-name><surname>Weenink</surname>, <given-names>D.</given-names></string-name> (<year>2007</year>). <article-title>Praat: Doing phonetics by computer [Computer program]</article-title>. (Version 6.0.30, <uri>http://www.praat.org</uri>)</mixed-citation></ref>
<ref id="B9"><label>9</label><mixed-citation publication-type="book"><string-name><surname>Brunner</surname>, <given-names>J.</given-names></string-name>, &amp; <string-name><surname>Zygis</surname>, <given-names>M.</given-names></string-name> (<year>2011</year>). <chapter-title>Why do glottal stops and low vowels like each other?</chapter-title> In <source>Proceedings of the 17th International Congress of Phonetic Sciences</source> (pp. <fpage>376</fpage>&#8211;<lpage>379</lpage>). <publisher-name>City University of Hong Kong</publisher-name>, <publisher-loc>Hong Kong</publisher-loc>.</mixed-citation></ref>
<ref id="B10"><label>10</label><mixed-citation publication-type="journal"><string-name><surname>Carignan</surname>, <given-names>C.</given-names></string-name> (<year>2014</year>). <article-title>An acoustic and articulatory examination of the &#8220;oral&#8221; in &#8220;nasal&#8221;: The oral articulations of French nasal vowels are not arbitrary</article-title>. <source>Journal of Phonetics</source>, <volume>46</volume>, <fpage>23</fpage>&#8211;<lpage>33</lpage>. DOI: <pub-id pub-id-type="doi">10.1016/j.wocn.2014.05.001</pub-id></mixed-citation></ref>
<ref id="B11"><label>11</label><mixed-citation publication-type="journal"><string-name><surname>Carignan</surname>, <given-names>C.</given-names></string-name> (<year>2018</year>). <article-title>Using ultrasound and nasalance to separate oral and nasal contributions to formant frequencies of nasalized vowels</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>143</volume>(<issue>5</issue>), <fpage>2588</fpage>&#8211;<lpage>2601</lpage>. DOI: <pub-id pub-id-type="doi">10.1121/1.5034760</pub-id></mixed-citation></ref>
<ref id="B12"><label>12</label><mixed-citation publication-type="journal"><string-name><surname>Carignan</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Shosted</surname>, <given-names>R. K.</given-names></string-name>, <string-name><surname>Fu</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Liang</surname>, <given-names>Z.-P.</given-names></string-name>, &amp; <string-name><surname>Sutton</surname>, <given-names>B. P.</given-names></string-name> (<year>2015</year>). <article-title>A real-time MRI investigation of the role of lingual and pharyngeal articulation in the production of the nasal vowel system of French</article-title>. <source>Journal of Phonetics</source>, <volume>50</volume>, <fpage>34</fpage>&#8211;<lpage>51</lpage>. DOI: <pub-id pub-id-type="doi">10.1016/j.wocn.2015.01.001</pub-id></mixed-citation></ref>
<ref id="B13"><label>13</label><mixed-citation publication-type="book"><string-name><surname>Catford</surname>, <given-names>J. C.</given-names></string-name> (<year>1983</year>). <chapter-title>Pharyngeal and laryngeal sounds in Caucasian languages</chapter-title>. In <string-name><given-names>D. M.</given-names> <surname>Bless</surname></string-name> &amp; <string-name><given-names>J. H.</given-names> <surname>Abbs</surname></string-name> (Eds.), (pp. <fpage>344</fpage>&#8211;<lpage>350</lpage>). <publisher-name>College-Hill Press</publisher-name>.</mixed-citation></ref>
<ref id="B14"><label>14</label><mixed-citation publication-type="thesis"><string-name><surname>Cooper</surname>, <given-names>G.</given-names></string-name> (<year>2005</year>). <source>Issues in the development of a writing system for the Kalasha language</source> (Unpublished doctoral dissertation). <publisher-name>Macquarie University</publisher-name>, <publisher-loc>Sydney</publisher-loc>.</mixed-citation></ref>
<ref id="B15"><label>15</label><mixed-citation publication-type="journal"><string-name><surname>de Boer</surname>, <given-names>B.</given-names></string-name> (<year>2000</year>). <article-title>Self-organization in vowel systems</article-title>. <source>Journal of phonetics</source>, <volume>28</volume>(<issue>4</issue>), <fpage>441</fpage>&#8211;<lpage>465</lpage>. DOI: <pub-id pub-id-type="doi">10.1006/jpho.2000.0125</pub-id></mixed-citation></ref>
<ref id="B16"><label>16</label><mixed-citation publication-type="journal"><string-name><surname>Delattre</surname>, <given-names>P.</given-names></string-name>, &amp; <string-name><surname>Freeman</surname>, <given-names>D. C.</given-names></string-name> (<year>1968</year>). <article-title>A dialect study of American r&#8217;s by x-ray motion picture</article-title>. <source>Linguistics</source>, <volume>6</volume>(<issue>44</issue>), <fpage>29</fpage>&#8211;<lpage>68</lpage>. DOI: <pub-id pub-id-type="doi">10.1515/ling.1968.6.44.29</pub-id></mixed-citation></ref>
<ref id="B17"><label>17</label><mixed-citation publication-type="journal"><string-name><surname>Di Carlo</surname>, <given-names>P.</given-names></string-name> (<year>2016</year>). <article-title>Retroflex vowels? phonetics, phonology, and history of unusual sounds in Kalasha and other languages of the Hindu Kush region</article-title>. <source>Archivio per L&#8217;Antropologia e la Etnologia, CXLVI</source>, <fpage>103</fpage>&#8211;<lpage>121</lpage>.</mixed-citation></ref>
<ref id="B18"><label>18</label><mixed-citation publication-type="book"><string-name><surname>Diehl</surname>, <given-names>R. L.</given-names></string-name>, <string-name><surname>Kluender</surname>, <given-names>K. R.</given-names></string-name>, <string-name><surname>Walsh</surname>, <given-names>M. A.</given-names></string-name>, &amp; <string-name><surname>Parker</surname>, <given-names>E. M.</given-names></string-name> (<year>1991</year>). <chapter-title>Auditory enhancement in speech perception and phonology</chapter-title>. In <string-name><given-names>R. R.</given-names> <surname>Hoffman</surname></string-name> &amp; <string-name><given-names>D. S.</given-names> <surname>Palermo</surname></string-name> (Eds.), <source>Cognition and the symbolic processes, vol 3: Applied and ecological perspectives</source> (pp. <fpage>59</fpage>&#8211;<lpage>76</lpage>). <publisher-loc>Hillsdale, NJ</publisher-loc>: <publisher-name>Erlbaum</publisher-name>.</mixed-citation></ref>
<ref id="B19"><label>19</label><mixed-citation publication-type="journal"><string-name><surname>Eskes</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Balm</surname>, <given-names>A. J.</given-names></string-name>, Van <string-name><surname>Alphen</surname>, <given-names>M. J.</given-names></string-name>, <string-name><surname>Smeele</surname>, <given-names>L. E.</given-names></string-name>, <string-name><surname>Stavness</surname>, <given-names>I.</given-names></string-name>, &amp; <string-name><surname>Van Der Heijden</surname>, <given-names>F.</given-names></string-name> (<year>2017</year>). <article-title>sEMG-assisted inverse modelling of 3D lip movement: a feasibility study towards person-specific modelling</article-title>. <source>Scientific Reports</source>, <volume>7</volume>(<issue>1</issue>), <fpage>1</fpage>&#8211;<lpage>14</lpage>. DOI: <pub-id pub-id-type="doi">10.1038/s41598-017-17790-4</pub-id></mixed-citation></ref>
<ref id="B20"><label>20</label><mixed-citation publication-type="journal"><string-name><surname>Esposito</surname>, <given-names>C. M.</given-names></string-name>, <string-name><surname>Sleeper</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Sch&#228;fer</surname>, <given-names>K.</given-names></string-name> (<year>2021</year>). <article-title>Examining the relationship between vowel quality and voice quality</article-title>. <source>Journal of the International Phonetic Association</source>, <volume>51</volume>(<issue>3</issue>), <fpage>361</fpage>&#8211;<lpage>392</lpage>. DOI: <pub-id pub-id-type="doi">10.1017/S0025100319000094</pub-id></mixed-citation></ref>
<ref id="B21"><label>21</label><mixed-citation publication-type="journal"><string-name><surname>Feng</surname>, <given-names>G.</given-names></string-name>, &amp; <string-name><surname>Castelli</surname>, <given-names>E.</given-names></string-name> (<year>1996</year>). <article-title>Some acoustic features of nasal and nasalized vowels: A target for vowel nasalization</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>99</volume>(<issue>6</issue>), <fpage>3694</fpage>&#8211;<lpage>3706</lpage>. DOI: <pub-id pub-id-type="doi">10.1121/1.414967</pub-id></mixed-citation></ref>
<ref id="B22"><label>22</label><mixed-citation publication-type="book"><string-name><surname>Flemming</surname>, <given-names>E. S.</given-names></string-name> (<year>2002</year>). <source>Auditory representations in phonology</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>Routledge</publisher-name>.</mixed-citation></ref>
<ref id="B23"><label>23</label><mixed-citation publication-type="journal"><string-name><surname>Fujimura</surname>, <given-names>O.</given-names></string-name>, &amp; <string-name><surname>Lindqvist</surname>, <given-names>J.</given-names></string-name> (<year>1971</year>). <article-title>Sweep-tone measurements of vocal-tract characteristics</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>49</volume>, <fpage>541</fpage>&#8211;<lpage>558</lpage>. DOI: <pub-id pub-id-type="doi">10.1121/1.1912385</pub-id></mixed-citation></ref>
<ref id="B24"><label>24</label><mixed-citation publication-type="book"><string-name><surname>Gick</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Wilson</surname>, <given-names>I.</given-names></string-name>, &amp; <string-name><surname>Derrick</surname>, <given-names>D.</given-names></string-name> (<year>2013</year>). <source>Articulatory phonetics</source>. <publisher-loc>Malden, MA</publisher-loc>: <publisher-name>Wiley-Blackwell</publisher-name>.</mixed-citation></ref>
<ref id="B25"><label>25</label><mixed-citation publication-type="book"><string-name><surname>Gussenhoven</surname>, <given-names>C.</given-names></string-name> (<year>2007</year>). <chapter-title>A vowel height split explained: Compensatory listening and speaker control</chapter-title>. In <string-name><given-names>J.</given-names> <surname>Cole</surname></string-name> &amp; <string-name><given-names>J.</given-names> <surname>Hualde</surname></string-name> (Eds.), <source>Papers in laboratory phonology 9: Change in phonology</source> (pp. <fpage>145</fpage>&#8211;<lpage>172</lpage>). <publisher-loc>Berlin</publisher-loc>: <publisher-name>Mouton de Gruyter</publisher-name>.</mixed-citation></ref>
<ref id="B26"><label>26</label><mixed-citation publication-type="book"><string-name><surname>Hamann</surname>, <given-names>S.</given-names></string-name> (<year>2003</year>). <source>The phonetics and phonology of retroflexes</source>. <publisher-loc>Utrecht, The Netherlands</publisher-loc>: <publisher-name>LOT</publisher-name>.</mixed-citation></ref>
<ref id="B27"><label>27</label><mixed-citation publication-type="book"><string-name><surname>Heeg&#229;rd</surname>, <given-names>J.</given-names></string-name>, &amp; <string-name><surname>M&#248;rch</surname>, <given-names>I. E.</given-names></string-name> (<year>2004</year>). <chapter-title>Retroflex vowels and other peculiarities in the Kalasha sound system</chapter-title>. In <string-name><given-names>A.</given-names> <surname>Saxena</surname></string-name> (Ed.), <source>Himalayan Languages: Past and Present</source> (pp. <fpage>57</fpage>&#8211;<lpage>76</lpage>). <publisher-loc>Berlin</publisher-loc>: <publisher-name>De Gruyter</publisher-name>.</mixed-citation></ref>
<ref id="B28"><label>28</label><mixed-citation publication-type="journal"><string-name><surname>Honda</surname>, <given-names>K.</given-names></string-name> (<year>1996</year>). <article-title>Organization of tongue articulation for vowels</article-title>. <source>Journal of Phonetics</source>, <volume>24</volume>(<issue>1</issue>), <fpage>39</fpage>&#8211;<lpage>52</lpage>. DOI: <pub-id pub-id-type="doi">10.1006/jpho.1996.0004</pub-id></mixed-citation></ref>
<ref id="B29"><label>29</label><mixed-citation publication-type="journal"><string-name><surname>Howson</surname>, <given-names>P. J.</given-names></string-name>, <string-name><surname>Moisik</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name><surname>&#379;ygis</surname>, <given-names>M.</given-names></string-name> (<year>2022</year>). <article-title>Lateral vocalization in Brazilian Portuguese</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>152</volume>(<issue>1</issue>), <fpage>281</fpage>&#8211;<lpage>294</lpage>. DOI: <pub-id pub-id-type="doi">10.1121/10.0012186</pub-id></mixed-citation></ref>
<ref id="B30"><label>30</label><mixed-citation publication-type="book"><string-name><surname>Hueber</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Chollet</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Denby</surname>, <given-names>B.</given-names></string-name>, &amp; <string-name><surname>Stone</surname>, <given-names>M.</given-names></string-name> (<year>2008</year>). <chapter-title>Acquisition of ultrasound, video and acoustic speech data for a silent-speech interface application</chapter-title>. In <source>Proceedings of the Eighth International Seminar on Speech Production</source> (pp. <fpage>365</fpage>&#8211;<lpage>369</lpage>). <publisher-name>Strasbourg</publisher-name>, <publisher-loc>France</publisher-loc>.</mixed-citation></ref>
<ref id="B31"><label>31</label><mixed-citation publication-type="book"><string-name><surname>Hughes</surname>, <given-names>T. J.</given-names></string-name> (<year>1987</year>). <source>The finite element method: Linear static and dynamic finite element analysis</source>. <publisher-loc>USA</publisher-loc>: <publisher-name>Prentice Hall</publisher-name>.</mixed-citation></ref>
<ref id="B32"><label>32</label><mixed-citation publication-type="journal"><string-name><surname>Hussain</surname>, <given-names>Q.</given-names></string-name>, &amp; <string-name><surname>Mielke</surname>, <given-names>J.</given-names></string-name> (<year>2020</year>). <article-title>Kalasha (Pakistan) &#8211; Language Snapshot</article-title>. <source>Language Documentation and Description</source>, <volume>17</volume>, <fpage>66</fpage>&#8211;<lpage>75</lpage>.</mixed-citation></ref>
<ref id="B33"><label>33</label><mixed-citation publication-type="journal"><string-name><surname>Hussain</surname>, <given-names>Q.</given-names></string-name>, &amp; <string-name><surname>Mielke</surname>, <given-names>J.</given-names></string-name> (<year>2021</year>). <article-title>An acoustic and articulatory study of rhotic and rhoticnasal vowels of Kalasha</article-title>. <source>Journal of Phonetics</source>, <volume>87</volume>, <fpage>1</fpage>&#8211;<lpage>45</lpage>. DOI: <pub-id pub-id-type="doi">10.1016/j.wocn.2020.101028</pub-id></mixed-citation></ref>
<ref id="B34"><label>34</label><mixed-citation publication-type="journal"><string-name><surname>Hussain</surname>, <given-names>Q.</given-names></string-name>, &amp; <string-name><surname>Mielke</surname>, <given-names>J.</given-names></string-name> (<year>2022</year>). <article-title>The emergence of bunched vowels from retroflex approximants in endangered Dardic languages</article-title>. <source>Linguistics Vanguard</source>, <volume>8</volume>(<issue>s5</issue>), <fpage>597</fpage>&#8211;<lpage>610</lpage>. DOI: <pub-id pub-id-type="doi">10.1515/lingvan-2021-0022</pub-id></mixed-citation></ref>
<ref id="B35"><label>35</label><mixed-citation publication-type="journal"><string-name><surname>Jackson</surname>, <given-names>M. T.-T.</given-names></string-name>, &amp; <string-name><surname>McGowan</surname>, <given-names>R. S.</given-names></string-name> (<year>2012</year>). <article-title>A study of high front vowels with articulatory data and acoustic simulations</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>131</volume>(<issue>4</issue>), <fpage>3017</fpage>&#8211;<lpage>3035</lpage>. DOI: <pub-id pub-id-type="doi">10.1121/1.3692246</pub-id></mixed-citation></ref>
<ref id="B36"><label>36</label><mixed-citation publication-type="journal"><string-name><surname>Jang</surname>, <given-names>H.</given-names></string-name> (<year>2022</year>). <article-title>A tutorial on articulatory muscles and ArtiSynth: Tongue and suprahyoid muscles, and 3D tongue model</article-title>. <source>Language and Linguistics Compass</source>, <volume>16</volume>(<issue>3</issue>), <elocation-id>e12447</elocation-id>. DOI: <pub-id pub-id-type="doi">10.1111/lnc3.12447</pub-id></mixed-citation></ref>
<ref id="B37"><label>37</label><mixed-citation publication-type="book"><string-name><surname>Jiang</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Chang</surname>, <given-names>Y.</given-names></string-name>, &amp; <string-name><surname>Hsieh</surname>, <given-names>F.</given-names></string-name> (<year>2019</year>). <chapter-title>An EMA study of Er-suffixation in Northeastern Mandarin monophthongs</chapter-title>. In <string-name><given-names>S.</given-names> <surname>Calhoun</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Escudero</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Tabain</surname></string-name>, &amp; <string-name><given-names>P.</given-names> <surname>Warren</surname></string-name> (Eds.), <source>Proceedings of the 19th International Congress of Phonetic Sciences</source>, (pp. <fpage>2149</fpage>&#8211;<lpage>2153</lpage>). <publisher-loc>Canberra</publisher-loc>: <publisher-name>Australasian Speech Science and Technology Association Inc</publisher-name>.</mixed-citation></ref>
<ref id="B38"><label>38</label><mixed-citation publication-type="journal"><string-name><surname>Kochetov</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Arsenault</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Petersen</surname>, <given-names>J. H.</given-names></string-name>, <string-name><surname>Kalas</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name><surname>Kalash</surname>, <given-names>T. K.</given-names></string-name> (<year>2021</year>). <article-title>Kalasha (Bumburet variety)</article-title>. <source>Journal of the International Phonetic Association</source>, <volume>51</volume>(<issue>3</issue>), <fpage>468</fpage>&#8211;<lpage>489</lpage>. DOI: <pub-id pub-id-type="doi">10.1017/S0025100319000367</pub-id></mixed-citation></ref>
<ref id="B39"><label>39</label><mixed-citation publication-type="journal"><string-name><surname>Krakow</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Beddor</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Goldstein</surname>, <given-names>L.</given-names></string-name>, &amp; <string-name><surname>Fowler</surname>, <given-names>C.</given-names></string-name> (<year>1988</year>). <article-title>Coarticulatory influences on the perceived height of nasal vowels</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>83</volume>, <fpage>1146</fpage>&#8211;<lpage>1158</lpage>. DOI: <pub-id pub-id-type="doi">10.1121/1.396059</pub-id></mixed-citation></ref>
<ref id="B40"><label>40</label><mixed-citation publication-type="journal"><string-name><surname>Kr&#225;msk&#7923;</surname>, <given-names>J.</given-names></string-name> (<year>1939</year>). <article-title>A study in the phonology of Modern Persian</article-title>. <source>Archiv Orient&#225;ln&#237;</source>, <volume>11</volume>(<issue>1</issue>), <fpage>66</fpage>.</mixed-citation></ref>
<ref id="B41"><label>41</label><mixed-citation publication-type="book"><string-name><surname>Laver</surname>, <given-names>J.</given-names></string-name> (<year>1980</year>). <source>The phonetic description of voice quality</source>. <publisher-loc>London</publisher-loc>: <publisher-name>Cambridge Studies in Linguistics</publisher-name>.</mixed-citation></ref>
<ref id="B42"><label>42</label><mixed-citation publication-type="book"><string-name><surname>Lehiste</surname>, <given-names>I.</given-names></string-name> (<year>1962</year>). <source>Acoustical characteristics of selected English consonants</source>. <publisher-loc>Ann Arbor</publisher-loc>: <publisher-name>The University of Michigan Communication Sciences Laboratory</publisher-name>.</mixed-citation></ref>
<ref id="B43"><label>43</label><mixed-citation publication-type="journal"><string-name><surname>Leidner</surname>, <given-names>D. R.</given-names></string-name> (<year>1976</year>). <article-title>The articulation of American English /l/: A study of gestural synergy and antagonism</article-title>. <source>Journal of Phonetics</source>, <volume>4</volume>(<issue>4</issue>), <fpage>327</fpage>&#8211;<lpage>335</lpage>. DOI: <pub-id pub-id-type="doi">10.1016/S0095-4470(19)31259-8</pub-id></mixed-citation></ref>
<ref id="B44"><label>44</label><mixed-citation publication-type="journal"><string-name><surname>Leitner</surname>, <given-names>G. W.</given-names></string-name> (<year>1880</year>). <article-title>A sketch of the Bashgali Kafirs and of their language</article-title>. <source>Journal of the United Service Institution of India</source>, <volume>IX</volume>(<issue>43</issue>), <fpage>143</fpage>&#8211;<lpage>190</lpage>.</mixed-citation></ref>
<ref id="B45"><label>45</label><mixed-citation publication-type="journal"><string-name><surname>Lindblom</surname>, <given-names>B.</given-names></string-name> (<year>1963</year>). <article-title>Spectrographic study of vowel reduction</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>35</volume>(<issue>11</issue>), <fpage>1773</fpage>&#8211;<lpage>1781</lpage>. DOI: <pub-id pub-id-type="doi">10.1121/1.1918816</pub-id></mixed-citation></ref>
<ref id="B46"><label>46</label><mixed-citation publication-type="book"><string-name><surname>Lindblom</surname>, <given-names>B.</given-names></string-name> (<year>1986</year>). <chapter-title>Phonetic universals in vowel systems</chapter-title>. In <string-name><given-names>J.</given-names> <surname>Ohala</surname></string-name> &amp; <string-name><given-names>J.</given-names> <surname>Jaeger</surname></string-name> (Eds.), <source>Experimental phonology</source> (pp. <fpage>13</fpage>&#8211;<lpage>44</lpage>). <publisher-loc>New York</publisher-loc>: <publisher-name>Academic Press</publisher-name>.</mixed-citation></ref>
<ref id="B47"><label>47</label><mixed-citation publication-type="book"><string-name><surname>Lindblom</surname>, <given-names>B.</given-names></string-name> (<year>1990</year>). <chapter-title>Explaining phonetic variation: A sketch of the H and H Theory</chapter-title>. In <string-name><given-names>W.</given-names> <surname>Hardcastle</surname></string-name> &amp; <string-name><given-names>A.</given-names> <surname>Marchal</surname></string-name> (Eds.), <source>Speech production and speech modelling</source> (pp. <fpage>403</fpage>&#8211;<lpage>439</lpage>). <publisher-loc>Dordrecht</publisher-loc>: <publisher-name>Kluwer</publisher-name>. DOI: <pub-id pub-id-type="doi">10.1007/978-94-009-2037-8_16</pub-id></mixed-citation></ref>
<ref id="B48"><label>48</label><mixed-citation publication-type="journal"><string-name><surname>Liu</surname>, <given-names>C.</given-names></string-name>, &amp; <string-name><surname>Kewley-Port</surname>, <given-names>D.</given-names></string-name> (<year>2004</year>). <article-title>Vowel formant discrimination for high-fidelity speech</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>116</volume>(<issue>2</issue>), <fpage>1224</fpage>&#8211;<lpage>1233</lpage>. DOI: <pub-id pub-id-type="doi">10.1121/1.1768958</pub-id></mixed-citation></ref>
<ref id="B49"><label>49</label><mixed-citation publication-type="book"><string-name><surname>Lloyd</surname>, <given-names>J. E.</given-names></string-name>, <string-name><surname>Stavness</surname>, <given-names>I.</given-names></string-name>, &amp; <string-name><surname>Fels</surname>, <given-names>S.</given-names></string-name> (<year>2012</year>). <chapter-title>Artisynth: A fast interactive biomechanical modeling toolkit combining multibody and finite element simulation</chapter-title>. In <string-name><given-names>Y.</given-names> <surname>Payan</surname></string-name> (Ed.), <source>Soft tissue biomechanical modeling for computer assisted surgery</source> (pp. <fpage>355</fpage>&#8211;<lpage>394</lpage>). <publisher-loc>Berlin</publisher-loc>: <publisher-name>Springer</publisher-name>. DOI: <pub-id pub-id-type="doi">10.1007/8415_2012_126</pub-id></mixed-citation></ref>
<ref id="B50"><label>50</label><mixed-citation publication-type="journal"><string-name><surname>Lotto</surname>, <given-names>A. J.</given-names></string-name>, <string-name><surname>Holt</surname>, <given-names>L. L.</given-names></string-name>, &amp; <string-name><surname>Kluender</surname>, <given-names>K. R.</given-names></string-name> (<year>1997</year>). <article-title>Effect of voice quality on perceived height of English vowels</article-title>. <source>Phonetica</source>, <volume>54</volume>(<issue>2</issue>), <fpage>76</fpage>&#8211;<lpage>93</lpage>. DOI: <pub-id pub-id-type="doi">10.1159/000262212</pub-id></mixed-citation></ref>
<ref id="B51"><label>51</label><mixed-citation publication-type="journal"><string-name><surname>MacNeilage</surname>, <given-names>P. F.</given-names></string-name>, &amp; <string-name><surname>DeClerk</surname>, <given-names>J. L.</given-names></string-name> (<year>1969</year>). <article-title>On the motor control of coarticulation in CVC monosyllables</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>45</volume>(<issue>5</issue>), <fpage>1217</fpage>&#8211;<lpage>1233</lpage>. DOI: <pub-id pub-id-type="doi">10.1121/1.1911593</pub-id></mixed-citation></ref>
<ref id="B52"><label>52</label><mixed-citation publication-type="book"><string-name><surname>Maddieson</surname>, <given-names>I.</given-names></string-name> (<year>1984</year>). <source>Patterns of sounds</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>. DOI: <pub-id pub-id-type="doi">10.1017/CBO9780511753459</pub-id></mixed-citation></ref>
<ref id="B53"><label>53</label><mixed-citation publication-type="webpage"><collab>MATLAB</collab>. (<year>2019</year>). <source>version 9.7.0.1261785 (r2019b)</source>. <publisher-loc>Natick, Massachusetts</publisher-loc>: <publisher-name>The Math-Works Inc</publisher-name>. Retrieved from <uri>https://www.mathworks.com/products/new_products/release2019b.html</uri></mixed-citation></ref>
<ref id="B54"><label>54</label><mixed-citation publication-type="journal"><string-name><surname>Mielke</surname>, <given-names>J.</given-names></string-name> (<year>2013</year>). <article-title>Ultrasound and corpus study of a change from below: Vowel rhoticity in Canadian French</article-title>. In <source>Penn Working Papers in Linguistics 19.2: Papers from NWAV 41</source> (pp. <fpage>141</fpage>&#8211;<lpage>150</lpage>).</mixed-citation></ref>
<ref id="B55"><label>55</label><mixed-citation publication-type="journal"><string-name><surname>Mielke</surname>, <given-names>J.</given-names></string-name> (<year>2015</year>). <article-title>An ultrasound study of Canadian French rhotic vowels with polar smoothing spline comparisons</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>137</volume>(<issue>5</issue>), <fpage>2858</fpage>&#8211;<lpage>2869</lpage>. DOI: <pub-id pub-id-type="doi">10.1121/1.4919346</pub-id></mixed-citation></ref>
<ref id="B56"><label>56</label><mixed-citation publication-type="journal"><string-name><surname>Mielke</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Baker</surname>, <given-names>A.</given-names></string-name>, &amp; <string-name><surname>Archangeli</surname>, <given-names>D.</given-names></string-name> (<year>2016</year>). <article-title>Individual-level contact limits phonological complexity: Evidence from bunched and retroflex /&#633;/</article-title>. <source>Language</source>, <volume>92</volume>(<issue>1</issue>), <fpage>101</fpage>&#8211;<lpage>140</lpage>. DOI: <pub-id pub-id-type="doi">10.1353/lan.2016.0019</pub-id></mixed-citation></ref>
<ref id="B57"><label>57</label><mixed-citation publication-type="thesis"><string-name><surname>Moisik</surname>, <given-names>S. R.</given-names></string-name> (<year>2013</year>). <source>The epilarynx in speech</source> (Unpublished doctoral dissertation). <publisher-name>University of Victoria</publisher-name>, <publisher-loc>Canada</publisher-loc>.</mixed-citation></ref>
<ref id="B58"><label>58</label><mixed-citation publication-type="webpage"><string-name><surname>Moran</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>McCloy</surname>, <given-names>D.</given-names></string-name>, &amp; <string-name><surname>Wright</surname>, <given-names>R.</given-names></string-name> (<year>2014</year>). <source>PHOIBLE Online</source>. (<publisher-loc>Leipzig</publisher-loc>: <publisher-name>Max Planck Institute for Evolutionary Anthropology</publisher-name>). <uri>http://phoible.org</uri> (accessed on 2023-01-30)</mixed-citation></ref>
<ref id="B59"><label>59</label><mixed-citation publication-type="journal"><string-name><surname>Morgenstierne</surname>, <given-names>G.</given-names></string-name> (<year>1954</year>). <article-title>The Waigali language</article-title>. <source>Norsk Tidsskrift for Sprogvidenskap</source>, <volume>17</volume>, <fpage>146</fpage>&#8211;<lpage>219</lpage>.</mixed-citation></ref>
<ref id="B60"><label>60</label><mixed-citation publication-type="book"><string-name><surname>Morgenstierne</surname>, <given-names>G.</given-names></string-name> (<year>1973</year>). <source>The Kalasha language (Indo-Iranian frontier languages, vol. 4)</source>. <publisher-loc>Oslo</publisher-loc>: <publisher-name>Universitetsforlaget</publisher-name>.</mixed-citation></ref>
<ref id="B61"><label>61</label><mixed-citation publication-type="journal"><string-name><surname>Nazari</surname>, <given-names>M. A.</given-names></string-name>, <string-name><surname>Perrier</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Chabanas</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Payan</surname>, <given-names>Y.</given-names></string-name> (<year>2010</year>). <article-title>Simulation of dynamic orofacial movements using a constitutive law varying with muscle activation</article-title>. <source>Computer Methods in Biomechanics and Biomedical Engineering</source>, <volume>13</volume>(<issue>4</issue>), <fpage>469</fpage>&#8211;<lpage>482</lpage>. DOI: <pub-id pub-id-type="doi">10.1080/10255840903505147</pub-id></mixed-citation></ref>
<ref id="B62"><label>62</label><mixed-citation publication-type="book"><string-name><surname>Nye</surname>, <given-names>G. E.</given-names></string-name> (<year>1955</year>). <source>The phonemes and morphemes of modern Persian: A descriptive study</source>. <publisher-name>University of Michigan</publisher-name>, <publisher-loc>USA</publisher-loc>.</mixed-citation></ref>
<ref id="B63"><label>63</label><mixed-citation publication-type="journal"><string-name><surname>Ong</surname>, <given-names>D.</given-names></string-name>, &amp; <string-name><surname>Stone</surname>, <given-names>M.</given-names></string-name> (<year>1998</year>). <article-title>Three-dimensional vocal tract shapes in /r/ and /l/: A study of MRI, ultrasound, electropalatography, and acoustics</article-title>. <source>Phonoscope</source>, <volume>1</volume>(<issue>1</issue>), <fpage>1</fpage>&#8211;<lpage>13</lpage>.</mixed-citation></ref>
<ref id="B64"><label>64</label><mixed-citation publication-type="journal"><string-name><surname>Padgett</surname>, <given-names>J.</given-names></string-name>, &amp; <string-name><surname>Tabain</surname>, <given-names>M.</given-names></string-name> (<year>2005</year>). <article-title>Adaptive dispersion theory and phonological vowel reduction in Russian</article-title>. <source>Phonetica</source>, <volume>62</volume>(<issue>1</issue>), <fpage>14</fpage>&#8211;<lpage>54</lpage>. DOI: <pub-id pub-id-type="doi">10.1159/000087223</pub-id></mixed-citation></ref>
<ref id="B65"><label>65</label><mixed-citation publication-type="thesis"><string-name><surname>Perder</surname>, <given-names>E.</given-names></string-name> (<year>2013</year>). <source>A grammatical description of Dameli</source> (Unpublished doctoral dissertation). <publisher-name>Stockholm University</publisher-name>, <publisher-loc>Stockholm, Sweden</publisher-loc>.</mixed-citation></ref>
<ref id="B66"><label>66</label><mixed-citation publication-type="journal"><string-name><surname>Sanders</surname>, <given-names>I.</given-names></string-name>, &amp; <string-name><surname>Mu</surname>, <given-names>L.</given-names></string-name> (<year>2013</year>). <article-title>A three-dimensional atlas of human tongue muscles</article-title>. <source>The Anatomical Record</source>, <volume>296</volume>(<issue>7</issue>), <fpage>1102</fpage>&#8211;<lpage>1114</lpage>. DOI: <pub-id pub-id-type="doi">10.1002/ar.22711</pub-id></mixed-citation></ref>
<ref id="B67"><label>67</label><mixed-citation publication-type="book"><string-name><surname>Scobbie</surname>, <given-names>J. M.</given-names></string-name>, <string-name><surname>Wrench</surname>, <given-names>A. A.</given-names></string-name>, &amp; <string-name><surname>van der Linden</surname>, <given-names>M.</given-names></string-name> (<year>2008</year>). <chapter-title>Head-probe stabilisation in ultrasound tongue imaging using a headset to permit natural head movement</chapter-title>. In <source>Proceedings of the 8th International Seminar on Speech Production</source> (pp. <fpage>373</fpage>&#8211;<lpage>376</lpage>). <publisher-name>Strasbourg</publisher-name>, <publisher-loc>France</publisher-loc>.</mixed-citation></ref>
<ref id="B68"><label>68</label><mixed-citation publication-type="journal"><string-name><surname>Serrurier</surname>, <given-names>A.</given-names></string-name>, &amp; <string-name><surname>Badin</surname>, <given-names>P.</given-names></string-name> (<year>2008</year>). <article-title>A three-dimensional articulatory model of the velum and nasopharyngeal wall based on MRI and CT data</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>123</volume>(<issue>4</issue>), <fpage>2335</fpage>&#8211;<lpage>2355</lpage>. DOI: <pub-id pub-id-type="doi">10.1121/1.2875111</pub-id></mixed-citation></ref>
<ref id="B69"><label>69</label><mixed-citation publication-type="journal"><string-name><surname>Smith</surname>, <given-names>K. K.</given-names></string-name>, &amp; <string-name><surname>Kier</surname>, <given-names>W. M.</given-names></string-name> (<year>1989</year>). <article-title>Trunks, tongues, and tentacles: Moving with skeletons of muscle</article-title>. <source>American Scientist</source>, <volume>77</volume>, <fpage>28</fpage>&#8211;<lpage>35</lpage>.</mixed-citation></ref>
<ref id="B70"><label>70</label><mixed-citation publication-type="journal"><string-name><surname>Stavness</surname>, <given-names>I.</given-names></string-name>, <string-name><surname>Gick</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Derrick</surname>, <given-names>D.</given-names></string-name>, &amp; <string-name><surname>Fels</surname>, <given-names>S.</given-names></string-name> (<year>2012</year>). <article-title>Biomechanical modeling of English /r/ variants</article-title>. <source>The Journal of the Acoustical Society of America &#8211; Express Letters</source>, <volume>131</volume>(<issue>5</issue>), <fpage>EL355</fpage>&#8211;<lpage>EL360</lpage>. DOI: <pub-id pub-id-type="doi">10.1121/1.3695407</pub-id></mixed-citation></ref>
<ref id="B71"><label>71</label><mixed-citation publication-type="journal"><string-name><surname>Stavness</surname>, <given-names>I.</given-names></string-name>, <string-name><surname>Lloyd</surname>, <given-names>J. E.</given-names></string-name>, &amp; <string-name><surname>Fels</surname>, <given-names>S.</given-names></string-name> (<year>2012</year>). <article-title>Automatic prediction of tongue muscle activations using a finite element model</article-title>. <source>Journal of Biomechanics</source>, <volume>45</volume>(<issue>16</issue>), <fpage>2841</fpage>&#8211;<lpage>2848</lpage>. DOI: <pub-id pub-id-type="doi">10.1016/j.jbiomech.2012.08.031</pub-id></mixed-citation></ref>
<ref id="B72"><label>72</label><mixed-citation publication-type="journal"><string-name><surname>Stavness</surname>, <given-names>I.</given-names></string-name>, <string-name><surname>Lloyd</surname>, <given-names>J. E.</given-names></string-name>, <string-name><surname>Payan</surname>, <given-names>Y.</given-names></string-name>, &amp; <string-name><surname>Fels</surname>, <given-names>S.</given-names></string-name> (<year>2011</year>). <article-title>Coupled hard&#8211;soft tissue simulation with contact and constraints applied to jaw&#8211;tongue&#8211;hyoid dynamics</article-title>. <source>International Journal for Numerical Methods in Biomedical Engineering</source>, <volume>27</volume>(<issue>3</issue>), <fpage>367</fpage>&#8211;<lpage>390</lpage>. DOI: <pub-id pub-id-type="doi">10.1002/cnm.1423</pub-id></mixed-citation></ref>
<ref id="B73"><label>73</label><mixed-citation publication-type="webpage"><string-name><surname>Strand</surname>, <given-names>R. F.</given-names></string-name> (<year>2011</year>). <source>The sound system of Ni&#353;ei-al&#226;</source>. (<uri>https://nuristan.info/lngFrameL.html</uri>).</mixed-citation></ref>
<ref id="B74"><label>74</label><mixed-citation publication-type="journal"><string-name><surname>Sussman</surname>, <given-names>H. M.</given-names></string-name>, <string-name><surname>MacNeilage</surname>, <given-names>P. F.</given-names></string-name>, &amp; <string-name><surname>Hanson</surname>, <given-names>R. J.</given-names></string-name> (<year>1973</year>). <article-title>Labial and mandibular dynamics during the production of bilabial consonants: Preliminary observations</article-title>. <source>Journal of Speech and Hearing Research</source>, <volume>16</volume>(<issue>3</issue>), <fpage>397</fpage>&#8211;<lpage>420</lpage>. DOI: <pub-id pub-id-type="doi">10.1044/jshr.1603.397</pub-id></mixed-citation></ref>
<ref id="B75"><label>75</label><mixed-citation publication-type="journal"><string-name><surname>Takano</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name><surname>Honda</surname>, <given-names>K.</given-names></string-name> (<year>2007</year>). <article-title>An MRI analysis of the extrinsic tongue muscles during vowel production</article-title>. <source>Speech Communication</source>, <volume>49</volume>(<issue>1</issue>), <fpage>49</fpage>&#8211;<lpage>58</lpage>. DOI: <pub-id pub-id-type="doi">10.1016/j.specom.2006.09.004</pub-id></mixed-citation></ref>
<ref id="B76"><label>76</label><mixed-citation publication-type="book"><string-name><surname>Thomas</surname>, <given-names>E. R.</given-names></string-name> (<year>2001</year>). <source>An acoustic analysis of vowel variation in New World English</source>. <publisher-loc>Durham, N.C.</publisher-loc>: <publisher-name>Duke University Press</publisher-name>. (Publication of the American Dialect Society 85).</mixed-citation></ref>
<ref id="B77"><label>77</label><mixed-citation publication-type="journal"><string-name><surname>Toosarvandani</surname>, <given-names>M. D.</given-names></string-name> (<year>2004</year>). <article-title>Vowel length in modern Farsi</article-title>. <source>Journal of the Royal Asiatic Society</source>, <volume>14</volume>(<issue>3</issue>), <fpage>241</fpage>&#8211;<lpage>251</lpage>. DOI: <pub-id pub-id-type="doi">10.1017/S1356186304004079</pub-id></mixed-citation></ref>
<ref id="B78"><label>78</label><mixed-citation publication-type="book"><string-name><surname>Trail</surname>, <given-names>R.</given-names></string-name>, &amp; <string-name><surname>Cooper</surname>, <given-names>G. R.</given-names></string-name> (<year>1985</year>). <source>Kalasha phonemic summary</source>. (ms., <publisher-name>Summer Institute of Linguistics</publisher-name>)</mixed-citation></ref>
<ref id="B79"><label>79</label><mixed-citation publication-type="journal"><string-name><surname>Walker</surname>, <given-names>R.</given-names></string-name>, &amp; <string-name><surname>Proctor</surname>, <given-names>M.</given-names></string-name> (<year>2019</year>). <article-title>The organisation and structure of rhotics in American English rhymes</article-title>. <source>Phonology</source>, <volume>36</volume>(<issue>3</issue>), <fpage>457</fpage>&#8211;<lpage>495</lpage>. DOI: <pub-id pub-id-type="doi">10.1017/S0952675719000228</pub-id></mixed-citation></ref>
<ref id="B80"><label>80</label><mixed-citation publication-type="book"><string-name><surname>Wells</surname>, <given-names>J. C.</given-names></string-name> (<year>1982</year>). <source>Accents of English: Volume 1</source>. <publisher-name>Cambridge University Press</publisher-name>. DOI: <pub-id pub-id-type="doi">10.1017/CBO9780511611759</pub-id></mixed-citation></ref>
<ref id="B81"><label>81</label><mixed-citation publication-type="journal"><string-name><surname>Winges</surname>, <given-names>S. A.</given-names></string-name>, <string-name><surname>Furuya</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Faber</surname>, <given-names>N. J.</given-names></string-name>, &amp; <string-name><surname>Flanders</surname>, <given-names>M.</given-names></string-name> (<year>2013</year>). <article-title>Patterns of muscle activity for digital coarticulation</article-title>. <source>Journal of Neurophysiology</source>, <volume>110</volume>(<issue>1</issue>), <fpage>230</fpage>&#8211;<lpage>242</lpage>. DOI: <pub-id pub-id-type="doi">10.1152/jn.00973.2012</pub-id></mixed-citation></ref>
<ref id="B82"><label>82</label><mixed-citation publication-type="journal"><string-name><surname>Wood</surname>, <given-names>S.</given-names></string-name> (<year>1979</year>). <article-title>A radiographic analysis of constriction locations for vowels</article-title>. <source>Journal of Phonetics</source>, <volume>7</volume>, <fpage>25</fpage>&#8211;<lpage>43</lpage>. DOI: <pub-id pub-id-type="doi">10.1016/S0095-4470(19)31031-9</pub-id></mixed-citation></ref>
<ref id="B83"><label>83</label><mixed-citation publication-type="book"><string-name><surname>Zemlin</surname>, <given-names>W. R.</given-names></string-name> (<year>1998</year>). <source>Speech and hearing science: Anatomy and physiology</source> (<edition>4th</edition> ed.). <publisher-loc>Boston, MA</publisher-loc>: <publisher-name>Allyn and Bacon</publisher-name>.</mixed-citation></ref>
<ref id="B84"><label>84</label><mixed-citation publication-type="journal"><string-name><surname>Zhou</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Espy-Wilson</surname>, <given-names>C. Y.</given-names></string-name>, <string-name><surname>Boyce</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Tiede</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Holland</surname>, <given-names>C.</given-names></string-name>, &amp; <string-name><surname>Choe</surname>, <given-names>A.</given-names></string-name> (<year>2008</year>). <article-title>A magnetic resonance imaging-based articulatory and acoustic study of &#8220;retroflex&#8221; and &#8220;bunched&#8221; American English /r/</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>123</volume>(<issue>6</issue>), <fpage>4466</fpage>&#8211;<lpage>4481</lpage>. DOI: <pub-id pub-id-type="doi">10.1121/1.2902168</pub-id></mixed-citation></ref>
</ref-list>
</back>
</article>