<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20120330//EN" "http://jats.nlm.nih.gov/publishing/1.2/JATS-journalpublishing1.dtd">
<!--<?xml-stylesheet type="text/xsl" href="Style/article.xsl"?>-->
<article article-type="research-article" dtd-version="1.2" xml:lang="en" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">
<front>
<journal-meta>
<journal-id journal-id-type="issn">1868-6354</journal-id>
<journal-title-group>
<journal-title>Laboratory Phonology: Journal of the Association for Laboratory Phonology</journal-title>
</journal-title-group>
<issn pub-type="epub">1868-6354</issn>
<publisher>
<publisher-name>Open Library of Humanities</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.16995/labphon.18890</article-id>
<article-categories>
<subj-group>
<subject>Journal article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Prosody and predictability in the timing of co-speech gestures: Evidence from Igbo &#8216;gesture shift&#8217;</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-7277-2339</contrib-id>
<name>
<surname>Franich</surname>
<given-names>Kathryn</given-names>
</name>
<email>kfranich@fas.harvard.edu</email>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0009-0009-8165-333X</contrib-id>
<name>
<surname>Nwosu</surname>
<given-names>Vincent</given-names>
</name>
<xref ref-type="aff" rid="aff-2">2</xref>
</contrib>
</contrib-group>
<aff id="aff-1"><label>1</label>Harvard University, Cambridge, MA, USA</aff>
<aff id="aff-2"><label>2</label>University of Calgary, Calgary, Alberta, Canada</aff>
<pub-date publication-format="electronic" date-type="pub" iso-8601-date="2026-06-03">
<day>03</day>
<month>06</month>
<year>2026</year>
</pub-date>
<pub-date pub-type="collection">
<year>2026</year>
</pub-date>
<volume>17</volume>
<issue>1</issue>
<fpage>1</fpage>
<lpage>39</lpage>
<permissions>
<copyright-statement>Copyright: &#x00A9; 2026 The Author(s)</copyright-statement>
<copyright-year>2026</copyright-year>
<license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/4.0/">
<license-p>This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International License (CC-BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. See <uri xlink:href="http://creativecommons.org/licenses/by/4.0/">http://creativecommons.org/licenses/by/4.0/</uri>.</license-p>
</license>
</permissions>
<self-uri xlink:href="http://www.journal-labphon.org/articles/10.16995/labphon.18890/"/>
<abstract>
<p>Though predictability plays an important role in conditioning speech patterns, its role in conditioning co-speech gestures has not been explored. Co-speech gestures are generally dispreferred on prosodically weak and phonetically reduced syllables cross-linguistically. Prosodic phrasing is one factor that can give rise to phonetic reduction; for example, phrase-medial words tend to be shorter than phrase-initial and phrase-final words. Predictability is another factor that can bring about reduction: words and phones that are more predictable are phonetically reduced across languages. In this paper, we examine how prosodic phrasing and predictability interact to shape the timing of co-speech gestures in conversation among four speakers of Igbo, a Niger-Congo language. While gestures tend to co-occur with word-final syllables in the language, they often occur earlier in the word when the word is phrase-medial. We hypothesize that these effects stem from a gradient process of vowel coalescence and reduction that targets word-final syllables, ultimately leading to a restructuring of prosodic prominence and an associated &#8216;shift&#8217; in gesture placement in the word. We test the hypothesis that both prosodic phrasing and predictability contribute to variability in gesture timing due to the similar effects they have on word durations. However, our findings reveal that gesture timing is only indirectly conditioned by durational patterns. Specifically, we demonstrate that certain predictability measures, notably lexical frequency and lexical contextual predictability (the probability of a word given a preceding or following word), do not affect gesture timing, despite influencing durations in Igbo. This suggests that the timing of co-speech gestures is fundamentally governed by prosodic planning mechanisms, rather than being sensitive to surface-level variations in duration. Finally, we find an effect of vowel-to-vowel co-occurrence probability (the probability of a vowel given the preceding or following vowel) on gesture timing, attributable to effects of phonological redundancy on prosodic planning.</p>
</abstract>
</article-meta>
</front>
<body>
<sec>
<title>1. Introduction</title>
<p>The study of multimodal communication has added a new dimension to our understanding of the processes that underlie speech production and planning. Gesture is now acknowledged to be conditioned by prosodic structure in ways similar to speech itself (Esteve-Gibert et al., 2017; <xref ref-type="bibr" rid="B40">Kendon, 1980</xref>, <xref ref-type="bibr" rid="B50">Leonard &amp; Cummins, 2011</xref>; Loehr, 2007; <xref ref-type="bibr" rid="B56">McNeill, 1995</xref>; <xref ref-type="bibr" rid="B66">Rochet-Capellan, 2008</xref>), suggesting that articulation across both modalities is subject to prosodic control (<xref ref-type="bibr" rid="B43">Krivokapi&#263;, 2014</xref>, <xref ref-type="bibr" rid="B44">Krivokapi&#263; et al., 2017</xref>) as well as other factors that condition speech and communicative timing (<xref ref-type="bibr" rid="B8">B&#246;gels &amp; Torreira, 2015</xref>; <xref ref-type="bibr" rid="B28">Garvin &amp; Franich, 2023</xref>; <xref ref-type="bibr" rid="B29">Garvin et al., 2025</xref>; <xref ref-type="bibr" rid="B79">Stoltmann &amp; Fuchs, 2017</xref>). So tightly coupled are speech and gesture in production that timing across the two domains is retained in the context of delayed auditory feedback (<xref ref-type="bibr" rid="B62">Pouw &amp; Dixon, 2019</xref>) and among speakers with cognitive impairments affecting speech production (<xref ref-type="bibr" rid="B35">Jenkins &amp; Pouw, 2023</xref>). Speech movements themselves are also more stable in the context of a co-speech gesture (<xref ref-type="bibr" rid="B29">Garvin et al., 2025</xref>). All this work suggests that the planning and control mechanisms responsible for the timing of speech are dually implicated in the timing of co-speech gesture.</p>
<p>As yet, gesture has received only a fraction of the attention that speech has with respect to the factors that influence its timing. The present paper focuses on interactions between prosodic structure and predictability&#8212;two factors that are known to influence speech production&#8212;on the timing of co-speech gestures within words. The language of study of the present work is Igbo, a Niger-Congo language spoken in Nigeria. In previous work (<xref ref-type="bibr" rid="B24">Franich et al., 2025</xref>), we show that the apexes of manual gestures of Igbo speakers gravitate to the word-final syllable, a position that is metrically prominent in Igbo based on phonological patterning (<xref ref-type="bibr" rid="B13">Clark, 1990</xref>). However, we also find that the timing of co-speech gestures varies as a function of phrase position, with gestures more likely to occur earlier in the word when the word is in medial position of an intonational phrase. We hypothesized that this effect is due to greater susceptibility of phrase-medial words to a gradient process of vowel coalescence in Igbo (<xref ref-type="bibr" rid="B89">Zsiga, 1997</xref>): when word-final syllables are prone to reduction, gestures shift earlier in the word. This hypothesis remains to be directly tested.</p>
<p>Apart from prosodic phrasing, predictability effects are known to influence the fine timing of words and segments (<xref ref-type="bibr" rid="B9">Bybee, 2006</xref>; <xref ref-type="bibr" rid="B20">Ernestus, 2011</xref>; <xref ref-type="bibr" rid="B61">Pierrehumbert, 2001</xref>; <xref ref-type="bibr" rid="B88">Zipf, 1929</xref>). Given that duration reduction occurs to a greater degree in more predictable words (<xref ref-type="bibr" rid="B3">Aylett &amp; Turk, 2004</xref>), the question arises as to whether predictability may be responsible for some of the observed variation in gesture timing in Igbo. In this study, we test the degree to which prosodic phrasing (specifically, the position of a word within an intonational phrase) and predictability a) influence word and segment durations in our Igbo corpus and b) influence the timing of co-speech gestures in our corpus. Our study takes into consideration multiple different measures of predictability on gesture timing, including lexical contextual predictability (the probability that a word will occur given the preceding or following word) and vowel co-occurrence probability (the probability of a vowel given the preceding or following vowel; see Section 1.3 for further detail). Our results provide insights not only on the mechanisms that underlie co-speech gesture timing, but also on the relationship of predictability to prosodic grammar more generally.</p>
<sec>
<title>1.1. Prosodic Structure and Co-Speech Gestures</title>
<p>There is a growing body of cross-linguistic research suggesting that co-speech gestures of the hands, arms, head, and other parts of the body are constrained in their timing according to various aspects of prosodic structure. In languages such as English, German, and Catalan, co-speech gestures are found to align with metrically prominent syllables, particularly stressed syllables that also bear phrase-level pitch accents (<xref ref-type="bibr" rid="B54">Loehr, 2012</xref>; <xref ref-type="bibr" rid="B21">Esteve-Gibert &amp; Prieto, 2013</xref>; <xref ref-type="bibr" rid="B47">K&#252;gler &amp; Gregori 2023</xref>). In lexical tone languages, gestures also co-occur with syllables bearing metrical prominence, though prominence is often cued through very different means compared with stress-based languages. For example, in Med&#649;mba, a Grassfields Bantu language spoken in Cameroon, co-speech gestures are drawn to foot heads, which are cued through a broader range of consonantal and vocalic contrasts than non-heads (<xref ref-type="bibr" rid="B23">Franich &amp; Keupdjio 2022</xref>; Franich et al., 2025). In Igbo, gestures tend to occur on word-final syllables (<xref ref-type="bibr" rid="B24">Franich, 2025</xref>), which have been analyzed as metrically prominent based on patterns of tonal downstep (<xref ref-type="bibr" rid="B13">Clark, 1990</xref>). Taken together, this research demonstrates that there is no universal acoustic or physiological cue (or set of cues) to which gestures consistently align across languages, making pure biomechanical accounts of co-speech gesture timing (e.g., <xref ref-type="bibr" rid="B64">Pouw et al., 2020</xref>; <xref ref-type="bibr" rid="B63">Pouw &amp; Fuchs, 2022</xref>; see also <xref ref-type="bibr" rid="B2">Ambrazaitis &amp; House, 2022</xref>) difficult to maintain.</p>
<p>The timing of co-speech gestures is also sensitive to the presence of prosodic phrase boundaries across multiple languages. For example, in Catalan (<xref ref-type="bibr" rid="B21">Esteve-Gibert &amp; Prieto, 2013</xref>) and Med&#649;mba (Franich et al., 2025), findings indicate that gestures occur relatively earlier in a word when the word is followed by a phrase boundary. Esteve-Gibert and Prieto (<xref ref-type="bibr" rid="B21">2013</xref>) compare this behavior to that of pitch accents, which are also known to retract away from a phrase boundary (Silverman &amp; Pierrehumbert, 1990; Prieto et al., 1995). Co-speech gestures are also found to lengthen in the vicinity of a phrase boundary, exhibiting a similar &#8216;clock-slowing&#8217; tendency at the boundary to consonants and vowels (<xref ref-type="bibr" rid="B43">Krivokapi&#263; et al., 2014</xref>; <xref ref-type="bibr" rid="B44">2017</xref>). Recent work has found that the timing of the onset of gesture strokes across languages correlates with the relative timing of cues to phrase-level prominence: in languages like English where focus-related durational lengthening influences both the consonant and vowel of the prosodically-prominent syllable (<xref ref-type="bibr" rid="B12">Cho 2006</xref>), gesture strokes start around the onset of the word. In languages where lengthening primarily influences the consonant and not the vowel, gesture strokes start before the word onset (<xref ref-type="bibr" rid="B24">Franich, 2025</xref>). These findings not only suggest a close relationship between prosodic structure and co-speech gesture timing, but also suggest that co-speech gesture timing is controlled as a part of the speaker&#8217;s prosodic grammar.</p>
</sec>
<sec>
<title>1.2. Variability in the Timing of Igbo Gestures: Prosody, Predictability, or Both?</title>
<p>Prosodic phrasing is known to modulate word and syllable durations cross-linguistically (<xref ref-type="bibr" rid="B85">Turk &amp; Shattuck-Hufnagel, 2007</xref>; <xref ref-type="bibr" rid="B59">Paschen et al., 2022</xref>; <xref ref-type="bibr" rid="B41">Kim et al. 2024</xref>), and we hypothesize that such effects may play a role in the timing of co-speech gestures in Igbo. However, variation in word and segment durations can also arise as a result of <italic>predictability effects</italic>, such as word frequency, lexical contextual probability (the probability of a word preceding or following another word), or phonotactic predictability (the predictability of a phone given a preceding or following phone) (<xref ref-type="bibr" rid="B6">Bell et al., 2003</xref>; <xref ref-type="bibr" rid="B9">Bybee, 2006</xref>; <xref ref-type="bibr" rid="B20">Ernestus, 2011</xref>; <xref ref-type="bibr" rid="B61">Pierrehumbert, 2001</xref>). While predictability has not been investigated explicitly for its relation to co-speech gesture timing, there are reasons to believe that predictability might be relevant for the planning and production of gestures. Predictability is known to influence word and segment durations: on the whole, more predictable words and segments are produced with shorter durations than less predictable words and segments (Section 1.3 for further discussion). Furthermore, and directly relevant to our Igbo data, there is the fact that both word frequency and lexical contextual probability increase the likelihood of coalescence patterns (<xref ref-type="bibr" rid="B45">Krug 1998</xref>).</p>
<p>Exactly how prosody and predictability might each influence co-speech gesture timing depends on the stage at which these phenomena influence the speech planning process. The locus of predictability effects within the speech chain is still a matter of some debate. Earlier work by Bybee &amp; Hopper (2001) proposed that phonetic effects of predictability arise relatively late in the speech planning process as a result of the greater ease of retrieving articulatory plans for more frequently uttered words or sequences. In contrast, Aylett and Turk&#8217;s (<xref ref-type="bibr" rid="B3">2004</xref>) <italic>smooth signal redundancy hypothesis</italic> proposes that predictability effects arise at an earlier stage of planning. Specifically, they propose that speakers endeavor to achieve &#8220;robust communication,&#8221; i.e., the successful transmission of relevant phonological contrasts to the listener, while minimizing speaker and listener effort in contexts of language redundancy (where lexical, syntactic, semantic, and pragmatic factors conspire to make words more predictable in context). In this theory (see also <xref ref-type="bibr" rid="B84">Turk, 2010</xref>), phonetic effects of predictability are implemented in a prosodic planning component that optimizes information flow in the speech signal. Later work by Seyfarth (<xref ref-type="bibr" rid="B73">2014</xref>) locates phonetic effects of predictability within the speaker&#8217;s exemplar-based lexicon, based on the fact that the average lexical contextual predictability of a word influences its duration even when the same word occurs in more and less predictable contexts. In contrast with this account, Shaw and Tang (<xref ref-type="bibr" rid="B83">2021</xref>) argue that lexical representations are not directly influenced by predictability effects, but that language redundancy (in an Aylett and Turk-style model) influences prosody, which, in turn, can &#8220;leak&#8221; into the memories of words. This is based on their finding that lexical contextual predictability was associated with not only variations in duration, but also variations in fundamental frequency and intensity&#8212;two other correlates of prosodic prominence&#8212;in a corpus of Mandarin speech.</p>
<p>The place of predictability effects in speech planning has implications for the larger architecture of speech planning and production, including planning of co-speech gesture. If durational variations resulting from prosody and predictability arise from the same processing stage, there is greater likelihood that both will influence the timing of co-speech gestures. If, on the other hand, phonetic effects of prosody and predictability are controlled (even to some degree) at separate levels of planning, this leaves room for different effects of prosody and predictability on gesture timing. To take one recent model of speech production, Shaw and Tang (<xref ref-type="bibr" rid="B77">2023</xref>) propose that speech planning proceeds through a selection process in which three components&#8212;the lexicon (with phonetically-detailed exemplars), phonological component, and prosodic component&#8212;serve to influence the chosen production output simultaneously and in parallel. A different approach is taken by Keating &amp; Shattuck-Hufnagel (<xref ref-type="bibr" rid="B39">2002</xref>), who see prosody as occurring at an earlier stage of speech planning relative to word form encoding. This model was designed to capture patterns such as the English &#8216;Rhythm Rule&#8217; (i.e., changes in stress and/or accent location in the context of a clash, as in Japane&#769;se a&#769;pples &#8594; Ja&#769;panese a&#769;pples), which evidences a kind of <italic>prosodic restructuring</italic> of words that happens in response to variables such as prominence and phrasing. The authors propose that speech planning proceeds in stages, such that certain details about a word&#8217;s phonetic pattern are not necessarily present at earlier stages of prosodic planning:</p>
<disp-quote>
<p>&#8220;Some aspects of this prosodic restructuring can be carried out in the absence of word form information; others require at least some information about word form such as number of syllables and stress pattern; and still others require full knowledge of the phonological segments of the words.&#8221; (p. 138)</p>
</disp-quote>
<p>Crucially for the present work, one way of viewing the variability of co-speech gesture timing in Igbo is through the lens of prosodic restructuring. As is discussed in Section 1, the planning of co-speech gestures is intricately connected with the process of prosodic planning. We propose that Igbo speakers, when anticipating prosodically-conditioned reduction processes, will &#8216;shift&#8217; their co-speech gestures earlier in the word, just as speakers are observed to shift speech-based prominence in response to the prosodic context (see Section 1.4 for additional arguments for &#8216;gesture shift&#8217; in Igbo). The model proposed by Keating and Shattuck-Hufnagel (in contrast to the model by <xref ref-type="bibr" rid="B77">Shaw and Tang, 2023</xref>) leaves open the possibility that the fine details of lexical representations&#8212;perhaps including those related to predictability effects&#8212;may not yet be available when prosodic restructuring takes place. As such, the model provides for the possibility that reduction effects linked to prosodic structure will have a different effect on the planning of co-speech gesture timing than reduction effects linked to predictability measures alone.</p>
<p>Aside from the influence of predictability on duration, another path by which co-speech gesture timing could be influenced by predictability effects is through the <italic>attraction</italic> of gestures to less predictable words and segments. Previous work has demonstrated for many languages that co-speech gestures are preferentially timed to words bearing new information in the discourse (<xref ref-type="bibr" rid="B34">Im &amp; Baumann, 2020</xref>; <xref ref-type="bibr" rid="B51">Levy &amp; McNeill, 1992</xref>; <xref ref-type="bibr" rid="B31">Gullberg, 1998</xref>; <xref ref-type="bibr" rid="B58">Mu&#241;oz-Coego et al., 2022</xref>; <xref ref-type="bibr" rid="B23">Franich, 2024</xref>). New information is also generally less predictable from the discourse context (<xref ref-type="bibr" rid="B22">Ferreira &amp; Lowder, 2016</xref>). A further possibility is that gestures, instead of (or in addition to) being driven away from prosodically/phonetically weak syllables, are <italic>attracted</italic> to less predictable syllables. If this is the case, then predictability may be found to have an effect on the timing of co-speech gestures even if it does not influence word and segment durations in the expected ways. In Section 1.4, we provide further details on the factors that are known to condition gesture timing in Igbo, and outline how predictability may fit in as an additional explanatory variable in co-speech gesture timing. First, however, we provide further details on how predictability can influence speech timing.</p>
</sec>
<sec>
<title>1.3. Predictability Effects on Speech Timing</title>
<p>While predictability has yet to be examined as a factor influencing co-speech gesture timing, there is a large body of existing work demonstrating its effects on speech timing. The term &#8216;predictability&#8217; can encompass many different quantitative measures (see <xref ref-type="bibr" rid="B76">Shaw &amp; Kawahara, 2018</xref> for a recent overview), but generally characterizes the probability with which a given phone or word will arise in speech. While word frequency has long been linked to phonetic reduction (<xref ref-type="bibr" rid="B9">Bybee, 2006</xref>; <xref ref-type="bibr" rid="B20">Ernestus, 2011</xref>; <xref ref-type="bibr" rid="B61">Pierrehumbert, 2001</xref>; <xref ref-type="bibr" rid="B88">Zipf, 1929</xref>) and negatively correlated with word duration, specifically (<xref ref-type="bibr" rid="B5">Bell et al., 2009</xref>; <xref ref-type="bibr" rid="B26">Gahl, 2008</xref>), other forms of <italic>contextual predictability</italic> are demonstrated to have an effect on word and segment durations (<xref ref-type="bibr" rid="B15">Cohen Priva, 2015</xref>; <xref ref-type="bibr" rid="B36">Jurafsky et al., 2001</xref>; <xref ref-type="bibr" rid="B73">Seyfarth, 2014</xref>). The best studied of these effects is forward lexical predictability (FLP), which is the probability that a word will occur given the preceding word (<xref ref-type="bibr" rid="B36">Jurafsky et al., 2001</xref>; <xref ref-type="bibr" rid="B45">Krug, 1998</xref>; <xref ref-type="bibr" rid="B71">Saffran et al. 1996</xref>). FLP has been shown to be an even better predictor of word durations than raw lexical frequency, with words with greater FLP found to exhibit relatively shorter durations than those with lower FLP. Recently, research has suggested that backward lexical predictability (BLP)&#8212;the predictability of a word given the word that follows it&#8212;also exerts an influence on word durations similar to what is observed for FLP (<xref ref-type="bibr" rid="B5">Bell, 2009</xref>; <xref ref-type="bibr" rid="B83">Tang &amp; Shaw, 2021</xref>). In addition to contextual lexical probability, contextual <italic>phonotactic</italic> probability has been explored in some languages. Evidence suggests that segments which are more predictable in their phonological contexts are more prone to lenition (<xref ref-type="bibr" rid="B15">Cohen Priva, 2015</xref>, <xref ref-type="bibr" rid="B16">2017</xref>; <xref ref-type="bibr" rid="B27">Gahl et al., 2012</xref>; Shaw and Kawahara, 2017; <xref ref-type="bibr" rid="B87">Whang, 2018</xref>). And though most work to date has examined phonotactic probability effects on neighboring segments, various researchers have cited predictability as a functional motivation for phonological vowel harmony patterns, in which (typically nonadjacent) vowels within a word or a particular prosodic domain will agree in terms of their phonological features (<xref ref-type="bibr" rid="B80">Suomi, 1983</xref>; <xref ref-type="bibr" rid="B38">Kaun, 1995</xref>; <xref ref-type="bibr" rid="B42">Kimper, 2017</xref>). Indeed, McCollum (<xref ref-type="bibr" rid="B55">2020</xref>) finds that harmonizing vowels that are noninitial in the harmony domain in Kyrgyz are prone to phonetic reduction. Importantly, effects of phonotactic probability on reduction are demonstrated to depend on language-specific pressures around contrast preservation (<xref ref-type="bibr" rid="B16">Cohen Priva, 2017</xref>), meaning that not all languages are expected to show effects of the same measures of phonotactic predictability. For example, in a cross-linguistic study of lenition patterns, Cohen Priva finds that the specific segments that languages opt to lenite is dependent on how contextually predictable target segments are within each language.</p>
<p>Given these findings, one of the goals of the present work is to examine how various forms of predictability interact with speech and gesture timing in Igbo. As will be discussed, the timing of co-speech gestures is especially variable when words occur in phrase-medial position, an environment where words are particularly prone to reduction due, among other things, to a process of vowel coalescence in the language. Predictability measures have been found to increase the likelihood of coalescence, segment reduction, and deletion in previous work (<xref ref-type="bibr" rid="B36">Jurafsky et al., 2001</xref>; <xref ref-type="bibr" rid="B45">Krug, 1998</xref>; <xref ref-type="bibr" rid="B71">Saffran et al., 1996</xref>), leading us to believe that predictability effects may also have an influence on the timing of co-speech gestures. In addition to lexical predictability effects, we focus in on one specific variety of phonotactic predictability&#8212;vowel-to-vowel co-occurrence probability&#8212;which we hypothesize may be especially important for speakers of Igbo, a language with vowel harmony. We predict that vowel-to-vowel predictability will have an influence on word and segment durations, which, in turn, may influence the timing of co-speech gestures. We now provide more general details on gesture patterning in Igbo, with a particular focus on prosodic phrasing and tonal melody.</p>
</sec>
<sec>
<title>1.4. Patterns of Igbo Prosodic Structure and Gesture</title>
<p>Igbo is a Niger-Congo language, part of the Volta-Niger subfamily (sometimes referred to as West Benue-Congo or Eastern Kwa) spoken in Nigeria by approximately 45 million speakers. One of the most extensive works to date on the phonology of Igbo is by Clark (<xref ref-type="bibr" rid="B13">1990</xref>). Among other things, Clark describes a process of floating tone-induced downstep in Igbo in the nominal associative construction which applies differently depending on whether the noun is disyllabic or trisyllabic: initial syllables of disyllabic words undergo downstep, whereas initial syllables of trisyllabic words resist it. Clark interprets this difference in behavior across the two types of nouns as related to their underlying metrical structures. She proposes that metrical prominence at the word level is assigned from right to left, with metrical prominence assigned to word-final syllables and alternating syllables moving leftward. The initial metrically prominent syllable in a trisyllabic word allows it to resist the tone docking that would trigger downstep in the associative construction. Metrical patterns for disyllabic high toned word <italic>ewu</italic> &#8216;goat&#8217; and trisyllabic word <italic>akw&#7909;kw&#7885;</italic> &#8216;book&#8217; are given in (1). Note that Igbo also has a process of ATR vowel harmony, such that vowels within the word must generally match in ATR value ([&#8211;ATR] vowels are marked below with the underdot diacritic).</p>
<list list-type="gloss">
<list-item>
<list list-type="wordfirst">
<list-item><p>(1)</p></list-item>
</list>
</list-item>
<list-item>
<list list-type="sentence-gloss">
<list-item>
<list list-type="final-sentence">
<list-item><p>Metrical patterns for Igbo words <italic>ewu</italic> &#8216;goat&#8217; and <italic>a&#803;kw&#7909;kw&#7885;</italic> &#8216;book.&#8217; Stars represent sites of metrical prominence within the word.</p></list-item>
</list>
</list-item>
</list>
</list-item>
</list>
<list list-type="gloss">
<list-item>
<list list-type="wordfirst">
<list-item><p>&#160;</p></list-item>
</list>
<list list-type="wordfirst">
<list-item><p>a.</p></list-item>
</list>
</list-item>
<list-item>
<list list-type="sentence-gloss">
<list-item>
<list list-type="final-sentence">
<list-item><p><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-17-18890-g10.png"/></p></list-item>
</list>
</list-item>
</list>
</list-item>
</list>
<list list-type="gloss">
<list-item>
<list list-type="wordfirst">
<list-item><p>&#160;</p></list-item>
</list>
<list list-type="wordfirst">
<list-item><p>b.</p></list-item>
</list>
</list-item>
<list-item>
<list list-type="sentence-gloss">
<list-item>
<list list-type="final-sentence">
<list-item><p><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-17-18890-g11.png"/></p></list-item>
</list>
</list-item>
</list>
</list-item>
</list>
<p>Independent evidence for the metrical status of word-final syllables in Igbo comes from our previous work on co-speech gesture timing, in which we show that word-final syllables are the most likely to co-occur with a gesture (Franich et al., 2025; see also <xref ref-type="fig" rid="F1">Figure 1</xref> below). However, word-final metrical prominence is not the only driving factor behind co-speech gesture timing in Igbo. We also find that gesture position is influenced by tone melody: specifically, words that contain a HH melody are more likely to have gesture aligned to the first high of the HH sequence. Thus, in a disyllabic word with a HH tone melody&#8212;attributable to a single, multiply linked high tone assigned from the left edge of the word (2)&#8212;gestures are significantly more likely to be attracted to the initial syllable of the word than the final syllable. Interestingly, in trisyllabic words with a HHH melody (again, attributable to a single multiply linked high tone), gestures gravitate to both the initial and final syllables, but less so to the medial syllable. This led Franich et al. (2025) to propose that the effects of tone melody are best explained through the mechanism of <italic>tonal feet</italic>, whereby binary groupings of high tones are assigned from left to right in the word; tonal feet can then compete with feet at the segmental level for the attraction of gestures. An example of a disyllabic HH word with both a segmental and a tonal foot is provided in (2); the position of the foot head at each level is underlined. For a trisyllabic HHH word, we propose that the word would be footed such that the initial two syllables form a tonal foot, and the third syllable initiates a new tonal foot, resulting in a parse of (<underline>H</underline>H)(<underline>H</underline>).</p>
<fig id="F1">
<caption>
<p><bold>Figure 1:</bold> Alignment of gesture apexes within Igbo disyllabic (left) and trisyllabic (right) words, organized in facets in terms of the position of target words within an intonational phrase (reproduced from Franich et al. 2025, pp. 21&#8211;24).</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-17-18890-g1.png"/>
</fig>
<list list-type="gloss">
<list-item>
<list list-type="wordfirst">
<list-item><p>(2)</p></list-item>
</list>
</list-item>
<list-item>
<list list-type="sentence-gloss">
<list-item>
<list list-type="final-sentence">
<list-item><p>Tonal and Segmental Foot Parses for <italic>oce</italic> &#8216;chair&#8217;</p></list-item>
<list-item><p><inline-graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-17-18890-g12.png"/></p></list-item>
</list>
</list-item>
</list>
</list-item>
</list>
<p>The relevance of tonal feet to gesture timing in Igbo is interesting in light of findings presented in Sections 3 and 4 on the effects of vowel co-occurrence probability on co-speech gesture timing. Briefly, our findings support the idea that co-speech gestures are avoided on <italic>phonologically redundant</italic> syllables; in other words, those that bear the same features as the preceding syllable. We discuss how tonal feet may fit into a broader pattern with respect to redundancy effects in shaping prosodic patterning.</p>
<p>Franich et al., (2025) also find an interaction between word and phrase position in predicting gesture timing, driven by the fact that words occurring in phrase-medial position are more likely to experience a &#8220;shift&#8221; in gesture timing from the word-final syllable to an earlier syllable in the word (initial position for disyllabic words, and medial position for trisyllabic words). <xref ref-type="fig" rid="F1">Figure 1</xref> shows results for the effect of phrase position on gesture timing in disyllabic and trisyllabic words in Igbo. Importantly, we find ample evidence that gesture timing varies within the <underline>same word</underline> according to phrase position. For example, one frequent word in the corpus, <italic>a&#768;&#803;na&#768;&#803;</italic>, meaning &#8216;land,&#8217; consistently showed alignment of gestures to word-final position when it occurred in phrase-initial and phrase-final position, but showed alignment of gestures to word-initial position about half the time when it occurred in phrase-medial position. Given the observed variation within the same words, the effects of phrase position do not seem to be reducible to differences in the prosodic structure of the words under investigation.</p>
<p>Franich et al. (2025) hypothesize that the difference in patterning across phrase positions may be attributable to the susceptibility of word-final syllables to phonetic reduction processes, including (but perhaps not limited to) a gradient process of vowel coalescence that targets word-final vowels when they occur in hiatus before another vowel. <xref ref-type="bibr" rid="B89">Zsiga 1997</xref> shows that this process&#8212;which is highly common due to the prevalence of words of shape VCV and VCVCV&#8212;can be modeled as a combination of final vowel shortening and articulatory overlap. She argues that the articulatory gestures (e.g., of the tongue and lips) of the word-final vowel are shortened in duration and heavily coarticulated with those of the following vowel.<xref ref-type="fn" rid="n1">1</xref></p>
<p>As discussed in Section 1.2, one interpretation of the effect of phrase position on gesture timing in Igbo is that words undergo metrical reorganization when word-final vowels are especially prone to reduction, and gesture shift reflects this reorganization. Reduction of a prominent syllable with an accompanying shift in prominence cues is not unexpected in the context of metrical patterning more generally: there are many documented cases of reduction-related stress shift across different languages (<xref ref-type="bibr" rid="B53">Lightner, 1972</xref>; <xref ref-type="bibr" rid="B32">Halle &amp; Vergnaud, 1987</xref>; <xref ref-type="bibr" rid="B1">Al-Mozainy et al., 1985</xref>). Reduction of the word-final vowel in Igbo is also consistent with the typology of vowel coalescence and assimilation patterns across the world&#8217;s languages. Word-final vowels&#8212;which are the initial vowels in the V#V sequence undergoing coalescence in this case&#8212;should reduce more than the word-initial vowels that follow them (<xref ref-type="bibr" rid="B10">Casali, 1997</xref>). In summary, &#8216;gesture shift&#8217; in Igbo bears many of the hallmarks of prosodic reorganization and stress shift cross-linguistically.</p>
<p>Given that coalescence in Igbo is a gradient phenomenon (<xref ref-type="bibr" rid="B89">Zsiga, 1997</xref>), and one that appears to be conditioned by factors that encourage vowel reduction, it is plausible that predictability could serve as an additional pressure leading to coalescence and &#8216;gesture shift&#8217; away from final position in the word. If predictability effects are largely implemented through prosodic planning, as suggested by models such as Aylett and Turk (<xref ref-type="bibr" rid="B3">2004</xref>) and Tang and Shaw (<xref ref-type="bibr" rid="B77">2023</xref>), then we expect that they should have a similar effect on co-speech gesture timing as other prosodic variables that influence word and segment durations, such as prosodic phrasing. In other words, if, like prosodic phrasing, predictability leads to shortening of word durations and a greater incidence of vowel coalescence, then greater predictability may be related to a greater incidence of gesture shift to earlier positions in the word. We term this the <italic>duration reduction hypothesis</italic>. If, however, phonetic reduction effects related to prosodic phrasing and predictability arise from different stages in speech planning (as is possible in the model by <xref ref-type="bibr" rid="B39">Keating &amp; Shattuck-Hufnagel, 2002</xref>), then co-speech gesture timing may be influenced by prosodic phrasing but not by predictability measures. We term this the <italic>prosodic reduction hypothesis</italic>. Finally, it is possible that predictability will exhibit effects on gesture timing that are less directly tied to duration shortening. For example, it could be the case that variation in gesture timing is partly driven by the attraction of gestures to less predictable, more information-rich syllables,<xref ref-type="fn" rid="n2">2</xref> which also happen to be less prone to durational reduction. In this case, vowel co-occurrence predictability may work in tandem with other prosodic factors in driving gestures away from word-final position when earlier syllables in the word are relatively less predictable/more informative. In this scenario, we would expect predictability measures, but not necessarily word durations, to have an influence on gesture timing. We call this the <italic>information hypothesis</italic>.</p>
</sec>
</sec>
<sec>
<title>2. Method</title>
<sec>
<title>2.1. Participants and Procedure</title>
<p>This research was granted ethics approval by the Institutional Review Board of Harvard University (Protocol IRB 22-1095). Four Igbo speakers from the greater Abuja area in Nigeria (3 men, 1 woman) participated in the study. The age range of the participants was between 35 and 55 years. Speakers were interviewed by a native Igbo speaker (the second author) in a quiet room or outdoor area. They were interviewed about aspects of language and cultural practices, such as traditions around marriage ceremonies and child naming.</p>
<p>Individual audio tracks were recorded for each participant using a Shure SM10 head-mounted microphone input to a Zoom Q8 video/audio recording device at an audio sampling rate of 48 kHz and a video sampling rate of 30 frames per second. Interviews were conducted for around 25&#8211;30 minutes, resulting in 9,070 words and around 25,000 phones recorded. Within this corpus, a total of 1073 <italic>gesture-aligned phones</italic>, or phones during which the manual apex of a gesture stroke occurred (see further details in Section 2.2), were annotated (annotation of gestures in the corpus is ongoing). Since the primary variable of interest was gesture position (by syllable) within the word, we exclude monosyllabic words from the present analysis. Furthermore, due to the relative sparsity of longer words in our corpus, as well as the tendency of longer words to be more complex in their structure, we limit our analysis to disyllabic and trisyllabic words. Our analysis focuses on the subset of these words whose tone melodies accounted for at least 2% of the total data, 478 words in all. The vast majority of the gestures obtained in the data (75%) were categorized as beat gestures, which constitute small, rhythmic hand and arm movements that typically align with prosodic emphasis but do not necessarily carry depictive meaning (e.g., as in the case of an iconic or a deictic gesture). We opt to include all gesture types in the present analysis in order to maximize our sample size, based on the fact that even iconic and deictic gestures are known to be constrained in their timing by prosodic structure (<xref ref-type="bibr" rid="B74">Shattuck-Hufnagel &amp; Prieto, 2019</xref>; <xref ref-type="bibr" rid="B67">Rohrer et al., 2023</xref>).</p>
</sec>
<sec>
<title>2.2. Processing of Audio and Video Data</title>
<p>Audio data were transcribed by a native speaker (the second author) and force-aligned using the FAVE aligner (<xref ref-type="bibr" rid="B68">Rosenfelder et al., 2022</xref>). Alignments were subsequently checked for accuracy by both authors. Manual gestures were coded by a team of trained annotators in ELAN software (<xref ref-type="bibr" rid="B19">ELAN 2024</xref>). Coding was carried out according to a set of criteria adapted from the MIT Speech Communications Group Gesture Studies Coding Manual (<xref ref-type="bibr" rid="B57">MIT Speech Communications Group, 2020</xref>). Coders kept the file&#8217;s audio muted to avoid auditory bias in coding decisions. Coders would observe which hand the speaker tended to gesture with most, and treated that hand as the dominant hand for the sake of coding. Coders were tasked with marking off intervals corresponding to <italic>communicatively intentional</italic> movements of the hands during communication. Movement intentionality is given by Shattuck-Hufnagel and Ren (<xref ref-type="bibr" rid="B75">2018</xref>) as a defining characteristic of co-speech gestures. Though the measure has yet to be defined kinematically, we note that it is likely indexed most closely by acceleration of the manual articulators (with faster acceleration for communicatively intentional gestures vs. fidgets). Part of the task of establishing reliability between coding partners in our study is accurately distinguishing &#8220;true&#8221; gestures from fidgets, based on impressions of intentionality. Our coders agreed over 95% of the time about which movements should be coded as gestures vs. fidgets in the dataset.</p>
<p>After identifying a gesture, coders marked various phases within it, including the <italic>preparation</italic> phase (if present), where the hands are brought into position before the gesture begins, the <italic>stroke</italic>, or the onset of intentional movement, any <italic>holds</italic> (pre- or poststroke, if present), where the hands rest in position after the stroke, and the <italic>recovery</italic> (if present), where the hands return to a rest position (<xref ref-type="fig" rid="F2">Figure 2</xref>; <xref ref-type="table" rid="T1">Table 1</xref>). All these intervals were marked on a single ELAN tier. On a separate tier, coders marked off an interval within the stroke phase which corresponded with the gesture <italic>apex</italic>. In prior work, the definition of the apex has varied, with some authors using a spatial definition (usually the point of maximum extension of articulators such as the fingers in a hand gesture) (<xref ref-type="bibr" rid="B54">Loehr, 2012</xref>), while others have relied on a definition based on gesture timing, such as the peak or minimum velocity of movement of an articulator (<xref ref-type="bibr" rid="B62">Pouw &amp; Dixon, 2019</xref>; <xref ref-type="bibr" rid="B64">Pouw, 2020</xref>). Given that peak velocity of the gesture has been found to closely align with prosodically prominent syllables, we opt to use peak velocity as our measure of apex timing. Apexes were coded in our data based on visual inspection, rather than through motion tracking data. Specifically, coders marked off the video frame in the data during which the hands moved the most, as reflected in visualized blurring of the hands across the video frame. We note that an alternative strategy of analyzing minimum velocity/peak extension of gestures did not yield substantially different results, likely because these landmarks tend to occur very close in time (at a latency much shorter than the average syllable). In prior work, we have demonstrated that this visualized point of peak velocity closely approximates the true moment of peak velocity as measured computationally (<xref ref-type="bibr" rid="B18">Dych et al., 2023</xref>).</p>
<fig id="F2">
<caption>
<p><bold>Figure 2:</bold> Key phases of a co-speech gesture. Pregesture phase reflects the position of the hands before the onset of the gesture. The preparation phase reflects the period where the hands are raised into position to execute the gesture. The stroke onset represents the beginning of intentional movement toward the gesture apex. Peak velocity of hand movements (reflected in blurring within video frame) is reached during the course of the stroke and shortly before the point of maximum extension of the hands.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-17-18890-g2.jpg"/>
</fig>
<table-wrap id="T1">
<caption>
<p><bold>Table 1:</bold> Gesture landmarks with kinematic definitions.</p>
</caption>
<table>
<tbody>
<tr>
<td align="left" valign="top"><bold>Gesture Landmark</bold></td>
<td align="left" valign="top"><bold>Kinematic Definition</bold></td>
<td align="left" valign="top"><bold>Required phase?</bold></td>
</tr>
<tr>
<td align="left" valign="top">Preparation onset</td>
<td align="left" valign="top">Start of movement of the hand into position, prior to stroke onset</td>
<td align="left" valign="top">No</td>
</tr>
<tr>
<td align="left" valign="top">Stroke onset</td>
<td align="left" valign="top">Start of intentional movement of the hand, regardless of direction</td>
<td align="left" valign="top">Yes</td>
</tr>
<tr>
<td align="left" valign="top">Stroke apex</td>
<td align="left" valign="top">Timing of peak velocity of manual movement, as observed from relative distance moved/blurriness from one frame to the next</td>
<td align="left" valign="top">Yes</td>
</tr>
<tr>
<td align="left" valign="top">Stroke offset</td>
<td align="left" valign="top">Endpoint of intentional movement of the hand</td>
<td align="left" valign="top">Yes</td>
</tr>
<tr>
<td align="left" valign="top">Recovery onset</td>
<td align="left" valign="top">Start of less intentional movement of the hand towards a rest position</td>
<td align="left" valign="top">No</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Coders were also tasked with establishing reliability of the timing of their gesture apexes. To do so, pairs of coders were assigned the same participant&#8217;s file, annotating the file independently and then comparing their results. Their goal was to code such that the apex annotations in their respective versions of the file were within a single video frame. After establishing reliability (above 85%) on a series of video samples, coders in each pair worked on a participant&#8217;s file together, continuing to check that their judgments about gesture timing aligned, and discussing any disagreements that arose between them. Where consensus between coders could not be reached, a third research team member was brought in for a tie-breaking judgment.</p>
</sec>
<sec>
<title>2.3. Data Postprocessing</title>
<p>Timestamps for gesture and speech data were extracted from ELAN .eaf files and Praat .TextGrid files, respectively, and merged using a Python script. Gestures were treated as &#8216;aligning&#8217; with a particular phone if the apex of the gesture overlapped with any part of the phone. Each syllable of a gesture-aligned word was then coded for position in the word, position in the intonational phrase, the tone with which the syllable was realized, and the tone melody of the word as a whole. Cues to intonational phrases in Igbo are still not fully defined; as such, for the purposes of the present analysis, we use the edges of &#8216;breath groups&#8217; (<xref ref-type="bibr" rid="B52">Lieberman, 1967</xref>) as an approximation of intonational phrase boundaries (<xref ref-type="bibr" rid="B60">Pierrehumbert, 1980</xref>). We observed a pattern of pitch reset&#8212;a common marker of an intonational phrase boundary across languages (<xref ref-type="bibr" rid="B17">Downing &amp; Rialland, 2017</xref>; <xref ref-type="bibr" rid="B46">K&#252;gler, 2017</xref>; <xref ref-type="bibr" rid="B48">Kula &amp; Hamann, 2017</xref>)&#8212;at the beginnings of units consistent with most breath groups, giving us confidence that these breath groups were generally coextensive with intonational phrases.</p>
</sec>
<sec>
<title>2.4. Measures of Predictability</title>
<p>Given prior evidence on the importance of lexical frequency and contextual lexical probabilities on word durations, we calculated both unigram frequency and forward and backward bigram probabilities for the words in our corpus. The equations for forward and backward bigram probabilities are given in (3).</p>
<list list-type="gloss">
<list-item>
<list list-type="wordfirst">
<list-item><p>(3)</p></list-item>
</list>
</list-item>
<list-item>
<list list-type="sentence-gloss">
<list-item>
<list list-type="final-sentence">
<list-item><p>Equations for Forward and Backward Bigram Probabilities</p></list-item>
</list>
</list-item>
<list-item>
<list list-type="word">
<list-item><p>Forward Bigram Probability:</p></list-item>
<list-item><p>P(<italic>b</italic>&#124;<italic>a</italic>) = P(<italic>ab</italic>)/P(<italic>a</italic>)</p></list-item>
</list>
<list list-type="word">
<list-item><p>Backward Bigram Probability:</p></list-item>
<list-item><p>P(<italic>a</italic>&#124;<italic>b</italic>) = P(<italic>ab</italic>)/P(<italic>b</italic>)</p></list-item>
</list>
</list-item>
</list>
</list-item>
</list>
<p>Following prior work (<xref ref-type="bibr" rid="B73">Seyfarth, 2014</xref>; <xref ref-type="bibr" rid="B83">Tang &amp; Shaw, 2021</xref>), lexical contextual predictability was calculated based on two language models which characterized the probability of each word in the corpus (containing a total of 9,070 words) given the word immediately preceding the target word or the word immediately following it. Language models were created using the <italic>kgrams</italic> package (<xref ref-type="bibr" rid="B30">Gherardi, 2024</xref>) for <italic>R</italic> statistical software (<xref ref-type="bibr" rid="B65">R Core Team 2024</xref>). Resulting model probabilities were then smoothed using the modified Kneser-Ney smoothing algorithm (<xref ref-type="bibr" rid="B11">Chen &amp; Goodman, 1999</xref>) initialized with default parameters specified in the <italic>kgrams</italic> package. We note that the reliability of our language models will be somewhat limited compared with other studies given the small size of our corpus; unfortunately, given that tone in Igbo is rarely marked in orthographic representations, there is currently little existing data on which a more reliable language model could be built, since the functional load for tone is quite high in Igbo.</p>
<p>In addition to these measures of lexical predictability, we also computed a measure of phonotactic probability that we hypothesize should be important for Igbo speakers: vowel-to-vowel co-occurrence probability. Given Igbo is a language with vowel harmony, featural dependencies across vowels form an important part of speech production planning. Not only that, but speakers of languages that utilize vowel harmony rely on harmony patterns in word recognition in a way that speakers of languages without vowel harmony do not (<xref ref-type="bibr" rid="B81">Suomi et al., 1997</xref>; <xref ref-type="bibr" rid="B86">Vroomen et al., 1998</xref>). Furthermore, vowel co-occurrence predictability based on vowel harmony has been shown to lead to phonetic reduction (<xref ref-type="bibr" rid="B55">McCollum, 2020</xref>), a critical variable in our study. Here, we utilize two measures of vowel-to-vowel predictability: forward contextual predictability (the probability that a vowel will occur based on the preceding vowel) and backward contextual predictability (the probability that a vowel will occur based on the following vowel). For trisyllabic words, we focus specifically on predictability between the medial and final vowels within the word, since final vowels are those which are most targeted for reduction in Igbo.</p>
</sec>
<sec>
<title>2.5. Hypotheses and Predictions</title>
<p>In our previous work, we observed that co-speech gestures were more variable in their timing when they occurred in phrase-medial position. We hypothesized that this was due to the greater susceptibility of word-final vowels to coalescence and durational reduction in this phrasal position, but did not evaluate that possibility empirically. Here, we test that hypothesis by examining whether a) words are shorter when they occur in phrase-medial position compared with other positions; and b) word duration is a significant predictor of co-speech gesture position within the word. We also test the hypothesis that words will be shorter when predictability&#8212;including lexical frequency, contextual lexical probability, and vowel co-occurrence probability&#8212;is high, and that predictability will have an impact on gesture timing. As proposed in Section 1.4, the <italic>duration reduction hypothesis</italic> proposes that the causal chain between predictability, phrase position, and gesture shift is as follows: predictability and prosodic phrasing lead to shortening of the word, which makes it more susceptible to vowel coalescence. This leads to a shift in prominence, with an accompanying shift of a gesture away from word-final position. Under this scenario, both phrase position and predictability would condition vowel coalescence and gesture shift because they both contribute to reduction of word durations. As such, words that are both phrase-medial and more predictable will be most likely to undergo phonetic reduction, and will therefore be the most likely to exhibit gesture shift to syllables earlier than the word-final syllable. However, recall from Section 1.2 that certain models of prosodic planning propose that prosodic restructuring&#8212;the phenomenon hypothesized to be implicated in Igbo gesture shift&#8212;may not be sensitive to all types of phonetic variation. In the event that the phonetic effects of predictability are not planned by the time that restructuring takes place, we would expect effects of predictability on word durations, but not effects of predictability on gesture timing (consistent with the <italic>prosodic reduction hypothesis</italic>).</p>
<p>Duration reduction may not be the only path through which predictability could influence gesture timing. The <italic>information hypothesis</italic> states that gestures will generally gravitate to syllables that are less predictable and more informative, regardless of durational patterning. Recall from Section 1.3 that certain kinds of predictability measures, such as phonotactic predictability, do not affect segment durations in all languages. Nonetheless, we might still expect such measures to influence co-speech gesture timing. There are two ways that the information hypothesis could be supported in our data. First, it would be supported if gesture timing is sensitive to predictability measures but segment durations are not conditioned by these same measures. Second, it would be supported if gesture timing is sensitive to predictability measures, but the effect is demonstrably not driven by duration reduction. For example, gestures could be more likely to shift in words with high vowel co-occurrence predictability (a measure that is theoretically linked with vowel harmony patterns), even if these words are not phonetically reduced.</p>
<p>Related to this last point, we incorporate tone melody into our statistical models both as a way of controlling for known influences on co-speech gesture timing and to investigate possible interactions of tone melody and predictability in influencing gesture timing. In our previous work, we have shown that co-speech gestures tend to shift to initial position of a HH sequence (Franich et al., 2025). We have analyzed these effects as deriving from prominence asymmetries at the level of the tonal foot, in which the leftmost high-toned syllable is considered the head of the constituent. Nonetheless, HH sequences also embody an <italic>informational</italic> asymmetry: the second syllable in the sequence carries a feature which is redundant with that of the preceding syllable. To the extent that the information hypothesis can capture some of the variability in co-speech gesture timing in our data, we might expect that tone melody could interact with predictability measures in influencing gesture timing.</p>
</sec>
<sec>
<title>2.6. Statistical Analysis</title>
<p>Three sets of statistical models were used to test our hypotheses. First, we constructed linear mixed effects models to investigate the effect of phrase position (three levels: initial, medial, and final) on both word duration and phone rate, with number of syllables (two vs. three) also included as a variable in the model. We opt to look at both word duration and phone rate in these models since they each provide a slightly different view of the effect of prosodic structure on speech timing. Word duration is a commonly used variable in the study of both prosodic and predictability-related effects on phonetic reduction (e.g., <xref ref-type="bibr" rid="B4">Baker &amp; Bradlow 2009</xref>; <xref ref-type="bibr" rid="B6">Bell et al. 2003</xref>). However, prior work by Ryan (<xref ref-type="bibr" rid="B69">2016</xref>; <xref ref-type="bibr" rid="B70">2019</xref>) has shown effects of prosodic &#8216;end-weight&#8217;&#8212;whereby heavier syllables and longer words tend to occur more frequently phrase-finally&#8212;which may be a confound in studying word duration effects. Phone rate is a measure of reduction that is robust to these effects, so it is useful as a complementary measure. An additional model investigated the interaction between phrase position, word position (again, with three levels: initial, medial, and final) and gesture presence (gesture vs. no gesture) on phone duration. For these three models, word and phone durations were log-transformed, and word and phrase position were sum-coded to allow us to compare how durations at each position compared with the grand mean. The next set of models examined the effects of the various predictability measures on word duration. All probability-related predictor variables were log-transformed, as was word duration. All continuous predictors were mean-centered to avoid collinearity (VIF values for fixed effects were all at 7 or below). Following previous work (e.g., <xref ref-type="bibr" rid="B5">Bell et al., 2009</xref>), we include two-way interactions between lexical frequency and our main predictability variables of interest, forward and backward lexical contextual probability and forward and backward vowel-to-vowel co-occurrence probability. We also include two-way interactions for our two lexical contextual probability measures and our two vowel-to-vowel probability measures. Finally, a mixed effects logistic regression model was used to investigate how phrase position, tone melody, predictability, and duration all contribute to the positioning of co-speech gestures within the word. The dependent variable here, word position, had two levels: final vs. nonfinal. For all but the final model (see Section 3.3), <italic>p-</italic>values were derived using Satterthwaite&#8217;s approximation. Post hoc comparisons were conducted in <italic>emmeans</italic> (<xref ref-type="bibr" rid="B49">Lenth et al. 2024</xref>) with <italic>p</italic>-values adjusted using Tukey&#8217;s HSD method. By-subject random intercepts were included for all models. Error bars on all plots reflect 95% confidence intervals.</p>
</sec>
</sec>
<sec>
<title>3. Results</title>
<sec>
<title>3.1. Durational Effects of Prosodic Phrasing and its Relation to Gesture Timing</title>
<p>Since no work to date has investigated effects of phrase position on word durations in Igbo, we first constructed a model investigating this effect. The model included phrase position and syllable count, as well as their interaction. Results are presented in <xref ref-type="table" rid="T2">Table 2</xref>. As expected, trisyllabic words had significantly longer durations than disyllabic words (<italic>p</italic> &lt; .001). Differences in duration between words found in phrase-medial vs. phrase-final position were also found, with phrase-medial words realized with shorter durations, a significant difference from the grand mean (<italic>p</italic> &lt; .05; <xref ref-type="fig" rid="F3">Figure 3</xref>). Meanwhile, word-final syllables were realized with significantly longer durations compared with the grand mean (<italic>p</italic> &lt; .001). Phrase-initial syllables did not differ significantly from the grand mean in their durations (<italic>p</italic> = .42). To ensure that the observed effects of phrase position on word durations was not simply an artifact of prosodic end weight, we also examined effects of phrase position on phone rate, measured as phones per second (<xref ref-type="table" rid="T3">Table 3</xref>). We find that phone rate for phrase-medial words is significantly higher than the grand mean (<italic>p</italic> &lt; .01), whereas it is significantly lower than the grand mean for phrase-final position (<italic>p</italic> &lt; .01) (<xref ref-type="fig" rid="F4">Figure 4</xref>). Phone rate in phrase-initial position did not differ significantly from the grand mean (<italic>p</italic> = .91). Taken together, these results provide evidence for phrase-final lengthening in Igbo, consistent with findings from other languages (Edwards et al., 1991; <xref ref-type="bibr" rid="B85">Turk &amp; Shattuck-Hufnagel, 2007</xref>; <xref ref-type="bibr" rid="B7">Berkovits, 1993</xref>; <xref ref-type="bibr" rid="B37">Katsika, 2016</xref>; <xref ref-type="bibr" rid="B72">Seo et al., 2019</xref>; <xref ref-type="bibr" rid="B59">Paschen et al., 2022</xref>; <xref ref-type="bibr" rid="B41">Kim et al., 2024</xref>). They also suggest, importantly, that phrase-medial position is subject to a greater level of phonetic reduction than phrase-initial or phrase-final position, consistent with the hypothesis proposed (but untested) in Franich et al. (2025).</p>
<table-wrap id="T2">
<caption>
<p><bold>Table 2:</bold> Influence of phrase position and syllable count on word duration.</p>
</caption>
<table>
<tbody>
<tr>
<td align="left" valign="top"><bold>Term</bold></td>
<td align="center" valign="top"><bold><italic>&#946;</italic></bold></td>
<td align="center" valign="top"><bold><italic>t-value</italic></bold></td>
<td align="center" valign="top"><bold><italic>p-value</italic></bold></td>
</tr>
<tr>
<td align="left" valign="top">(Intercept)</td>
<td align="right" valign="top">&#8211;0.953</td>
<td align="right" valign="top">&#8211;18.998</td>
<td align="left" valign="top">&lt;.001***</td>
</tr>
<tr>
<td align="left" valign="top">Phrase Position (Initial)</td>
<td align="right" valign="top">&#8211;0.440</td>
<td align="right" valign="top">&#8211;0.808</td>
<td align="left" valign="top">=.42</td>
</tr>
<tr>
<td align="left" valign="top">Phrase Position (Medial)</td>
<td align="right" valign="top">&#8211;0.105</td>
<td align="right" valign="top">&#8211;2.563</td>
<td align="left" valign="top">&lt;.05*</td>
</tr>
<tr>
<td align="left" valign="top">Phrase Position (Final)</td>
<td align="right" valign="top">0.148</td>
<td align="right" valign="top">3.313</td>
<td align="left" valign="top">&lt;.001***</td>
</tr>
<tr>
<td align="left" valign="top">Syllables</td>
<td align="right" valign="top">0.035</td>
<td align="right" valign="top">5.457</td>
<td align="left" valign="top">&lt;.001***</td>
</tr>
<tr>
<td align="left" valign="top">Phrase Position (Initial) * Syllables</td>
<td align="right" valign="top">0.006</td>
<td align="right" valign="top">0.052</td>
<td align="left" valign="top">=.96</td>
</tr>
<tr>
<td align="left" valign="top">Phrase Position (Medial) * Syllables</td>
<td align="right" valign="top">&#8211;0.011</td>
<td align="right" valign="top">&#8211;0.159</td>
<td align="left" valign="top">=.87</td>
</tr>
<tr>
<td align="left" valign="top">Phrase Position (Final) * Syllables</td>
<td align="right" valign="top">0.005</td>
<td align="right" valign="top">0.068</td>
<td align="left" valign="top">=.95</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F3">
<caption>
<p><bold>Figure 3:</bold> Word duration by phrase position and number of syllables.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-17-18890-g3.png"/>
</fig>
<table-wrap id="T3">
<caption>
<p><bold>Table 3:</bold> Influence of phrase position and syllable count on phone rate.</p>
</caption>
<table>
<tbody>
<tr>
<td align="left" valign="top"><bold>Term</bold></td>
<td align="center" valign="top"><bold><italic>&#946;</italic></bold></td>
<td align="center" valign="top"><bold><italic>t-value</italic></bold></td>
<td align="center" valign="top"><bold><italic>p-value</italic></bold></td>
</tr>
<tr>
<td align="left" valign="top">(Intercept)</td>
<td align="right" valign="top">14.525</td>
<td align="right" valign="top">23.226</td>
<td align="left" valign="top">&lt;.001***</td>
</tr>
<tr>
<td align="left" valign="top">Phrase Position (Initial)</td>
<td align="right" valign="top">0.122</td>
<td align="right" valign="top">0.115</td>
<td align="left" valign="top">=.91</td>
</tr>
<tr>
<td align="left" valign="top">Phrase Position (Medial)</td>
<td align="right" valign="top">1.701</td>
<td align="right" valign="top">2.740</td>
<td align="left" valign="top">&lt;.01**</td>
</tr>
<tr>
<td align="left" valign="top">Phrase Position (Final)</td>
<td align="right" valign="top">&#8211;1.823</td>
<td align="right" valign="top">&#8211;2.779</td>
<td align="left" valign="top">&lt;.01**</td>
</tr>
<tr>
<td align="left" valign="top">Syllables</td>
<td align="right" valign="top">0.513</td>
<td align="right" valign="top">0.904</td>
<td align="left" valign="top">=.37</td>
</tr>
<tr>
<td align="left" valign="top">Phrase Position (Initial) * Syllables</td>
<td align="right" valign="top">0.463</td>
<td align="right" valign="top">0.440</td>
<td align="left" valign="top">=.66</td>
</tr>
<tr>
<td align="left" valign="top">Phrase Position (Medial) * Syllables</td>
<td align="right" valign="top">&#8211;0.105</td>
<td align="right" valign="top">&#8211;0.169</td>
<td align="left" valign="top">=.87</td>
</tr>
<tr>
<td align="left" valign="top">Phrase Position (Final) * Syllables</td>
<td align="right" valign="top">&#8211;0.359</td>
<td align="right" valign="top">&#8211;0.547</td>
<td align="left" valign="top">=.58</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F4">
<caption>
<p><bold>Figure 4:</bold> Phone rate by phrase position and syllable count.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-17-18890-g4.png"/>
</fig>
<p>Looking now at how duration interacts with gesture presence, we constructed a mixed effects logistic regression model aimed at evaluating whether word durations (log-transformed) predicted gesture position within the word (final vs. nonfinal). The model also included syllable count as well as an interaction between syllable count and word duration. Results indicated a significant effect of word duration on gesture position (<italic>p</italic> &lt; .01; <xref ref-type="table" rid="T4">Table 4</xref>), as well as a significant effect of syllable count (<italic>p</italic> &lt; .001). As seen in <xref ref-type="fig" rid="F5">Figure 5</xref>, gestures were more likely to occur word-finally when word durations were longer. Though this relationship was observed to have a more positive slope for trisyllabic words, the interaction between word duration and syllable count did not reach significance (<italic>p</italic> = .10).</p>
<table-wrap id="T4">
<caption>
<p><bold>Table 4:</bold> Influence of word duration and syllable count on gesture position.</p>
</caption>
<table>
<tbody>
<tr>
<td align="left" valign="top"><bold>Term</bold></td>
<td align="center" valign="top"><bold><italic>&#946;</italic></bold></td>
<td align="center" valign="top"><bold><italic>z-value</italic></bold></td>
<td align="center" valign="top"><bold><italic>p-value</italic></bold></td>
</tr>
<tr>
<td align="left" valign="top">(Intercept)</td>
<td align="right" valign="top">0.672</td>
<td align="right" valign="top">3.464</td>
<td align="left" valign="top">&lt;.001***</td>
</tr>
<tr>
<td align="left" valign="top">Log Word Duration</td>
<td align="right" valign="top">0.379</td>
<td align="right" valign="top">2.817</td>
<td align="left" valign="top">&lt;.01**</td>
</tr>
<tr>
<td align="left" valign="top">Syllables</td>
<td align="right" valign="top">&#8211;1.042</td>
<td align="right" valign="top">&#8211;5.140</td>
<td align="left" valign="top">&lt;.001***</td>
</tr>
<tr>
<td align="left" valign="top">Log Word Duration * Syllables</td>
<td align="right" valign="top">&#8211;0.327</td>
<td align="right" valign="top">&#8211;1.643</td>
<td align="left" valign="top">=.10</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F5">
<caption>
<p><bold>Figure 5:</bold> Co-speech gesture position by word position and syllable count.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-17-18890-g5.png"/>
</fig>
<p>To better understand the relationship between durations and gesture presence, we also examined phone durations across different word and phrase positions with or without an accompanying co-speech gesture. In <xref ref-type="fig" rid="F6">Figure 6</xref>, we see that while phrase-final syllables usually have longer durations than initial and medial syllables, this pattern changes in phase-medial position, where word-final syllables tend to have some of the shortest durations. This interaction between phrase position and word position was significant (<italic>p</italic> &lt; .05; <xref ref-type="table" rid="T5">Table 5</xref>). These findings are again consistent with the idea that word-final syllables are reduced to a larger degree in phrase-medial position. It is also the case that syllables in all positions have longer durations in the presence of a gesture than when they do not carry a gesture (<italic>p</italic> &lt; .001). This is consistent with the idea that syllables associated with co-speech gestures bear greater prosodic prominence than those that do not.<xref ref-type="fn" rid="n3">3</xref> Although collinearity effects precluded the incorporation of two- and three-way interactions in our model between word position, phrase position, and co-speech gesture presence, visual comparison of results in <xref ref-type="fig" rid="F6">Figure 6</xref> shows that the effects of co-speech gesture are not uniform across word and phrase positions when a gesture is present. Of note is the fact that word-medial syllables receive an especially large boost in duration in the presence of a co-speech gesture, precisely when the word occurs in phrase-medial position. This is consistent with the idea that gesture shift in phrase-medial position occurs in response to a shift in prosodic prominence within the word.</p>
<fig id="F6">
<caption>
<p><bold>Figure 6:</bold> Phone duration by word position, phrase position, and gesture presence.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-17-18890-g6.png"/>
</fig>
<table-wrap id="T5">
<caption>
<p><bold>Table 5:</bold> Influence of phrase position, word position, and gesture on phone duration.</p>
</caption>
<table>
<tbody>
<tr>
<td align="left" valign="top"><bold>Term</bold></td>
<td align="center" valign="top"><bold><italic>&#946;</italic></bold></td>
<td align="center" valign="top"><bold><italic>t-value</italic></bold></td>
<td align="center" valign="top"><bold><italic>p-value</italic></bold></td>
</tr>
<tr>
<td align="left" valign="top">(Intercept)</td>
<td align="right" valign="top">1.891</td>
<td align="right" valign="top">80.899</td>
<td align="left" valign="top">&lt;.001***</td>
</tr>
<tr>
<td align="left" valign="top">Phrase Position (Medial)</td>
<td align="right" valign="top">&#8211;0.038</td>
<td align="right" valign="top">&#8211;1.968</td>
<td align="left" valign="top">=.05</td>
</tr>
<tr>
<td align="left" valign="top">Phrase Position (Final)</td>
<td align="right" valign="top">0.073</td>
<td align="right" valign="top">3.381</td>
<td align="left" valign="top">&lt;.001***</td>
</tr>
<tr>
<td align="left" valign="top">Word Position (Medial)</td>
<td align="right" valign="top">&#8211;0.046</td>
<td align="right" valign="top">&#8211;1.668</td>
<td align="left" valign="top">=.10</td>
</tr>
<tr>
<td align="left" valign="top">Word Position (Final)</td>
<td align="right" valign="top">0.025</td>
<td align="right" valign="top">1.127</td>
<td align="left" valign="top">=.26</td>
</tr>
<tr>
<td align="left" valign="top">Gesture</td>
<td align="right" valign="top">0.132</td>
<td align="right" valign="top">5.686</td>
<td align="left" valign="top">&lt;.001***</td>
</tr>
<tr>
<td align="left" valign="top">Phrase Position (Medial) * Word Position (Medial)</td>
<td align="right" valign="top">0.036</td>
<td align="right" valign="top">1.144</td>
<td align="left" valign="top">=.26</td>
</tr>
<tr>
<td align="left" valign="top">Phrase Position (Final) * Word Position (Medial)</td>
<td align="right" valign="top">&#8211;0.027</td>
<td align="right" valign="top">&#8211;0.756</td>
<td align="left" valign="top">=.45</td>
</tr>
<tr>
<td align="left" valign="top">Phrase Position (Medial) * Word Position (Final)</td>
<td align="right" valign="top">&#8211;0.061</td>
<td align="right" valign="top">&#8211;2.449</td>
<td align="left" valign="top">&lt;.05*</td>
</tr>
<tr>
<td align="left" valign="top">Phrase Position (Final) * Word Position (Final)</td>
<td align="right" valign="top">0.018</td>
<td align="right" valign="top">0.667</td>
<td align="left" valign="top">=.51</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>3.2. Predictability Effects on Duration</title>
<p>We now turn to look at various predictability effects on word duration within our corpus. A summary of these results is provided in <xref ref-type="table" rid="T6">Table 6</xref>. Four of our predictability measures&#8212;forward lexical probability (<italic>p</italic> &lt; .05), forward V-to-V probability (<italic>p</italic> &lt; .001), backward V-to-V probability (<italic>p</italic> &lt; .05), and lexical frequency (<italic>p</italic> &lt; .001)&#8212;were significant predictors of word duration in the model. As has been found in previous work, all three variables were negatively related to word duration. There was also a significant interaction between forward and backward V-to-V probability in predicting word duration (<italic>p</italic> &lt; .001): specifically, word durations were even shorter where both forward and backward V-to-V probability were high. Finally, forward lexical probability and forward V-to-V probability both interacted with word frequency, but the effects were in opposite directions: word durations were shorter for words with both high forward V-to-V probability and high lexical frequency (<italic>p</italic> &lt; .05). However, word durations were only shorter for words with high forward lexical probability when lexical frequency was low (<italic>p</italic> &lt; .001). This latter finding may be attributable to differences between content and function words in terms of how predictability modulates word duration (most content words being less frequent, on the whole, than function words). For example, Tang and Bennett (<xref ref-type="bibr" rid="B82">2018</xref>) find that forward lexical bigram predictability was associated with shorter word durations for content words, but not for function words, in Kaqchikel Mayan. Such findings differ from prior work on English by Bell et al. (<xref ref-type="bibr" rid="B5">2009</xref>) who found a greater effect of backward lexical predictability for both content and function words. Tang and Bennett attribute their findings to differences in the morphological complexity of Kaqchikel words; we note that Igbo, like Kaqchikel, has a more agglutinating word structure than English, featuring a rich set of both inflectional and derivational affixes. Despite this typological variation, our findings in Igbo are generally consistent with cross-linguistic findings that greater predictability leads to shorter word durations.</p>
<table-wrap id="T6">
<caption>
<p><bold>Table 6:</bold> Influence of predictability measures on word duration.</p>
</caption>
<table>
<tbody>
<tr>
<td align="left" valign="top"><bold>Term</bold></td>
<td align="center" valign="top"><bold><italic>&#946;</italic></bold></td>
<td align="center" valign="top"><bold><italic>t-value</italic></bold></td>
<td align="center" valign="top"><bold><italic>p-value</italic></bold></td>
</tr>
<tr>
<td align="left" valign="top">(Intercept)</td>
<td align="right" valign="top">&#8211;1.058</td>
<td align="right" valign="top">&#8211;17.759</td>
<td align="left" valign="top">&lt;.001</td>
</tr>
<tr>
<td align="left" valign="top">logBackVProb</td>
<td align="right" valign="top">&#8211;0.099</td>
<td align="right" valign="top">&#8211;2.512</td>
<td align="left" valign="top">&lt;.05*</td>
</tr>
<tr>
<td align="left" valign="top">logFwdVProb</td>
<td align="right" valign="top">&#8211;0.214</td>
<td align="right" valign="top">&#8211;3.684</td>
<td align="left" valign="top">&lt;.001***</td>
</tr>
<tr>
<td align="left" valign="top">logBckLexProb</td>
<td align="right" valign="top">0.028</td>
<td align="right" valign="top">1.173</td>
<td align="left" valign="top">=.24</td>
</tr>
<tr>
<td align="left" valign="top">logFwdLexProb</td>
<td align="right" valign="top">&#8211;0.078</td>
<td align="right" valign="top">&#8211;2.012</td>
<td align="left" valign="top">&lt;.05*</td>
</tr>
<tr>
<td align="left" valign="top">logLexFreq</td>
<td align="right" valign="top">&#8211;0.125</td>
<td align="right" valign="top">&#8211;3.379</td>
<td align="left" valign="top">&lt;.001***</td>
</tr>
<tr>
<td align="left" valign="top">Syllables</td>
<td align="right" valign="top">0.233</td>
<td align="right" valign="top">3.910</td>
<td align="left" valign="top">&lt;.001***</td>
</tr>
<tr>
<td align="left" valign="top">logBackVProb:logLexFreq</td>
<td align="right" valign="top">0.050</td>
<td align="right" valign="top">1.302</td>
<td align="left" valign="top">=.19</td>
</tr>
<tr>
<td align="left" valign="top">logFwdVProb: logLexFreq</td>
<td align="right" valign="top">&#8211;0.097</td>
<td align="right" valign="top">&#8211;1.993</td>
<td align="left" valign="top">&lt;.05*</td>
</tr>
<tr>
<td align="left" valign="top">logBckLexProb:logLexFreq</td>
<td align="right" valign="top">&#8211;0.041</td>
<td align="right" valign="top">&#8211;1.160</td>
<td align="left" valign="top">=.25</td>
</tr>
<tr>
<td align="left" valign="top">logFwdLexProb:logLexFreq</td>
<td align="right" valign="top">0.057</td>
<td align="right" valign="top">2.187</td>
<td align="left" valign="top">&lt;.05*</td>
</tr>
<tr>
<td align="left" valign="top">logBckLexProb:logFwdLexProb</td>
<td align="right" valign="top">&#8211;0.018</td>
<td align="right" valign="top">&#8211;0.484</td>
<td align="left" valign="top">=.63</td>
</tr>
<tr>
<td align="left" valign="top">logBckVProb:logFwdVProb</td>
<td align="right" valign="top">&#8211;0.042</td>
<td align="right" valign="top">&#8211;2.437</td>
<td align="left" valign="top">=.01*</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>3.3. Prosodic Structure and Predictability as Factors in Co-Speech Gesture Timing</title>
<p>The final model we present aims to examine how predictability effects influence co-speech gesture timing, alongside factors already known to influence gesture timing, including phrase position and tonal melody. Due to the large number of variables to be tested in the analysis and the potential for collinearity effects, we use forward nested model comparison to determine which factors contribute significantly to model fit. Starting with a model including only phrase position, tone melody, and word duration (number of syllables was omitted since this information is already coded in tone melody, e.g., HH vs. HHH), we added in each predictability measure to check for significant differences in model fit, measured through Chi-square tests of goodness of fit. Including lexical frequency did not lead to a significantly better fit of the model (<italic>&#967;</italic><sup>2</sup>(478,1) = 1.854, <italic>p</italic> = .17), nor did inclusion of forward lexical probability (<italic>&#967;</italic><sup>2</sup>(478,1) = 2.308, <italic>p</italic> = .13) or backward lexical probability (<italic>&#967;</italic><sup>2</sup>(478,1) = 2.733, <italic>p</italic> = .10). Adding in forward V-to-V predictability did improve model fit (<italic>&#967;</italic><sup>2</sup>(478,1) = 4.766, <italic>p</italic> &lt; .05), as did adding in backward V-to-V predictability (<italic>&#967;</italic><sup>2</sup>(478,1) = 12.045, <italic>p</italic> &lt; .001). However, a model including both terms did not provide a significantly better fit than a model with only backward V-to-V predictability (<italic>&#967;</italic><sup>2</sup>(478,1) = 0.404, <italic>p</italic> = .53); thus, forward V-to-V predictability was removed from the model. Model fit was significantly improved with the addition of an interaction between tone melody and backward V-to-V predictability (<italic>&#967;</italic><sup>2</sup>(478,14) = 31.820, <italic>p</italic> &lt; .01), but not between phrase position and backward V-to-V predictability (<italic>&#967;</italic><sup>2</sup>(478,2) = 0.504, <italic>p</italic> = .78). In our last two model comparisons, we incorporated a) a two-way interaction between phrase position and word duration; and b) a two-way interaction between backward V-to-V probability and word duration in predicting gesture timing. Neither the first interaction (<italic>&#967;</italic><sup>2</sup>(478,2) = 0.135, <italic>p</italic> = .93) nor the second interaction (<italic>&#967;</italic><sup>2</sup>(478,1) = 0.740, <italic>p</italic> = .39) improved model fit.</p>
<p>Final model results are presented in <xref ref-type="table" rid="T7">Table 7</xref>. First off, as expected, both phrase position and tone melody are significant predictors of gesture timing (<italic>p</italic> &lt; .001). Post hoc comparisons indicated that gestures were less likely to target word-final position when a word occurred phrase-medially vs. phrase-finally (<italic>&#946;</italic> = &#8211;0.828, <italic>z</italic> = &#8211;3.025, <italic>p</italic> &lt; .01) or phrase-initially (<italic>&#946;</italic> = 1.261, <italic>z</italic> = 2.594, <italic>p</italic> &lt; .05). As found by Franich et al. (2025), gestures are also less likely to target word-final position in words containing a single H tone span, as in a disyllabic HH or trisyllabic HHH melody. For example, gestures were significantly more likely to occur nonfinally in HH words compared with HL words (<italic>&#946;</italic> = 1.436, <italic>z</italic> = 3.694, <italic>p</italic> &lt; .05). Backward V-to-V probability was a significant predictor of gesture timing, with gestures less likely to target word-final vowels when this predictability measure was high. Interestingly, the significant interaction between backward V-to-V probability and tone melody indicated that gestures were even more likely to occur in nonfinal position when a word had <italic>both</italic> a HH sequence <italic>and</italic> had high backward V-to-V probability. Post hoc comparisons indicated a significantly less negative slope in the relationship between word-finality and backward V-to-V probability for HL words compared with HH words (<italic>&#946;</italic> = 1.151, <italic>z</italic> = 3.113, <italic>p</italic> &lt; .05), while differences in slopes were not significant between HL and LH words (<italic>&#946;</italic> = 0.335, <italic>z</italic> = 0.946, <italic>p</italic> = .99) or between HL and LL words (<italic>&#946;</italic> = 0.842, <italic>z</italic> = 1.054, <italic>p</italic> = .99). Though no significant differences emerged for slopes between different trisyllabic melodies, HHH words displayed numerically more negative slopes than HLH, HHL, and LHH melodies (the three other most common trisyllabic tone melodies in the dataset).</p>
<table-wrap id="T7">
<caption>
<p><bold>Table 7:</bold> Influence of phrase position, tone melody, word duration, and backward V-to-V probability on gesture alignment within a word (final vs. nonfinal position). Categorical variables of Phrase Position and Tone Melody are contrast-coded with reference levels of Phrase Position (Final) and Tone Melody (HH), respectively.</p>
</caption>
<table>
<tbody>
<tr>
<td align="left" valign="top"><bold>Term</bold></td>
<td align="center" valign="top"><bold><italic>&#946;</italic></bold></td>
<td align="center" valign="top"><bold><italic>z-value</italic></bold></td>
<td align="center" valign="top"><bold><italic>p-value</italic></bold></td>
</tr>
<tr>
<td align="left" valign="top">(Intercept)</td>
<td align="right" valign="top">0.221</td>
<td align="right" valign="top">0.227</td>
<td align="left" valign="top">=.82</td>
</tr>
<tr>
<td align="left" valign="top">log Word Duration</td>
<td align="right" valign="top">0.170</td>
<td align="right" valign="top">1.346</td>
<td align="left" valign="top">=.18</td>
</tr>
<tr>
<td align="left" valign="top">logBackVProb</td>
<td align="right" valign="top">&#8211;0.964</td>
<td align="right" valign="top">&#8211;3.463</td>
<td align="left" valign="top">&lt;.001**</td>
</tr>
<tr>
<td align="left" valign="top">Phrase Position (Initial)</td>
<td align="right" valign="top">0.433</td>
<td align="right" valign="top">0.841</td>
<td align="left" valign="top">=.400</td>
</tr>
<tr>
<td align="left" valign="top">Phrase Position (Medial)</td>
<td align="right" valign="top">&#8211;0.828</td>
<td align="right" valign="top">&#8211;3.025</td>
<td align="left" valign="top">&lt;.001***</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (HL)</td>
<td align="right" valign="top">1.428</td>
<td align="right" valign="top">3.678</td>
<td align="left" valign="top">&lt;.001***</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (LH)</td>
<td align="right" valign="top">1.139</td>
<td align="right" valign="top">3.058</td>
<td align="left" valign="top">&lt;.01**</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (LL)</td>
<td align="right" valign="top">0.968</td>
<td align="right" valign="top">1.447</td>
<td align="left" valign="top">=.15</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (HHH)</td>
<td align="right" valign="top">0.636</td>
<td align="right" valign="top">0.564</td>
<td align="left" valign="top">=.57</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (HLH)</td>
<td align="right" valign="top">&#8211;0.299</td>
<td align="right" valign="top">&#8211;0.289</td>
<td align="left" valign="top">=.77</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (HLL)</td>
<td align="right" valign="top">1.259</td>
<td align="right" valign="top">0.670</td>
<td align="left" valign="top">=.50</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (HHL)</td>
<td align="right" valign="top">0.206</td>
<td align="right" valign="top">0.200</td>
<td align="left" valign="top">=.84</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (LHH)</td>
<td align="right" valign="top">&#8211;0.574</td>
<td align="right" valign="top">&#8211;0.506</td>
<td align="left" valign="top">=.61</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (LHL)</td>
<td align="right" valign="top">0.440</td>
<td align="right" valign="top">0.321</td>
<td align="left" valign="top">=.75</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (LLH)</td>
<td align="right" valign="top">1.938</td>
<td align="right" valign="top">1.225</td>
<td align="left" valign="top">=.22</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (H!HH)</td>
<td align="right" valign="top">&#8211;0.513</td>
<td align="right" valign="top">&#8211;0.466</td>
<td align="left" valign="top">=.64</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (H!HL)</td>
<td align="right" valign="top">&#8211;5.310</td>
<td align="right" valign="top">&#8211;0.725</td>
<td align="left" valign="top">=.47</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (HL) * logBackVProb</td>
<td align="right" valign="top">1.134</td>
<td align="right" valign="top">3.055</td>
<td align="left" valign="top">&lt;.01**</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (LH) * logBackVProb</td>
<td align="right" valign="top">0.864</td>
<td align="right" valign="top">2.274</td>
<td align="left" valign="top">&lt;.05*</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (LL) * logBackVProb</td>
<td align="right" valign="top">0.422</td>
<td align="right" valign="top">0.524</td>
<td align="left" valign="top">=.60</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (HHH) * logBackVProb</td>
<td align="right" valign="top">&#8211;1.169</td>
<td align="right" valign="top">&#8211;1.178</td>
<td align="left" valign="top">=.23</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (HLH) * logBackVProb</td>
<td align="right" valign="top">1.040</td>
<td align="right" valign="top">1.946</td>
<td align="left" valign="top">=.05</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (HLL) * logBackVProb</td>
<td align="right" valign="top">&#8211;0.878</td>
<td align="right" valign="top">&#8211;0.560</td>
<td align="left" valign="top">=.58</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (HHL) * logBackVProb</td>
<td align="right" valign="top">0.900</td>
<td align="right" valign="top">1.873</td>
<td align="left" valign="top">=.06</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (LHH) * logBackVProb</td>
<td align="right" valign="top">1.804</td>
<td align="right" valign="top">1.891</td>
<td align="left" valign="top">=.06</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (LHL) * logBackVProb</td>
<td align="right" valign="top">&#8211;0.046</td>
<td align="right" valign="top">&#8211;0.042</td>
<td align="left" valign="top">=.97</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (LLH) * logBackVProb</td>
<td align="right" valign="top">&#8211;1.186</td>
<td align="right" valign="top">&#8211;0.916</td>
<td align="left" valign="top">=.36</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (H!HH) * logBackVProb</td>
<td align="right" valign="top">0.526</td>
<td align="right" valign="top">0.815</td>
<td align="left" valign="top">=.42</td>
</tr>
<tr>
<td align="left" valign="top">Tone Melody (H!HL) * logBackVProb</td>
<td align="right" valign="top">&#8211;16.181</td>
<td align="right" valign="top">&#8211;0.745</td>
<td align="left" valign="top">=.46</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="fig" rid="F7">Figure 7</xref> demonstrates the interaction of V-to-V probability and tone melody for disyllabic words. As can be seen, there was a stronger negative relationship between word position and backward V-to-V probability for HH words compared with words bearing other tone melodies. <xref ref-type="fig" rid="F8">Figure 8</xref>, in contrast, presents the data with gesture position on the <italic>x</italic> axis and backward V-to-V probability on the <italic>y</italic> axis.</p>
<fig id="F7">
<caption>
<p><bold>Figure 7:</bold> Influence of Backward V-to-V Probability on gesture position, disyl. words.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-17-18890-g7.png"/>
</fig>
<fig id="F8">
<caption>
<p><bold>Figure 8:</bold> Influence of gesture position on Backward V-to-V Probability, disyl. words.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-17-18890-g8.png"/>
</fig>
<p>In sum, though our results in Section 3.1 support the idea of prosodic restructuring and gesture shift in phrase-medial position (where words are shorter and segments more reduced), results from Sections 3.2 and 3.3 do not provide support for the duration reduction hypothesis, as not all variables that influence word durations also impact gesture position. Notably, lexical frequency and forward lexical contextual predictability both had a negative influence on word durations; however, these factors did not play a role in influencing co-speech gesture timing. Furthermore, while backward vowel co-occurrence probability did influence both segment durations and gesture timing, our prediction of cumulative effects between predictability and phrase position was not borne out. More specifically, we anticipated an interaction between phrase position and V-to-V probability in predicting gesture timing, but this was not found, suggesting that these two factors influence gesture timing independently. Furthermore, taking a median split of the data for our backward V-to-V probability measure, it was not the case that words with higher V-to-V probability displayed nonword-final gestures more frequently in phrase-medial position. In fact, it was phrase-initial position that showed the greatest rate of nonword-final gestures among those words with the highest backward V-to-V probability. We now move on to discuss the implications of these results for models of co-speech gesture timing and speech planning.</p>
</sec>
</sec>
<sec>
<title>4. Discussion</title>
<p>The goal of this work is to understand the factors that influence the timing of co-speech gestures in Igbo, and how patterns of co-speech gesture may inform our understanding of the speech planning process more broadly. Our earlier work indicated that co-speech gestures in Igbo were more variable in their timing in phrase-medial words (Franich et al. 2025). A possible explanation for this effect is the fact that word-final syllables, the syllables to which gestures typically align in Igbo, undergo more extensive phonetic reduction and articulatory overlap in phrase-medial position, leading to a form of prosodic restructuring and a corresponding &#8216;gesture shift.&#8217; By hypothesis, any variable that could lead to this type of phonetic reduction could, in principle, influence the timing of co-speech gestures&#8212;we termed this the <italic>duration reduction hypothesis</italic> of gesture timing. To test this hypothesis, we investigated whether various measures of predictability&#8212;lexical frequency, lexical contextual probability, and vowel-to-vowel co-occurrence probability&#8212;influenced co-speech gesture timing, given that these measures of predictability tend to correlate with phonetic reduction patterns cross-linguistically (<xref ref-type="bibr" rid="B6">Bell et al., 2003</xref>; <xref ref-type="bibr" rid="B9">Bybee, 2006</xref>; <xref ref-type="bibr" rid="B20">Ernestus, 2011</xref>; <xref ref-type="bibr" rid="B61">Pierrehumbert, 2001</xref>; <xref ref-type="bibr" rid="B82">Tang &amp; Bennett 2018</xref>). Specifically, we proposed that words with higher predictability would be produced with shorter durations, leading them to undergo prosodic restructuring similar to the kind discussed by Keating &amp; Shattuck-Hufnagel (<xref ref-type="bibr" rid="B39">2002</xref>), in which default prominence within a word shifts in response to certain variables during the course of speech planning. In our proposal, restructuring leads not only to a shift in prominence location, but also to a concomitant shift in manual gesture timing.</p>
<p>Our findings on word and phone durations support the idea that phrase-medial position is subject to a greater degree of phonetic reduction of words compared with other prosodic positions. Furthermore, when examined in isolation, word duration was a significant predictor of gesture location in Igbo, with words occurring more frequently in word-final position when words were longer. However, several of our results undermine the idea of a direct relationship between duration reduction and gesture shift. Most notably, though phrase-medial position and higher predictability were both found to be linked to durational shortening, the relationship between predictability and gesture timing was less straightforward than the duration reduction hypothesis would have predicted. In particular, we find no influence of lexical contextual predictability effects on gesture timing, though lexical contextual predictability did condition word durations in our corpus (in line with findings from previous languages, e.g., <xref ref-type="bibr" rid="B6">Bell et al., 2003</xref>; <xref ref-type="bibr" rid="B82">Tang &amp; Bennett, 2018</xref>). Meanwhile, vowel co-occurrence probability did influence both word durations and gesture timing. However, contrary to predictions from the duration reduction hypothesis, words with high backward vowel co-occurrence probability did not display a greater degree of gesture shift in phrase-medial position compared with other phrasal positions. Instead, it appears that phrase position and vowel co-occurrence probability operate independently in conditioning co-speech gesture timing.</p>
<p>We propose that apparent effects of phonetic reduction on gesture timing are better seen as stemming from a prosodic planning component that is shaped by prosodically conditioned contextual variation, but that is not directly sensitive to durational variations. As such, manual gestures are shifted in prosodic environments that tend to be accompanied by phonetic reduction, but phonetic reduction itself does not play a role in conditioning the speaker/gesturer&#8217;s planning of gesture timing. One recent model of speech planning that could capture such patterns is the XT/3C model of speech planning by Turk and Shattuck-Hufnagel (2020), which has similarities to Keating &amp; Shattuck-Hufnagel&#8217;s (<xref ref-type="bibr" rid="B39">2002</xref>) model of prosodic planning. A key similarity between the two models is that prosodic structure is planned relatively early. Turk and Shattuck-Hufnagel propose that the first stage of speech production planning takes place in the Phonological Planning Component, which is responsible for planning prosodic structure (<xref ref-type="fig" rid="F9">Figure 9</xref>). Prosody itself is shaped by usage-based factors such as predictability, as in Aylett and Turk&#8217;s (<xref ref-type="bibr" rid="B3">2004</xref>) smooth signal redundancy framework. However, despite the incorporation of relative speech rate and speech style into this component, the authors argue that this stage of processing is entirely abstract and symbolic, and does not contain &#8220;specific spectral, spatial, and/or temporal information&#8221; (p. 151). As such, only those predictability effects that have been incorporated into the prosodic grammar are able to be planned at this stage. It is only at a later stage, in the Phonetic Planning Component, that quantitative targets are matched to abstract phonetic/gestural goals. Furthermore, achievement of these quantitative targets in the Phonetic Planning Component is mediated by competing constraints between achieving the specified quantitative target and minimizing the motor costs associated with achievement of those targets. Thus, the Phonetic Component provides additional opportunity for phonetic reduction to arise, even if it is not specified as a part of an acoustic target within the Phonological Component.</p>
<fig id="F9">
<caption>
<p><bold>Figure 9:</bold> XT/3C model of speech planning by Turk and Shattuck-Hufnagel (2020), with co-speech gesture incorporated.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-17-18890-g9.png"/>
</fig>
<p>We propose that co-speech gestures are planned alongside other aspects of prosodic structure and phonology during the more abstract Phonological Planning Component stage. Given the relatively early stage at which prosodic planning takes place, prosodic restructuring and associated co-speech gesture shift are only sensitive to abstract units, rather than to quantitative acoustic targets with continuous durations. Nonetheless, since abstract prosodic structure is influenced in the model by factors such as utterance length and relative speech rate (which both have implications for phrase-medial reduction processes, especially when combined with prosodic phrase structure), the observation that manual gestures are avoided in just those prosodic environments where words tend to be reduced/coalesced can be explained straightforwardly. Meanwhile, the lack of effects of lexical frequency and lexical contextual predictability on co-speech gesture timing can be accounted for by assuming that these effects do not uniformly form a part of the prosodic grammar in Igbo. Instead, phonetic reduction on frequent and contextually predictable words could result from target undershoot within the Phonetic Planning Component, since reducing more predictable words incurs less motor cost, and generally has minimal impact on the robustness of communication. Ultimately, such effects could eventually be reinterpreted on the part of the listener as a part of the prosodic grammar itself (see <xref ref-type="bibr" rid="B3">Aylett and Turk, 2004</xref> p. 35&#8211;36 for discussion of how language use may influence prosodic structure in diachronic terms). To sum up, a key takeaway from our results is the observation that predictability effects are not synonymous with prosodic effects. To capture our results, models of speech planning must treat these effects distinctly, at least to some extent.</p>
<p>In the context of our account, the effects of vowel co-occurrence probability on gesture timing should be reinterpreted as an effect of prosody, rather than predictability, <italic>per se</italic>. Vowel co-occurrence probability is tied to the phonological process of vowel harmony in Igbo, since vowels that harmonize are more mutually predictable. It is therefore appropriate that this type of predictability would be incorporated as a part of Turk and Shattuck-Hufnagel&#8217;s Phonological Planning Component. As discussed previously, words in Igbo with higher backward vowel co-occurrence probability were more likely to have co-speech gestures occurring on nonfinal vowels. As evidence of a link between backward vowel co-occurrence probability and vowel harmony, we find that a third of the nonloan words with the lowest V-to-V probability are disharmonic words (words in which at least one vowel does not match in harmony), including <italic>b&#236;&#803;a&#803;&#768;ru&#769;te&#768;</italic> &#8216;arrive,&#8217; <italic>nw&#225;&#803;nn&#232;</italic> &#8216;sibling,&#8217; and <italic>bu&#769;kwa&#803;&#768;</italic> &#8216;carry.&#8217; Such words are overall quite rare in the language (<xref ref-type="bibr" rid="B78">Smolek 2010</xref>). We also find that the ten words with values on either side of the median V-to-V probability score contained vowels that were consistently [&#8211;ATR], while those with the highest V-to-V probability score in the corpus contained vowels that were consistently [+ATR]. This falls out from the fact that [&#8211;ATR] words are overall more common in Igbo, meaning that their vowel sequences will be overall less predictable (<xref ref-type="bibr" rid="B78">Smolek, 2010</xref>).</p>
<p>However, looking more closely at the words that had the highest V-to-V contextual probability, we find an interesting pattern: these words were not only harmonic words, but they contained at least two adjacent identical vowels. Examples of these words include <italic>mmi&#769;li&#768;</italic> &#8216;water,&#8217; <italic>&#236;z&#236;z&#236;</italic> &#8216;first&#8217;, <italic>&#243;w&#232;r&#232;</italic> &#8216;Owerri,&#8217; <italic>&#232;k&#232;ne&#769;</italic> &#8216;give thanks,&#8217; <italic>e&#769;be&#769;ke&#768;</italic> &#8216;his place&#8217; and <italic>a&#769;&#803;na&#768;&#803;</italic> &#8216;land,&#8217; and <italic>ma&#769;&#803;ra&#769;&#803;</italic> &#8216;they know.&#8217; Of course, Igbo does not require vowels within the domain of harmony to be identical in all features, only that they match in ATR value (See Section 1.4). Furthermore, we observed an interaction between backward V-to-V co-occurrence probability and tone melody in predicting gesture timing, reflecting that gestures were even less likely to occur in word-final position when the word had both a contextually predictable final vowel and also contained an HH sequence. In other words, the greater the featural identity across adjacent vowels and tones in a word, the more likely the gesture was to shift to an earlier position. It seems, therefore, that vowel co-occurrence probability not only captures harmonic patterns, but also captures a more specific pattern of <italic>phonological redundancy</italic> between vowels. In line with our information hypothesis (see Section 2.5), gestures appear to be targeting syllables with less redundant vowels and tones within our dataset. While we believe that patterns of redundancy are rooted in some of the same perceptual and articulatory mechanisms that give rise to harmony patterns more generally (e.g., <xref ref-type="bibr" rid="B80">Suomi, 1983</xref>; <xref ref-type="bibr" rid="B38">Kaun, 1995</xref>; <xref ref-type="bibr" rid="B25">Gafos and Benu&#353; 2006</xref>; <xref ref-type="bibr" rid="B42">Kimper, 2017</xref>), we ultimately see the observed relationship between co-speech gesture and phonological redundancy as prosodically-mediated. Specifically, in line with Turk and Shattuck-Hufnagel&#8217;s proposal, we argue that phonological redundancy serves to influence the prosodic grammar by shaping patterns of prosodic prominence. To give a more concrete example, we have argued previously that reference must be made to a &#8216;tonal foot&#8217; in order to capture co-speech gesture patterns in HH and HHH words (<xref ref-type="bibr" rid="B24">Franich et al., 2025</xref>). While these patterns bear the hallmark of phonological metrical structure (e.g., binary constraints on foot structure; see Section 1.4), redundancy may provide a functional motivation for the incorporation of such structures in the grammar. Results of the current work suggest that similar prosodic structures may be needed to capture gesture patterns in at least some words with phonologically redundant vowels in Igbo; future work will need to investigate this possibility in more detail.</p>
<p>McCollum (<xref ref-type="bibr" rid="B55">2020</xref>) finds that noninitial harmonizing vowels in Kyrgyz (where vowels harmonize for both backness and rounding) are subject to phonetic reduction relative to initial harmonizing vowels. Our results provide further evidence that harmony-based redundancy factors are relevant for speech planning in typologically unrelated languages which utilize vowel harmony. Ultimately, our findings may also speak to the language-specific nature of predictability effects. Previous work has argued that phonetic reduction related to phonotactic predictability is mediated by language-specific pressures to maintain phonological contrasts (<xref ref-type="bibr" rid="B16">Cohen Priva, 2017</xref>; <xref ref-type="bibr" rid="B82">Tang &amp; Bennett, 2018</xref>). We hypothesize, with McCollum (<xref ref-type="bibr" rid="B55">2020</xref>), that vowel-to-vowel co-occurrence probability should be especially active in shaping phonetic patterns in languages with vowel harmony, since predictability is especially high within vowel sequences for such languages. Future work will need to evaluate possible differences in the effects of vowel co-occurrence probability on speech production across languages with and without vowel harmony.</p>
<p>Finally, our findings on the influence of phonological redundancy on co-speech gesture timing underscore important links between predictability and informational factors in driving gesture occurrence and timing. Prior work illustrates that gestures are not only more likely to occur on prosodically prominent syllables, but also more likely to occur on words which are new within the discourse context (Im &amp; Baumann, 2021). Words bearing new information are usually also less predictable from the discourse context (<xref ref-type="bibr" rid="B22">Ferreira &amp; Lowder, 2016</xref>), just as disharmonic and nonredundant vowels are less predictable within a word in our data. We can conclude, then, that information-related constraints on co-speech gestures are present at multiple levels of linguistic structure: gestures are enacted where new information is being communicated, and their timing is further refined according to word-level prominence, which is conditioned, in part, by patterns of phonological redundancy within the word.</p>
</sec>
<sec>
<title>5. Conclusion</title>
<p>We have sought to understand the mechanisms that govern the timing of co-speech gestures in Igbo, a Niger-Congo language. Our findings contribute to a growing body of work suggesting that co-speech gestures are planned within the same component of grammar that is responsible for prosodic planning in speech production. In that vein, our results reveal that prosodic restructuring and accompanying gesture shift constitute one source of variability in co-speech gesture timing in Igbo. This process is prosodically controlled, and only indirectly sensitive to patterns of phonetic reduction. However, our results demonstrate that prosodic restructuring is only one source of variation in co-speech gesture timing in the language. Another source of variation comes from phonological redundancy effects, which, while indexed by vowel co-occurrence predictability (as well as tone melody), also constitute patterns around which abstract prosodic structure is based. These findings highlight the importance of information at the sub-word level in shaping prosodic patterning. Our findings also suggest that predictability effects, though important in shaping prosodic grammar, must be separated from true prosodic prominence asymmetries at the level of speech planning and production.</p>
</sec>
</body>
<back>
<sec>
<title>Data accessibility statement</title>
<p>Data for analyses presented in this paper can be accessed at the following link: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" ext-link-type="uri" xlink:href="https://osf.io/4dy3z/?view_only=40dc905eca294df1b256b7d0c20ff68d">https://osf.io/4dy3z/?view_only=40dc905eca294df1b256b7d0c20ff68d</ext-link>.</p>
</sec>
<sec>
<title>Ethics and consent</title>
<p>This research was granted ethics approval by the Institutional Review Board of Harvard University (Protocol IRB 22-1095).</p>
</sec>
<sec>
<title>Funding information</title>
<p>This work was supported by National Science Foundation Linguistics Program Grant No. BCS- 2018003 (PI: Kathryn Franich). The National Science Foundation does not necessarily endorse the ideas and claims in this research.</p>
</sec>
<sec>
<title>Acknowledgements</title>
<p>We are grateful to two anonymous reviewers and to Taehong Cho and Lisa Davidson for their extremely valuable feedback that has helped us to improve the paper immensely. We would like to thank the Igbo participants for their time in participating in this study and to Echezona Ozoigbo for assisting with data collection. Thanks, also, to Eliana Spradling for help with data coding and management of our coding team, and the many research assistants who contributed to the coding of the corpus for this study, including Aaron Arlanza, Clarissa Briasco-Stewart, Walter Dych, Jordan Mitchell, Luc De Nardi, Sam Lyczkowski, Vitor Lacerda Siqueira, and Lexi Williams. Thank you to Karee Garvin for assistance with coding coordination and scripts for data wrangling. Finally, we thank audiences at the 2023 Annual Conference on African Linguistics, the 2024 Annual Conference on Phonology, the 2024 Symposium Series on Multimodal Communication, Princeton, UMass Amherst, Rutgers, and Brown for helpful feedback and discussion. All mistakes are our own.</p>
</sec>
<sec>
<title>Competing interests</title>
<p>The authors have no competing interests to declare.</p>
</sec>
<sec>
<title>Author contributions</title>
<p>KF contributed to conceptualization of the study, data preparation, data analysis, and writing. VN contributed to conceptualization of the study, data collection, data preparation, and writing.</p>
</sec>
<fn-group>
<fn id="n1"><p>Zsiga points to several other factors influencing the likelihood of coalescence, such as the vowel qualities involved (coalescence is less likely when transitioning from a high vowel to a non-high vowel; see also Ihuini &amp; Kenstowicz <xref ref-type="bibr" rid="B33">1994</xref>) and their occurrence within the same phonological phrase.</p></fn>
<fn id="n2"><p>See Cohen Priva (<xref ref-type="bibr" rid="B14">2008</xref>) for further discussion of the relationship between predictability and informativity at the level of the syllable.</p></fn>
<fn id="n3"><p>See also Garvin et al. (<xref ref-type="bibr" rid="B29">2025</xref>) for evidence that gestures themselves can induce some level of durational lengthening.</p></fn>
</fn-group>
<ref-list>
<ref id="B1"><mixed-citation publication-type="journal"><string-name><surname>Al-Mozainy</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Bley-Vroman</surname>, <given-names>R.</given-names></string-name>, &amp; <string-name><surname>McCarthy</surname>, <given-names>J.</given-names></string-name> (<year>1985</year>). <article-title>Stress shift and metrical structure</article-title>. <source>Linguistic Inquiry</source>, <volume>16</volume>, <fpage>135</fpage>&#8211;<lpage>144</lpage>.</mixed-citation></ref>
<ref id="B2"><mixed-citation publication-type="journal"><string-name><surname>Ambrazaitis</surname>, <given-names>G.</given-names></string-name>, &amp; <string-name><surname>House</surname>, <given-names>D.</given-names></string-name> (<year>2022</year>) <article-title>Probing effects of lexical prosody on speech-gesture integration in prominence production by Swedish news presenters</article-title>. <source>Laboratory Phonology</source>, <volume>24</volume>(<issue>1</issue>). <pub-id pub-id-type="doi">10.16995/labphon.6430</pub-id></mixed-citation></ref>
<ref id="B3"><mixed-citation publication-type="journal"><string-name><surname>Aylett</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Turk</surname>, <given-names>A.</given-names></string-name> (<year>2004</year>). <article-title>The smooth signal redundancy hypothesis: A functional explanation for relationships between redundancy, prosodic prominence, and duration in spontaneous speech</article-title>. <source>Language and Speech</source>, <volume>47</volume>(<issue>1</issue>). <pub-id pub-id-type="doi">10.1177/00238309040470010201</pub-id></mixed-citation></ref>
<ref id="B4"><mixed-citation publication-type="journal"><string-name><surname>Baker</surname>, <given-names>R. E.</given-names></string-name>, &amp; <string-name><surname>Bradlow</surname>, <given-names>A. R.</given-names></string-name> (<year>2009</year>). <article-title>Variability in word duration as a function of probability, speech style, and prosody</article-title>. <source>Language and Speech</source>, <volume>52</volume>(<issue>4</issue>), <fpage>391</fpage>&#8211;<lpage>413</lpage>. PMID: 20121039; PMCID: PMC2841971. <pub-id pub-id-type="doi">10.1177/0023830909336575</pub-id></mixed-citation></ref>
<ref id="B5"><mixed-citation publication-type="journal"><string-name><surname>Bell</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Brenier</surname>, <given-names>J. M.</given-names></string-name>, <string-name><surname>Gregory</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Girand</surname>, <given-names>C.</given-names></string-name>, &amp; <string-name><surname>Jurafsky</surname>, <given-names>D.</given-names></string-name> (<year>2009</year>). <article-title>Predictability effects on durations of content and function words in conversational English</article-title>. <source>Journal of Memory and Language</source>, <volume>60</volume>(<issue>1</issue>), <fpage>92</fpage>&#8211;<lpage>111</lpage>. <pub-id pub-id-type="doi">10.1016/j.jml.2008.06.003</pub-id></mixed-citation></ref>
<ref id="B6"><mixed-citation publication-type="journal"><string-name><surname>Bell</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Jurafsky</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Fosler-Lussier</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Girand</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Gregory</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Gildea</surname>, <given-names>D.</given-names></string-name> (<year>2003</year>). <article-title>Effects of disfluencies, predictability, and utterance position on word form variation in English conversation</article-title>. <source>Journal of the Acoustical Society of America</source>, <volume>113</volume>, <fpage>1001</fpage>&#8211;<lpage>1024</lpage>.</mixed-citation></ref>
<ref id="B7"><mixed-citation publication-type="journal"><string-name><surname>Berkovits</surname>, <given-names>R.</given-names></string-name> (<year>1993</year>). <article-title>Utterance-final lengthening and the duration of final-stop closures</article-title>. <source>Journal of Phonetics</source>, <volume>21</volume>, <fpage>479</fpage>&#8211;<lpage>489</lpage>. <pub-id pub-id-type="doi">10.1016/S0095-4470(19)30231-1</pub-id></mixed-citation></ref>
<ref id="B8"><mixed-citation publication-type="journal"><string-name><surname>B&#246;gels</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name><surname>Torreira</surname>, <given-names>F.</given-names></string-name> (<year>2015</year>). <article-title>Listeners use intonational phrase boundaries to project turn ends in spoken interaction</article-title>. <source>Journal of Phonetics</source>, <volume>52</volume>, <fpage>46</fpage>&#8211;<lpage>57</lpage>. <pub-id pub-id-type="doi">10.1016/j.wocn.2015.04.004</pub-id></mixed-citation></ref>
<ref id="B9"><mixed-citation publication-type="journal"><string-name><surname>Bybee</surname>, <given-names>J. L.</given-names></string-name> (<year>2006</year>). <article-title>From Usage to Grammar: The Mind&#8217;s Response to Repetition</article-title>. <source>Language</source>, <volume>82</volume>(<issue>4</issue>), <fpage>711</fpage>&#8211;<lpage>733</lpage>. <pub-id pub-id-type="doi">10.1353/lan.2006.0186</pub-id></mixed-citation></ref>
<ref id="B10"><mixed-citation publication-type="journal"><string-name><surname>Casali</surname>, <given-names>R. F.</given-names></string-name> (<year>1997</year>). <article-title>Vowel elision in hiatus contexts: Which vowel goes?</article-title> <source>Language</source>, <volume>73</volume>(<issue>3</issue>), <fpage>493</fpage>&#8211;<lpage>533</lpage>. <pub-id pub-id-type="doi">10.2307/415882</pub-id></mixed-citation></ref>
<ref id="B11"><mixed-citation publication-type="journal"><string-name><surname>Chen</surname>, <given-names>S. F.</given-names></string-name>, &amp; <string-name><surname>Goodman</surname>, <given-names>J.</given-names></string-name> (<year>1999</year>). <article-title>An empirical study of smoothing techniques for language modeling</article-title>. <source>Computer Speech &amp; Language</source>, <volume>13</volume>(<issue>4</issue>), <fpage>359</fpage>&#8211;<lpage>393</lpage>. <pub-id pub-id-type="doi">10.1006/csla.1999.0128</pub-id></mixed-citation></ref>
<ref id="B12"><mixed-citation publication-type="book"><string-name><surname>Cho</surname>, <given-names>T.</given-names></string-name> (<year>2006</year>). <chapter-title>Manifestation of prosodic structure in articulatory variation: Evidence from lip kinematics in English</chapter-title>. <source>Laboratory Phonology</source>, <volume>8</volume>, <fpage>519</fpage>&#8211;<lpage>548</lpage>. Edited by <string-name><given-names>L.</given-names> <surname>Goldstein</surname></string-name>, <string-name><given-names>D. H.</given-names> <surname>Whalen</surname></string-name> and <string-name><given-names>C. T.</given-names> <surname>Best</surname></string-name>. <publisher-loc>Berlin, New York</publisher-loc>: <publisher-name>De Gruyter Mouton</publisher-name>. <pub-id pub-id-type="doi">10.1515/9783110197211.3.519</pub-id></mixed-citation></ref>
<ref id="B13"><mixed-citation publication-type="book"><string-name><surname>Clark</surname>, <given-names>M. M.</given-names></string-name> (<year>1990</year>): <source>The tonal system of Igbo</source>. <publisher-loc>Dordrecht</publisher-loc>, <publisher-name>Foris</publisher-name>.</mixed-citation></ref>
<ref id="B14"><mixed-citation publication-type="book"><string-name><surname>Cohen Priva</surname>, <given-names>U.</given-names></string-name> (<year>2008</year>). <chapter-title>Using information content to predict phone deletion</chapter-title>. In <string-name><given-names>N.</given-names> <surname>Abner</surname></string-name> &amp; <string-name><given-names>J.</given-names> <surname>Bishop</surname></string-name> (Eds.) <source>Proceedings of the 27th West Coast Conference on Formal Linguistics</source> (pp. <fpage>90</fpage>&#8211;<lpage>98</lpage>). <publisher-name>Cascadilla Proceedings Project</publisher-name>.</mixed-citation></ref>
<ref id="B15"><mixed-citation publication-type="journal"><string-name><surname>Cohen Priva</surname>, <given-names>U.</given-names></string-name> (<year>2015</year>). <article-title>Informativity affects consonant duration and deletion rates</article-title>. <source>Laboratory Phonology</source>, <volume>6</volume>(<issue>2</issue>). <pub-id pub-id-type="doi">10.1515/lp-2015-0008</pub-id></mixed-citation></ref>
<ref id="B16"><mixed-citation publication-type="journal"><string-name><surname>Cohen Priva</surname>, <given-names>U.</given-names></string-name> (<year>2017</year>). <article-title>Informativity and the actuation of lenition</article-title>. <source>Language</source>, <volume>93</volume>(<issue>3</issue>), <fpage>569</fpage>&#8211;<lpage>597</lpage>. <pub-id pub-id-type="doi">10.1353/lan.2017.0037</pub-id></mixed-citation></ref>
<ref id="B17"><mixed-citation publication-type="book"><string-name><surname>Downing</surname>, <given-names>L. J.</given-names></string-name>, &amp; <string-name><surname>Rialland</surname>, <given-names>A.</given-names></string-name> (<year>2017</year>). <source>Intonation in African Tone Languages</source>. <publisher-loc>Berlin, Boston</publisher-loc>: <publisher-name>De Gruyter Mouton</publisher-name>. <pub-id pub-id-type="doi">10.1515/9783110503524</pub-id></mixed-citation></ref>
<ref id="B18"><mixed-citation publication-type="journal"><string-name><surname>Dych</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Garvin</surname>, <given-names>K.</given-names></string-name>, &amp; <string-name><surname>Franich</surname>, <given-names>K.</given-names></string-name> (<year>2023</year>). <article-title>Comparing manual vs. semi-automated methods for the coding of co-speech gestures</article-title>. <source>Proceedings of the 20th International Congress of Phonetic Sciences (ICPhS)</source>, <fpage>4170</fpage>&#8211;<lpage>4174</lpage>.</mixed-citation></ref>
<ref id="B19"><mixed-citation publication-type="webpage"><collab>ELAN (Version 6.9) [Computer software].</collab> (<year>2024</year>). <article-title>Nijmegen: Max Planck Institute for Psycholinguistics, The Language Archive</article-title>. Retrieved from <uri>https://archive.mpi.nl/tla/elan</uri></mixed-citation></ref>
<ref id="B20"><mixed-citation publication-type="journal"><string-name><surname>Ernestus</surname>, <given-names>M.</given-names></string-name> (<year>2011</year>). <article-title>An introduction to reduced pronunciation variants</article-title>. <source>Journal of Phonetics</source>, <volume>39</volume>(<issue>3</issue>), <fpage>253</fpage>&#8211;<lpage>260</lpage>. <pub-id pub-id-type="doi">10.1016/S0095-4470(11)00055-6</pub-id></mixed-citation></ref>
<ref id="B21"><mixed-citation publication-type="journal"><string-name><surname>Esteve-Gibert</surname>, <given-names>N.</given-names></string-name>, &amp; <string-name><surname>Prieto</surname>, <given-names>P.</given-names></string-name> (<year>2013</year>). <article-title>Prosodic Structure Shapes the Temporal Realization of Intonation and Manual Gesture Movements</article-title>. <source>Journal of Speech, Language, and Hearing Research</source>, <volume>56</volume>(<issue>3</issue>), <fpage>850</fpage>&#8211;<lpage>864</lpage>. <pub-id pub-id-type="doi">10.1044/1092-4388(2012/12-0049</pub-id>)</mixed-citation></ref>
<ref id="B22"><mixed-citation publication-type="book"><string-name><surname>Ferreira</surname>, <given-names>F.</given-names></string-name>, &amp; <string-name><surname>Lowder</surname>, <given-names>M. W.</given-names></string-name> (<year>2016</year>). <chapter-title>Prediction, information structure, and good-enough language processing</chapter-title>. In <string-name><given-names>B. H.</given-names> <surname>Ross</surname></string-name> (Ed.), <source>The psychology of learning and motivation</source> (pp. <fpage>217</fpage>&#8211;<lpage>247</lpage>). <publisher-name>Elsevier Academic Press</publisher-name>.</mixed-citation></ref>
<ref id="B23"><mixed-citation publication-type="book"><string-name><surname>Franich</surname>, <given-names>K.</given-names></string-name> (<year>2024</year>). <chapter-title>The phonological status of stem-initial prominence in Grassfields Bantu (and beyond)</chapter-title>. <source>Plenary presentation at Second Annual Conference on Bantoid Languages</source>, <publisher-name>University of Yaound&#233; I</publisher-name>, <publisher-loc>Cameroon</publisher-loc>, <month>June</month> <day>6</day>.</mixed-citation></ref>
<ref id="B24"><mixed-citation publication-type="book"><string-name><surname>Franich</surname>, <given-names>K.</given-names></string-name> (<year>2025</year>). <chapter-title>Cross-linguistic variation in co-speech gesture timing: Evidence for grammatical control</chapter-title>. <source>Presentation at The International Society for Gesture Studies &#8211; Catalonia</source>, <publisher-loc>Spain</publisher-loc>, <month>November</month> <day>21</day>.</mixed-citation></ref>
<ref id="B25"><mixed-citation publication-type="journal"><string-name><surname>Gafos</surname>, <given-names>A. I.</given-names></string-name>, &amp; <string-name><surname>Be&#328;u&#353;</surname>, <given-names>&#352;.</given-names></string-name> (<year>2006</year>). <article-title>Dynamics of phonological cognition</article-title>. <source>Cognitive Science</source>, <volume>30</volume>, <fpage>905</fpage>&#8211;<lpage>943</lpage>. <pub-id pub-id-type="doi">10.1207/s15516709cog0000_80</pub-id></mixed-citation></ref>
<ref id="B26"><mixed-citation publication-type="journal"><string-name><surname>Gahl</surname>, <given-names>S.</given-names></string-name> (<year>2008</year>). <article-title>Time and Thyme Are not Homophones: The Effect of Lemma Frequency on Word Durations in Spontaneous Speech</article-title>. <source>Language</source>, <volume>84</volume>(<issue>3</issue>), <fpage>474</fpage>&#8211;<lpage>496</lpage>. <pub-id pub-id-type="doi">10.1353/lan.0.0035</pub-id></mixed-citation></ref>
<ref id="B27"><mixed-citation publication-type="journal"><string-name><surname>Gahl</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Yao</surname>, <given-names>Y.</given-names></string-name>, &amp; <string-name><surname>Johnson</surname>, <given-names>K.</given-names></string-name> (<year>2012</year>). <article-title>Why reduce? Phonological neighborhood density and phonetic reduction in spontaneous speech</article-title>. <source>Journal of Memory and Language</source>, <volume>66</volume>(<issue>4</issue>), <fpage>789</fpage>&#8211;<lpage>806</lpage>. <pub-id pub-id-type="doi">10.1016/j.jml.2011.11.006</pub-id></mixed-citation></ref>
<ref id="B28"><mixed-citation publication-type="journal"><string-name><surname>Garvin</surname>, <given-names>K.</given-names></string-name>, &amp; <string-name><surname>Franich</surname>, <given-names>K.</given-names></string-name> (<year>2023</year>). <article-title>Gestural alignment and accommodation in speaker-listener head gestures</article-title>. <source>Proceedings of the 20th International Congress of Phonetic Sciences (ICPhS)</source>, <fpage>4160</fpage>&#8211;<lpage>4164</lpage>.</mixed-citation></ref>
<ref id="B29"><mixed-citation publication-type="journal"><string-name><surname>Garvin</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Spradling</surname>, <given-names>E.</given-names></string-name>, &amp; <string-name><surname>Franich</surname>, <given-names>K.</given-names></string-name> (<year>2025</year>). <article-title>Co-speech gestures influence the magnitude and stability of articulatory movements: Evidence for coupling-based enhancement</article-title>. <source>Scientific Reports</source>, <volume>15</volume>(<issue>157</issue>). <pub-id pub-id-type="doi">10.1038/s41598-024-84097-6</pub-id></mixed-citation></ref>
<ref id="B30"><mixed-citation publication-type="webpage"><string-name><surname>Gherardi</surname>, <given-names>V.</given-names></string-name> (<year>2024</year>). <source>Kgrams package</source>. R package version 0.2.1. <uri>https://cran.r-project.org/web/packages/kgrams/kgrams.pdf</uri></mixed-citation></ref>
<ref id="B31"><mixed-citation publication-type="book"><string-name><surname>Gullberg</surname>, <given-names>M.</given-names></string-name> (<year>1998</year>). <source>Gesture as a communication strategy in second language discourse: A study of learners of French and Swedish</source>. <publisher-name>Lund University Press</publisher-name>.</mixed-citation></ref>
<ref id="B32"><mixed-citation publication-type="book"><string-name><surname>Halle</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Vergnaud</surname>, <given-names>J.-R.</given-names></string-name> (<year>1987</year>). <source>An essay on stress</source>. (Current Studies in Linguistics 15). <publisher-name>MIT Press</publisher-name>.</mixed-citation></ref>
<ref id="B33"><mixed-citation publication-type="book"><string-name><surname>Ihiunu</surname>, <given-names>P.</given-names></string-name>, &amp; <string-name><surname>Kenstowicz</surname>, <given-names>M.</given-names></string-name> (<year>1994</year>). <chapter-title>Two notes on Igbo vowels [Unpublished manuscript]</chapter-title>. <publisher-name>MIT</publisher-name>.</mixed-citation></ref>
<ref id="B34"><mixed-citation publication-type="journal"><string-name><surname>Im</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name><surname>Baumann</surname>, <given-names>S.</given-names></string-name> (<year>2020</year>). <article-title>Probabilistic relation between co-speech gestures, pitch accents and information status</article-title>. <source>Proceedings of the Linguistic Society of America</source>, <volume>5</volume>(<issue>1</issue>). <pub-id pub-id-type="doi">10.3765/plsa.v5i1.4755</pub-id></mixed-citation></ref>
<ref id="B35"><mixed-citation publication-type="journal"><string-name><surname>Jenkins</surname>, <given-names>T.</given-names></string-name>, &amp; <string-name><surname>Pouw</surname>, <given-names>W.</given-names></string-name> (<year>2023</year>). <article-title>Gesture&#8211;speech coupling in persons with aphasia: A kinematic-acoustic analysis</article-title>. <source>Journal of Experimental Psychology: General</source>, <volume>152</volume>(<issue>5</issue>), <fpage>1469</fpage>&#8211;<lpage>1483</lpage>. <pub-id pub-id-type="doi">10.1037/xge0001346</pub-id></mixed-citation></ref>
<ref id="B36"><mixed-citation publication-type="book"><string-name><surname>Jurafsky</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Bell</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Gregory</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Raymond</surname>, <given-names>W. D.</given-names></string-name> (<year>2001</year>). <chapter-title>Probabilistic relations between words: Evidence from reduction in lexical production</chapter-title>. In <string-name><given-names>J.</given-names> <surname>Bybee</surname></string-name> &amp; <string-name><given-names>P.</given-names> <surname>Hopper</surname></string-name> (Eds.), <source>Frequency and the Emergence of Linguistic Structure</source> (Typological Studies in Language) (pp. <fpage>229</fpage>&#8211;<lpage>254</lpage>). <publisher-name>John Benjamins</publisher-name>.</mixed-citation></ref>
<ref id="B37"><mixed-citation publication-type="journal"><string-name><surname>Katsika</surname>, <given-names>A.</given-names></string-name> (<year>2016</year>). <article-title>The role of prominence in determining the scope of boundary-related lengthening in Greek</article-title>. <source>Journal of Phonetics</source>, <volume>55</volume>, <fpage>149</fpage>&#8211;<lpage>181</lpage>. <pub-id pub-id-type="doi">10.1016/j.wocn.2015.12.003</pub-id></mixed-citation></ref>
<ref id="B38"><mixed-citation publication-type="thesis"><string-name><surname>Kaun</surname>, <given-names>A.</given-names></string-name> (<year>1995</year>). <source>The typology of rounding harmony: An Optimality Theoretic approach</source> [Doctoral dissertation, <publisher-name>UCLA</publisher-name>]. ProQuest Dissertations &amp; Theses Global.</mixed-citation></ref>
<ref id="B39"><mixed-citation publication-type="journal"><string-name><surname>Keating</surname>, <given-names>P.</given-names></string-name>, &amp; <string-name><surname>Shattuck-Hufnagel</surname>, <given-names>S.</given-names></string-name> (<year>2002</year>). <article-title>A prosodic view of word form encoding</article-title>. <source>UCLA Working Papers in Phonology</source>, <fpage>112</fpage>&#8211;<lpage>156</lpage>.</mixed-citation></ref>
<ref id="B40"><mixed-citation publication-type="book"><string-name><surname>Kendon</surname>, <given-names>A.</given-names></string-name> (<year>1980</year>). <chapter-title>Gesticulation and speech: Two aspects of the process of utterance</chapter-title>. In <string-name><given-names>M. R.</given-names> <surname>Key</surname></string-name> (Ed.), <source>The relationship of verbal and nonverbal communication</source> (pp. <fpage>207</fpage>&#8211;<lpage>228</lpage>). <publisher-name>De Gruyter Mouton</publisher-name>. <pub-id pub-id-type="doi">10.1515/9783110813098.207</pub-id></mixed-citation></ref>
<ref id="B41"><mixed-citation publication-type="journal"><string-name><surname>Kim</surname>, <given-names>J. J.</given-names></string-name>, <string-name><surname>Kim</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name><surname>Cho</surname>, <given-names>T.</given-names></string-name> (<year>2024</year>). <article-title>Preboundary lengthening and articulatory strengthening in Korean as an edge-prominence language</article-title>. <source>Laboratory Phonology</source>, <volume>15</volume>(<issue>1</issue>). <pub-id pub-id-type="doi">10.16995/labphon.9880</pub-id></mixed-citation></ref>
<ref id="B42"><mixed-citation publication-type="journal"><string-name><surname>Kimper</surname>, <given-names>W.</given-names></string-name> (<year>2017</year>). <article-title>Not crazy after all these years? Perceptual grounding for long-distance vowel harmony</article-title>. <source>Laboratory Phonology: Journal of the Association for Laboratory Phonology</source>, <volume>8</volume>(<issue>1</issue>), <elocation-id>19</elocation-id>. <pub-id pub-id-type="doi">10.5334/labphon.47</pub-id></mixed-citation></ref>
<ref id="B43"><mixed-citation publication-type="journal"><string-name><surname>Krivokapi&#263;</surname>, <given-names>J.</given-names></string-name> (<year>2014</year>). <article-title>Gestural coordination at prosodic boundaries and its role for prosodic structure and speech planning processes</article-title>. <source>Philosophical Transactions of the Royal Society B: Biological Sciences</source>, <volume>369</volume>(<issue>1658</issue>), <elocation-id>20130397</elocation-id>.</mixed-citation></ref>
<ref id="B44"><mixed-citation publication-type="journal"><string-name><surname>Krivokapi&#263;</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Tiede</surname>, <given-names>M. K.</given-names></string-name>, &amp; <string-name><surname>Tyrone</surname>, <given-names>M. E.</given-names></string-name> (<year>2017</year>). <article-title>A Kinematic Study of Prosodic Structure in Articulatory and Manual Gestures: Results from a Novel Method of Data Collection</article-title>. <source>Laboratory Phonology</source>, <volume>8</volume>(<issue>1</issue>), Article <elocation-id>1</elocation-id>. <pub-id pub-id-type="doi">10.5334/labphon.75</pub-id></mixed-citation></ref>
<ref id="B45"><mixed-citation publication-type="journal"><string-name><surname>Krug</surname>, <given-names>M.</given-names></string-name> (<year>1998</year>). <article-title>String Frequency: A Cognitive Motivating Factor in Coalescence, Language Processing, and Linguistic Change</article-title>. <source>Journal of English Linguistics</source>, <volume>26</volume>(<issue>4</issue>), <fpage>286</fpage>&#8211;<lpage>320</lpage>. <pub-id pub-id-type="doi">10.1177/007542429802600402</pub-id></mixed-citation></ref>
<ref id="B46"><mixed-citation publication-type="book"><string-name><surname>K&#252;gler</surname>, <given-names>F.</given-names></string-name> (<year>2017</year>). <chapter-title>Tone and intonation in Akan</chapter-title>. In <string-name><given-names>L. J.</given-names> <surname>Downing</surname></string-name> &amp; <string-name><given-names>A.</given-names> <surname>Rialland</surname></string-name> (Eds.), <source>Intonation in African tone languages</source> (pp. <fpage>89</fpage>&#8211;<lpage>130</lpage>). <publisher-name>De Gruyter Mouton</publisher-name>. <pub-id pub-id-type="doi">10.1515/9783110503524-004</pub-id></mixed-citation></ref>
<ref id="B47"><mixed-citation publication-type="journal"><string-name><surname>K&#252;gler</surname>, <given-names>F.</given-names></string-name>, &amp; <string-name><surname>Gregori</surname>, <given-names>A.</given-names></string-name> (<year>2023</year>). <article-title>Iconic gestures in focus &#8211; synchronization of prosody and gestures in prominence</article-title>. <source>Proceedings of the International Congress of Phonetic Sciences</source>, <fpage>4125</fpage>&#8211;<lpage>4129</lpage>.</mixed-citation></ref>
<ref id="B48"><mixed-citation publication-type="book"><string-name><surname>Kula</surname>, <given-names>N.</given-names></string-name>, &amp; <string-name><surname>Hamman</surname>, <given-names>S.</given-names></string-name> (<year>2017</year>). <chapter-title>Intonation in Bemba</chapter-title>. In <string-name><given-names>L. J.</given-names> <surname>Downing</surname></string-name> &amp; <string-name><given-names>A.</given-names> <surname>Rialland</surname></string-name> (Eds.), <source>Intonation in African tone languages</source> (pp. <fpage>321</fpage>&#8211;<lpage>364</lpage>). <publisher-name>De Gruyter Mouton</publisher-name>. <pub-id pub-id-type="doi">10.1515/9783110503524-004</pub-id></mixed-citation></ref>
<ref id="B49"><mixed-citation publication-type="webpage"><string-name><surname>Lenth</surname>, <given-names>R. V.</given-names></string-name> (<year>2024</year>). <source>emmeans: Estimated Marginal Means, aka Least-Squares Means</source>. R package version 1.6.1. <uri>https://CRAN.R-project.org/package=emmeans</uri></mixed-citation></ref>
<ref id="B50"><mixed-citation publication-type="journal"><string-name><surname>Leonard</surname>, <given-names>T.</given-names></string-name>, &amp; <string-name><surname>Cummins</surname>, <given-names>F.</given-names></string-name> (<year>2011</year>). <article-title>The temporal relation between beat gestures and speech</article-title>. <source>Language and Cognitive Processes</source>, <volume>26</volume>(<issue>10</issue>). <pub-id pub-id-type="doi">10.1080/01690965.2010.500218</pub-id></mixed-citation></ref>
<ref id="B51"><mixed-citation publication-type="journal"><string-name><surname>Levy</surname>, <given-names>E. T.</given-names></string-name>, &amp; <string-name><surname>McNeill</surname>, <given-names>D.</given-names></string-name> (<year>1992</year>). <article-title>Speech, gesture, and discourse</article-title>. <source>Discourse Processes</source>, <volume>15</volume>(<issue>3</issue>), <fpage>277</fpage>&#8211;<lpage>301</lpage>. <pub-id pub-id-type="doi">10.1080/01638539209544813</pub-id></mixed-citation></ref>
<ref id="B52"><mixed-citation publication-type="book"><string-name><surname>Lieberman</surname>, <given-names>P.</given-names></string-name> (<year>1967</year>). <source>Intonation, perception, and language</source>. <publisher-name>MIT Press</publisher-name>.</mixed-citation></ref>
<ref id="B53"><mixed-citation publication-type="book"><string-name><surname>Lightner</surname>, <given-names>T. M.</given-names></string-name> (<year>1972</year>). <source>Problems in the theory of phonology: Russian phonology and Turkish phonology</source>. <publisher-name>Linguistic Research, Inc</publisher-name>.</mixed-citation></ref>
<ref id="B54"><mixed-citation publication-type="journal"><string-name><surname>Loehr</surname>, <given-names>D. P.</given-names></string-name> (<year>2012</year>). <article-title>Temporal, structural, and pragmatic synchrony between intonation and gesture</article-title>. <source>Laboratory Phonology</source>, <volume>3</volume>(<issue>1</issue>). <pub-id pub-id-type="doi">10.1515/lp-2012-0006</pub-id></mixed-citation></ref>
<ref id="B55"><mixed-citation publication-type="journal"><string-name><surname>McCollum</surname>, <given-names>A.</given-names></string-name> (<year>2020</year>). <article-title>Vowel harmony and positional variation in Kyrgyz</article-title>. <source>Laboratory Phonology</source> <volume>11</volume>(<issue>1</issue>), <elocation-id>25</elocation-id>. <pub-id pub-id-type="doi">10.5334/labphon.247</pub-id></mixed-citation></ref>
<ref id="B56"><mixed-citation publication-type="book"><string-name><surname>McNeill</surname>, <given-names>D.</given-names></string-name> (<year>1995</year>). <source>Hand and mind: What gestures reveal about thought</source>. <publisher-name>University of Chicago Press</publisher-name>.</mixed-citation></ref>
<ref id="B57"><mixed-citation publication-type="webpage"><collab>MIT Speech Communications Group.</collab> (<year>2020</year>). <source>MIT Speech Communications Group gesture studies coding manual</source>. Retrieved <month>June</month> 2020 from <uri>http://web.mit.edu/pelire/www/manual/</uri></mixed-citation></ref>
<ref id="B58"><mixed-citation publication-type="book"><string-name><surname>Mu&#241;oz-Coego</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Florit-Pons</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Rohrer</surname>, <given-names>P. L.</given-names></string-name>, <string-name><surname>Vil&#224;-Gim&#233;nez</surname>, <given-names>I.</given-names></string-name>, &amp; <string-name><surname>Prieto</surname>, <given-names>P.</given-names></string-name> (<year>2022</year>). <chapter-title>The prosodic and gestural marking of the information status of referents in children&#8217;s narrative speech: a longitudinal study</chapter-title>. <source>Proceedings of the 8th International Conference on Speech Prosody (ICPhS)</source>, <publisher-loc>Lisbon</publisher-loc>, <fpage>401</fpage>&#8211;<lpage>405</lpage>.</mixed-citation></ref>
<ref id="B59"><mixed-citation publication-type="journal"><string-name><surname>Paschen</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Fuchs</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name><surname>Seifart</surname>, <given-names>F.</given-names></string-name> (<year>2022</year>). <article-title>Final Lengthening and vowel length in 25 languages</article-title>. <source>Journal of Phonetics</source>, <volume>94</volume>, <elocation-id>101179</elocation-id>. <pub-id pub-id-type="doi">10.1016/j.wocn.2022.101179</pub-id></mixed-citation></ref>
<ref id="B60"><mixed-citation publication-type="thesis"><string-name><surname>Pierrehumbert</surname>, <given-names>J.</given-names></string-name> (<year>1980</year>). <source>The phonology and phonetics of English intonation</source> [Doctoral dissertation, <publisher-name>MIT</publisher-name>]. ProQuest Dissertations &amp; Theses Global.</mixed-citation></ref>
<ref id="B61"><mixed-citation publication-type="book"><string-name><surname>Pierrehumbert</surname>, <given-names>J.</given-names></string-name> (<year>2001</year>). <chapter-title>Probabilistic relations between words: Evidence from reduction in lexical production</chapter-title>. In <string-name><given-names>J.</given-names> <surname>Bybee</surname></string-name> &amp; <string-name><given-names>P.</given-names> <surname>Hopper</surname></string-name> (Eds.) <source>Frequency and the Emergence of Linguistic Structure</source> (Typological Studies in Language) (pp. <fpage>137</fpage>&#8211;<lpage>158</lpage>). <publisher-name>John Benjamins</publisher-name>.</mixed-citation></ref>
<ref id="B62"><mixed-citation publication-type="journal"><string-name><surname>Pouw</surname>, <given-names>W.</given-names></string-name>, &amp; <string-name><surname>Dixon</surname>, <given-names>J. A.</given-names></string-name> (<year>2019</year>). <article-title>Entrainment and Modulation of Gesture&#8211;Speech Synchrony Under Delayed Auditory Feedback</article-title>. <source>Cognitive Science</source>, <volume>43</volume>(<issue>3</issue>), <elocation-id>e12721</elocation-id>. <pub-id pub-id-type="doi">10.1111/cogs.12721</pub-id></mixed-citation></ref>
<ref id="B63"><mixed-citation publication-type="journal"><string-name><surname>Pouw</surname>, <given-names>W.</given-names></string-name>, &amp; <string-name><surname>Fuchs</surname>, <given-names>S.</given-names></string-name> (<year>2022</year>). <article-title>Origins of vocal-entangled gesture</article-title>. <source>Neuroscience &amp; Biobehavioral Reviews</source>, <volume>141</volume>, <elocation-id>104836</elocation-id>. <pub-id pub-id-type="doi">10.1016/j.neubiorev.2022.104836</pub-id></mixed-citation></ref>
<ref id="B64"><mixed-citation publication-type="journal"><string-name><surname>Pouw</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Harrison</surname>, <given-names>S. J.</given-names></string-name>, &amp; <string-name><surname>Dixon</surname>, <given-names>J. A.</given-names></string-name> (<year>2020</year>). <article-title>Gesture&#8211;speech physics: The biomechanical basis for the emergence of gesture&#8211;speech synchrony</article-title>. <source>Journal of Experimental Psychology: General</source>, <volume>149</volume>(<issue>2</issue>), Article <elocation-id>2</elocation-id>. <pub-id pub-id-type="doi">10.1037/xge0000646</pub-id></mixed-citation></ref>
<ref id="B65"><mixed-citation publication-type="webpage"><collab>R Core Team</collab>. (<year>2024</year>). <source>R: A language and environment for statistical computing</source>. <publisher-name>R Foundation for Statistical Computing</publisher-name>, <publisher-loc>Vienna, Austria</publisher-loc>. <uri>https://www.R-project.org/</uri></mixed-citation></ref>
<ref id="B66"><mixed-citation publication-type="journal"><string-name><surname>Rochet-Capellan</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Laboissi&#232;re</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Galv&#225;n</surname>, <given-names>A.</given-names></string-name>, &amp; <string-name><surname>Schwartz</surname>, <given-names>J.-L.</given-names></string-name> (<year>2008</year>). <article-title>The speech focus position effect on jaw&#8211;finger coordination in a pointing task</article-title>. <source>Journal of Speech, Language, and Hearing Research</source>, <volume>51</volume>(<issue>6</issue>). <pub-id pub-id-type="doi">10.1044/1092-4388(2008/07-0173)</pub-id></mixed-citation></ref>
<ref id="B67"><mixed-citation publication-type="journal"><string-name><surname>Rohrer</surname>, <given-names>P. L.</given-names></string-name>, <string-name><surname>Delais-Roussarie</surname>, <given-names>E.</given-names></string-name>, &amp; <string-name><surname>Prieto</surname>, <given-names>P.</given-names></string-name> (<year>2023</year>). <article-title>Visualizing prosodic structure: Manual gestures as highlighters of prosodic heads and edges in English academic discourses</article-title>. <source>Lingua</source>, <volume>293</volume>, <elocation-id>103583</elocation-id>. <pub-id pub-id-type="doi">10.1016/j.lingua.2023.103583</pub-id></mixed-citation></ref>
<ref id="B68"><mixed-citation publication-type="book"><string-name><surname>Rosenfelder</surname>, <given-names>I.</given-names></string-name>, <string-name><surname>Fruehwald</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Evanini</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Seyfarth</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Gorman</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Prichard</surname>, <given-names>H.</given-names></string-name>, &amp; <string-name><surname>Yuan</surname>, <given-names>J.</given-names></string-name> (<year>2022</year>). <source>Fave: Speaker Fix</source> [Computer software]. <publisher-name>Zenodo</publisher-name>. <pub-id pub-id-type="doi">10.5281/ZENODO.22281</pub-id></mixed-citation></ref>
<ref id="B69"><mixed-citation publication-type="journal"><string-name><surname>Ryan</surname>, <given-names>K. M.</given-names></string-name> (<year>2016</year>). <article-title>Phonological weight</article-title>. <source>Language and Linguistics Compass</source>, <volume>10</volume>, <fpage>720</fpage>&#8211;<lpage>733</lpage>.</mixed-citation></ref>
<ref id="B70"><mixed-citation publication-type="journal"><string-name><surname>Ryan</surname>, <given-names>K. M.</given-names></string-name> (<year>2019</year>). <article-title>Prosodic end-weight reflects phrasal stress</article-title>. <source>Natural Language &amp; Linguistic Theory</source>, <volume>37</volume>(<issue>1</issue>), <fpage>315</fpage>&#8211;<lpage>356</lpage>. <pub-id pub-id-type="doi">10.1007/s11049-018-9411-6</pub-id></mixed-citation></ref>
<ref id="B71"><mixed-citation publication-type="journal"><string-name><surname>Saffran</surname>, <given-names>J. R.</given-names></string-name>, <string-name><surname>Aslin</surname>, <given-names>R. N.</given-names></string-name>, &amp; <string-name><surname>Newport</surname>, <given-names>E. L.</given-names></string-name> (<year>1996</year>). <article-title>Statistical cues in language acquisition: Word segmentation by infants</article-title>. <source>COGSCI-96</source>, <fpage>376</fpage>&#8211;<lpage>380</lpage>.</mixed-citation></ref>
<ref id="B72"><mixed-citation publication-type="journal"><string-name><surname>Seo</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Kim</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Kubozono</surname>, <given-names>H.</given-names></string-name>, &amp; <string-name><surname>Cho</surname>, <given-names>T.</given-names></string-name> (<year>2019</year>). <article-title>Preboundary lengthening in Japanese: To what extent do lexical pitch accent and moraic structure matter?</article-title> <source>Journal of the Acoustical Society of America</source>, <volume>146</volume>, <fpage>1817</fpage>&#8211;<lpage>1823</lpage>. <pub-id pub-id-type="doi">10.1121/1.5122191</pub-id></mixed-citation></ref>
<ref id="B73"><mixed-citation publication-type="journal"><string-name><surname>Seyfarth</surname>, <given-names>S.</given-names></string-name> (<year>2014</year>). <article-title>Word informativity influences acoustic duration: Effects of contextual predictability on lexical representation</article-title>. <source>Cognition</source>, <volume>133</volume>(<issue>1</issue>), <fpage>140</fpage>&#8211;<lpage>155</lpage>. <pub-id pub-id-type="doi">10.1016/j.cognition.2014.06.013</pub-id></mixed-citation></ref>
<ref id="B74"><mixed-citation publication-type="book"><string-name><surname>Shattuck-Hufnagel</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name><surname>Prieto</surname>, <given-names>P.</given-names></string-name> (<year>2019</year>). <chapter-title>Dimensionalizing co-speech gestures</chapter-title>. In <string-name><given-names>S.</given-names> <surname>Calhoun</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Escudero</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Tabain</surname></string-name>, &amp; <string-name><given-names>P.</given-names> <surname>Warren</surname></string-name> (Eds.), <source>Proceedings of the 19th international congress of Phonetic Sciences</source>, <publisher-loc>Melbourne, Australia</publisher-loc> (pp. <fpage>1490</fpage>&#8211;<lpage>1494</lpage>). <publisher-name>Australasian Speech Science and Technology Association</publisher-name>.</mixed-citation></ref>
<ref id="B75"><mixed-citation publication-type="journal"><string-name><surname>Shattuck-Hufnagel</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name><surname>Ren</surname>, <given-names>A.</given-names></string-name> (<year>2018</year>). <article-title>The Prosodic Characteristics of Non-referential Co-speech Gestures in a Sample of Academic-Lecture-Style Speech</article-title>. <source>Frontiers in Psychology</source>, <volume>9</volume>, <elocation-id>1514</elocation-id>. <pub-id pub-id-type="doi">10.3389/fpsyg.2018.01514</pub-id></mixed-citation></ref>
<ref id="B76"><mixed-citation publication-type="journal"><string-name><surname>Shaw</surname>, <given-names>J.</given-names></string-name>, &amp; <string-name><surname>Kawahara</surname>, <given-names>S.</given-names></string-name> (<year>2018</year>). <article-title>Predictability and phonology: Past, present and future</article-title>. <source>Linguistics Vanguard</source>, <volume>4</volume>(<issue>s2</issue>), <elocation-id>20180042</elocation-id>. <pub-id pub-id-type="doi">10.1515/lingvan-2018-0042</pub-id></mixed-citation></ref>
<ref id="B77"><mixed-citation publication-type="book"><string-name><surname>Shaw</surname>, <given-names>J.</given-names></string-name>, &amp; <string-name><surname>Tang</surname>, <given-names>K.</given-names></string-name> (<year>2023</year>). <chapter-title>A dynamic neural field model of leaky prosody: proof of concept</chapter-title>. <source>Proceedings of the 2022 Annual Meeting in Phonology</source>. <publisher-name>University of California</publisher-name> at <publisher-loc>Los Angeles</publisher-loc>.</mixed-citation></ref>
<ref id="B78"><mixed-citation publication-type="book"><string-name><surname>Smolek</surname>, <given-names>A.</given-names></string-name> (<year>2010</year>). <chapter-title>Vowel harmony in Tuvan and Igbo: Statistical and Optimality Theoretic analyses</chapter-title>. [Unpublished BA thesis]. <publisher-name>Swarthmore College</publisher-name>.</mixed-citation></ref>
<ref id="B79"><mixed-citation publication-type="journal"><string-name><surname>Stoltmann</surname>, <given-names>K.</given-names></string-name>, &amp; <string-name><surname>Fuchs</surname>, <given-names>S.</given-names></string-name> (<year>2017</year>). <article-title>The influence of handedness and pointing direction on deictic gestures and speech interaction: Evidence from motion capture data on Polish counting-out rhymes</article-title>. <source>The 14th International Conference on Auditory-Visual Speech Processing</source>, <fpage>21</fpage>&#8211;<lpage>25</lpage>. <pub-id pub-id-type="doi">10.21437/AVSP.2017-5</pub-id></mixed-citation></ref>
<ref id="B80"><mixed-citation publication-type="journal"><string-name><surname>Suomi</surname>, <given-names>K.</given-names></string-name> (<year>1983</year>). <article-title>Palatal vowel harmony: A perceptually motivated phenomenon?</article-title> <source>Nordic Journal of Linguistics</source>, <volume>6</volume>(<issue>1</issue>), <fpage>1</fpage>&#8211;<lpage>35</lpage>. <pub-id pub-id-type="doi">10.1017/S0332586500000949</pub-id></mixed-citation></ref>
<ref id="B81"><mixed-citation publication-type="journal"><string-name><surname>Suomi</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>McQueen</surname>, <given-names>J. M.</given-names></string-name>, &amp; <string-name><surname>Cutler</surname>, <given-names>A.</given-names></string-name> (<year>1997</year>). <article-title>Vowel harmony and speech segmentation in Finnish</article-title>. <source>Journal of Memory and Language</source>, <volume>36</volume>, <fpage>422</fpage>&#8211;<lpage>444</lpage>.</mixed-citation></ref>
<ref id="B82"><mixed-citation publication-type="journal"><string-name><surname>Tang</surname>, <given-names>K.</given-names></string-name>, &amp; <string-name><surname>Bennett</surname>, <given-names>R.</given-names></string-name> (<year>2018</year>). <article-title>Contextual predictability influences word and morpheme duration in a morphologically complex language (Kaqchikel Mayan)</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>144</volume>, <fpage>997</fpage>&#8211;<lpage>1017</lpage>. <pub-id pub-id-type="doi">10.1121/1.5046095</pub-id></mixed-citation></ref>
<ref id="B83"><mixed-citation publication-type="journal"><string-name><surname>Tang</surname>, <given-names>K.</given-names></string-name>, &amp; <string-name><surname>Shaw</surname>, <given-names>J. A.</given-names></string-name> (<year>2021</year>). <article-title>Prosody leaks into the memories of words</article-title>. <source>Cognition</source>, <volume>210</volume>, <elocation-id>104601</elocation-id>. <pub-id pub-id-type="doi">10.1016/j.cognition.2021.104601</pub-id></mixed-citation></ref>
<ref id="B84"><mixed-citation publication-type="journal"><string-name><surname>Turk</surname>, <given-names>A.</given-names></string-name> (<year>2010</year>). <article-title>Does prosodic constituency signal relative predictability? A Smooth Signal Redundancy hypothesis</article-title>. <source>Laboratory Phonology</source>, <volume>1</volume>(<issue>2</issue>). <pub-id pub-id-type="doi">10.1515/labphon.2010.012</pub-id></mixed-citation></ref>
<ref id="B85"><mixed-citation publication-type="journal"><string-name><surname>Turk</surname>, <given-names>A.</given-names></string-name>, &amp; <string-name><surname>Shattuck-Hufnagel</surname>, <given-names>S.</given-names></string-name> (<year>2007</year>). <article-title>Multiple targets of phrase-final lengthening in American English words</article-title>. <source>Journal of Phonetics</source>, <volume>35</volume>, <fpage>445</fpage>&#8211;<lpage>472</lpage>.</mixed-citation></ref>
<ref id="B86"><mixed-citation publication-type="journal"><string-name><surname>Vroomen</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Tuomainen</surname>, <given-names>J.</given-names></string-name>, &amp; <string-name><surname>de Gelder</surname>, <given-names>B.</given-names></string-name> (<year>1998</year>). <article-title>The roles of word stress and vowel harmony in speech segmentation</article-title>. <source>Journal of Memory and Language</source>, <volume>38</volume>(<issue>2</issue>), <fpage>133</fpage>&#8211;<lpage>149</lpage>.</mixed-citation></ref>
<ref id="B87"><mixed-citation publication-type="journal"><string-name><surname>Whang</surname>, <given-names>J.</given-names></string-name> (<year>2018</year>). <article-title>Recoverability-driven coarticulation: Acoustic evidence from Japanese high vowel devoicing</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>143</volume>(<issue>2</issue>), <fpage>1159</fpage>&#8211;<lpage>1172</lpage>. <pub-id pub-id-type="doi">10.1121/1.5024893</pub-id></mixed-citation></ref>
<ref id="B88"><mixed-citation publication-type="journal"><string-name><surname>Zipf</surname>, <given-names>G. K.</given-names></string-name> (<year>1929</year>). <article-title>Relative frequency as a determinant of phonetic change</article-title>. <source>Harvard Studies in Classical Philology</source>, <volume>40</volume>, <fpage>1</fpage>&#8211;<lpage>95</lpage>. <pub-id pub-id-type="doi">10.2307/310585</pub-id></mixed-citation></ref>
<ref id="B89"><mixed-citation publication-type="journal"><string-name><surname>Zsiga</surname>, <given-names>E. C.</given-names></string-name> (<year>1997</year>). <article-title>Features, gestures and Igbo assimilation: An approach to the phonology-phonetics interface</article-title>. <source>Language</source>, <volume>73</volume>(<issue>2</issue>), <fpage>227</fpage>&#8211;<lpage>274</lpage>.</mixed-citation></ref>
</ref-list>
</back>
</article>