<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.0 20120330//EN" "http://jats.nlm.nih.gov/publishing/1.0/JATS-journalpublishing1.dtd">
<!--<?xml-stylesheet type="text/xsl" href="article.xsl"?>-->
<article article-type="research-article" dtd-version="1.0" xml:lang="en" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">
<front>
<journal-meta>
<journal-id journal-id-type="issn">1868-6354</journal-id>
<journal-title-group>
<journal-title>Laboratory Phonology</journal-title>
</journal-title-group>
<issn pub-type="epub">1868-6354</issn>
<publisher>
<publisher-name>Ubiquity Press</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.5334/labphon.92</article-id>
<article-categories>
<subj-group>
<subject>Journal article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>The Interpretation of Prosodic Variability in the Context of Accompanying Sociophonetic Cues</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Warren</surname>
<given-names>Paul</given-names>
</name>
<email>paul.warren@vuw.ac.nz</email>
<xref ref-type="aff" rid="aff-1"/>
</contrib>
</contrib-group>
<aff id="aff-1">School of Linguistics and Applied Language Studies, Victoria University of Wellington, Wellington, NZ</aff>
<pub-date publication-format="electronic" date-type="pub" iso-8601-date="2017-05-04">
<day>04</day>
<month>05</month>
<year>2017</year>
</pub-date>
<volume>8</volume>
<issue>1</issue>
<elocation-id>11</elocation-id>
<permissions>
<copyright-statement>Copyright: &#x00A9; 2017 The Author(s)</copyright-statement>
<copyright-year>2017</copyright-year>
<license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/4.0/">
<license-p>This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International License (CC-BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. See <uri xlink:href="http://creativecommons.org/licenses/by/4.0/">http://creativecommons.org/licenses/by/4.0/</uri>.</license-p>
</license>
</permissions>
<self-uri xlink:href="http://www.journal-labphon.org/articles/10.5334/labphon.92/"/>
<abstract>
<p>Production data have shown that one of the features distinguishing uptalk rises from question rises in New Zealand English (NZE) is the alignment point of the rise start, which is earlier in question utterances realized by younger speakers. Previous research has indicated that listeners are sensitive to this distinction in making a forced-choice decision as to whether an utterance is a statement or a question. NZE is also characterized by an ongoing merger of the <sc>NEAR</sc> and <sc>SQUARE</sc> diphthongs, with younger speakers more likely to realize the vowel in a word such as <italic>care</italic> with a closer starting point (as in [i&#600;], overlapping with their realization of the <sc>NEAR</sc> vowel), whereas older speakers would have more open starting point (as in [e&#600;]). The current study uses the mouse-tracking paradigm to provide evidence that the realization of <sc>SQUARE</sc> with an innovative vs. a conservative variant in a word early in an utterance affects NZE listeners&#8217; sensitivity firstly to a rise as a potential signal of an uptalked statement and secondly to the early alignment of the rise as a signal of a question. This finding indicates that the interpretation of prosodic variability can depend on speaker characteristics imputed from other sociophonetic cues.</p>
</abstract>
<kwd-group>
<kwd>uptalk</kwd>
<kwd>rising intonation</kwd>
<kwd>New Zealand English</kwd>
<kwd>sociophonetics</kwd>
<kwd>speech perception</kwd>
<kwd>mousetracking</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec>
<title>1 Introduction</title>
<p>This study considers whether listeners&#8217; interpretation of prosodic variables, namely the use of an utterance-final rise to indicate an uptalked statement and the alignment of the starting point of that rise to indicate whether the utterance is indeed uptalk or a question, depends on socially-conditioned expectations linked to a segmental phonetic cue earlier in the utterance. The response required of participants in the experiment reported below is whether an utterance is intended as a question or as a statement, a distinction that has been linked in New Zealand English (NZE) to the difference between early and late starting points respectively for the final rise. The segmental sociophonetic cue used in the experiment as a potential indicator of social grouping is the realization of a single vowel in a word that occurs earlier in each test utterance, prior to the intonational cue. This is the realization of the <sc>SQUARE</sc> diphthong,<xref ref-type="fn" rid="n1">1</xref> which for many speakers of the same dialect is merged with the <sc>NEAR</sc> diphthong. For non-merging speakers, the <sc>SQUARE</sc> and <sc>NEAR</sc> vowels might be transcribed phonemically as /e&#600;/ and /i&#600;/ respectively. For merging speakers, these two centering diphthongs both have a closer starting point that is more typical of the realization of <sc>NEAR</sc> for non-merging speakers, i.e., [i&#600;]. Uptalk use, the distinction between early and late rises for questions and statements and the <sc>NEAR-SQUARE</sc> merger are all linked in this language variety with younger speakers (predominantly teenagers and those in their twenties), especially females, although not all young New Zealand women regularly use uptalk or merge their <sc>NEAR</sc> and <sc>SQUARE</sc> diphthongs. The research question is whether manipulation of the vowel causes a change in listeners&#8217; expectations about the speaker, resulting in a shift in the interpretation of the intonation. Background information on both the intonational pattern and the vowel cue will now be presented, in order to contextualize the experimental design.</p>
<p>Uptalk has been documented for New Zealand English (largely under the label HRT, for high-rising terminal) since the mid-1960s (<xref ref-type="bibr" rid="B7">Benton, 1965</xref>). Early systematic studies indicated that it was most prevalent amongst the young, and particularly young women (<xref ref-type="bibr" rid="B6">Bell &amp; Johnson, 1997</xref>; <xref ref-type="bibr" rid="B10">Britain, 1992</xref>), and there remains a strong association of uptalk with younger speakers. An important factor for the experiment reported below is the extent of the overlap of uptalk with question intonation. Since intonational rises can also mark questions, there is a frequent lay perception that uptalk indicates a questioning nature, and that uptalkers are therefore insecure or lacking in confidence (see discussion in <xref ref-type="bibr" rid="B46">Warren, 2016</xref>). This association is not entirely surprising, given that it has been claimed that uptalk rises and question rises may be phonetically indistinguishable (<xref ref-type="bibr" rid="B24">Guy et al., 1986</xref>; <xref ref-type="bibr" rid="B32">Ladd, 1996</xref>; <xref ref-type="bibr" rid="B34">Lakoff, 1973</xref>). However, more recent research indicates emerging differences between these rise types in a number of varieties of English, with the nature of the differences dependent on variety (<xref ref-type="bibr" rid="B48">Warren &amp; Fletcher, 2016</xref>). In Australian English, for instance, uptalk rises tend to start from a lower pitch, possibly as part of a fall-rise contour (<xref ref-type="bibr" rid="B15">Fletcher &amp; Harrington, 2001</xref>; <xref ref-type="bibr" rid="B36">McGregor, 2005</xref>). Other research has suggested that question rises have earlier onsets than statement rises, in particular in NZE (<xref ref-type="bibr" rid="B14">Fletcher et al., 2005</xref>; <xref ref-type="bibr" rid="B45">Warren, 2005</xref>; <xref ref-type="bibr" rid="B47">Warren &amp; Daly, 2005</xref>), South African English (<xref ref-type="bibr" rid="B12">Dorrington, 2010</xref>), and Southern Californian (<xref ref-type="bibr" rid="B40">Ritchart &amp; Arvaniti, 2014</xref>). It should be stressed that early and late are relative concepts, and refer to the alignment of the rise onset with regard to the nuclear accent and following material, rather than in the utterance as a whole. In NZE, this alignment difference has been shown to be related to speaker age, with younger speakers more likely to distinguish between an early rise start for questions and a later start for uptalk utterances, while older speakers use later onsets for both sentence types (<xref ref-type="bibr" rid="B14">Fletcher et al., 2005</xref>; <xref ref-type="bibr" rid="B45">Warren, 2005</xref>).</p>
<p>A number of perceptual studies have investigated the cueing value of such differences. In one study, Fletcher and Loakes (<xref ref-type="bibr" rid="B16">2010</xref>) found that the most statement-like utterances in a set of rising Australian English utterances were those with a low onset, while questions were signalled more reliably by a high onset, reflecting the production data for that variety. In a small-scale study in NZE (<xref ref-type="bibr" rid="B45">Warren, 2005</xref>; <xref ref-type="bibr" rid="B52">Zwartz &amp; Warren, 2003</xref>), participants were asked to classify variants of a single utterance. The utterance in question had a six-syllable nucleus + tail sequence (<italic>basketball stadium</italic>), and the final rise was resynthesized firstly so that its onset was temporally aligned at a range of five equally distant points from the initial accented syllable (<italic>ba</italic>-) through to the utterance-final syllable (-<italic>um</italic>), and secondly so that rises either progressed linearly towards a final high point at the end of the utterance, or involved a sharp rise followed by a high plateau. The results indicated that questions were more clearly signalled by the sharp early rise, and statements by the sharp late rise, reflecting respectively the convex and concave contours found in production data from NZE speakers (<xref ref-type="bibr" rid="B45">Warren, 2005</xref>).</p>
<p>The current study aims to build on these findings by additionally investigating whether the listeners&#8217; perceptions of uptalk, and of early and late rises in NZE as signals of questions and statements respectively, are also dependent on whether the speaker is a likely &#8216;uptalker.&#8217; As indicated above, it does so indirectly by manipulating a segmental variable that has a social distribution similar to that of uptalk, i.e., the realization of the <sc>SQUARE</sc> diphthong. In a longitudinal study of the NZE <sc>NEAR-SQUARE</sc> merger, Gordon and Maclagan (<xref ref-type="bibr" rid="B23">2001</xref>) examined production data based on words containing these vowels, read in sentence and word-list contexts by adolescents from the same demographic, sampled every five years. While the diphthongs were still both widely present in their first sample from 1983, they showed significant overlap in adolescent speech by 1998. The merger has been towards a closer starting point for <sc>SQUARE</sc>, i.e., a merger-by-approximation towards <sc>NEAR</sc>. In the current paper, this innovative realization of <sc>SQUARE</sc> will be given the label [i&#600;], which is the transcription recommended for the <sc>NEAR</sc> diphthong in NZE by Bauer and Warren (<xref ref-type="bibr" rid="B5">2004</xref>). The more conservative variant will be labelled [e&#600;], the transcription they give for <sc>SQUARE</sc>, although it should be noted that single labels cannot do justice to the variation that is found in both the closer and the more open realization of the diphthong. Like the use of uptalk intonation, the use of a variant of <sc>SQUARE</sc> with a closer starting point continues to be commented on in the media and in letters-to-the-editor, particularly as a possible source of ambiguity (in this case resulting in perceived homophony of words like <italic>beer</italic> and <italic>bare</italic> or <italic>cheer</italic> and <italic>chair</italic>). It also continues to be associated with the speech of younger speakers, with the diphthongs remaining unmerged for many, especially older speakers.</p>
<p>Based on a series of experiments involving the production and perception of <sc>NEAR</sc> and <sc>SQUARE</sc>, Warren et al. (<xref ref-type="bibr" rid="B49">2007</xref>) concluded that the social indexing of the merger in NZE meant that lexical items containing a diphthong with the closer starting point of [i&#600;] were least likely to be identified as containing the <sc>SQUARE</sc> vowel when the speaker was perceived to be older and male. This was demonstrated in a study employing four voices (two of each sex) that were independently rated for probable speaker age. In another study Hay et al. (<xref ref-type="bibr" rid="B28">2006b</xref>) manipulated age through photographs and showed that perceived age of the speaker influenced identification accuracy of <sc>NEAR</sc> and <sc>SQUARE</sc> word pairs. The photographs were presented as though they were pictures of the speakers being listened to, and accuracy was greatest after photographs of older individuals, even though the same male and female speech tokens were presented after each of the male and female photographs respectively. The consequences of the merger for lexical access have also been investigated in semantic priming experiments (<xref ref-type="bibr" rid="B49">Warren et al., 2007</xref>; <xref ref-type="bibr" rid="B50">Warren et al., 2003</xref>), which showed that there was an asymmetry in the priming exhibited by young NZE-speaking participants, but that listeners were again sensitive to the age of the speaker. When the stimuli were from a younger voice, the form [&#679;i&#600;] primed both <italic>shout</italic> (an associate of <italic>cheer</italic>) and <italic>sit</italic> (an associate of <italic>chair</italic>), but the form [&#679;e&#600;] only primed <italic>sit</italic> (the associate of <italic>chair</italic>). When an older speaker was used, the priming was not asymmetrical&#8212;[&#679;i&#600;] primed only <italic>shout</italic> and [&#679;e&#600;] primed only <italic>sit</italic>.</p>
<p>The brief summaries above of research on uptalk and on the <sc>NEAR-SQUARE</sc> merger in NZE indicate some changes-in-progress that have been occurring over a similar timeframe, and which are similarly socially stratified. Younger speakers, particularly but not exclusively women, are more likely to use uptalk, to distinguish uptalk rises and question rises through the alignment of the start of the rise, and to have a closer starting point for the <sc>SQUARE</sc> diphthong.</p>
<p>The results reported above from the studies of <sc>NEAR</sc> and <sc>SQUARE</sc> conducted by Hay et al. (<xref ref-type="bibr" rid="B28">2006b</xref>) and by Warren et al. (<xref ref-type="bibr" rid="B49">2007</xref>) are part of a growing body of research that shows that social characteristics associated with speakers can affect the interpretation of phonetic information. In some of the early perceptual work in this area, Strand and Johnson (<xref ref-type="bibr" rid="B44">1996</xref>) found that participants&#8217; categorization of fricatives on a [&#643;]-[s] continuum was influenced by the putative sex of the speaker as indicated by video clips with which the audio signals were aligned. Subsequently, Johnson et al. (<xref ref-type="bibr" rid="B31">1999</xref>) showed that listeners&#8217; categorizations of vowels on a continuum from [&#650;] to [&#652;] were affected not only by a visually presented face (male or female), but also, in a separate experiment, by the imagined sex of the speaker (participants in one group were told that the speaker was female and asked to imagine a female speaker while doing the experiment, while the other group were told to do the same for a male speaker).</p>
<p>In other research, the speaker&#8217;s putative dialect origin has been shown to affect perceptual responses. Niedzielski (<xref ref-type="bibr" rid="B37">1999</xref>) asked participants, all residents of Detroit, to indicate which of a set of resynthesized vowels best matched a vowel in a sentence they had heard. The results showed that participants who were led to believe that the speaker was from Canada chose a different resynthesized vowel from those who were led to believe that the speaker was from Detroit, despite the fact that both groups of participants heard precisely the same sentences and resynthesized vowels. Hay et al. (<xref ref-type="bibr" rid="B27">2006a</xref>) asked their participants to identify which of a series of resynthesized vowels best matched an /&#618;/ target vowel. Again, the experimental manipulation was the supposed dialect origin of the speaker, but in this case this was signalled not explicitly through instructions to the participants, but by means of the appearance of the words &#8216;Australian&#8217; or &#8216;New Zealander&#8217; at the top of the response sheets. While the results for male participants were inconclusive, female participants were more likely to select a raised, more Australian token from the continuum when they were in the &#8216;Australian&#8217; condition. In a subsequent study, Hay and Drager (<xref ref-type="bibr" rid="B26">2010</xref>) found that the mere presence of a stuffed toy kangaroo (indicating Australia) or kiwi (indicating New Zealand) could influence listeners&#8217; responses.</p>
<p>As well as the age effects shown in the studies of NZE <sc>NEAR</sc> and <sc>SQUARE</sc> by Hay et al. (<xref ref-type="bibr" rid="B28">2006b</xref>) and Warren et al. (<xref ref-type="bibr" rid="B49">2007</xref>), Drager (<xref ref-type="bibr" rid="B13">2011</xref>) found that the speaker&#8217;s perceived age influenced the categorization of vowels on a <sc>DRESS-TRAP</sc> continuum in NZE, a variety which has a well-established pattern of raising of the short front vowels. However, she found this effect only for older participants, which she conjectures may be linked to their greater experience of a range of speakers from different generations as well as to their greater exposure to the progression of the sound change.</p>
<p>The studies reviewed above have exploited the potential social indexicality of phonetic cues and have shown that segmental phonetic perception can be affected by the perceived characteristics of the speaker (as prompted by photographs, movie clips, dialect region labels, or even stuffed toys). The current study builds on these links between speaker characteristics and phonetic properties at the segmental level, and investigates whether segmental differences can impact on the interpretation of cues at the suprasegmental level. The interaction of segmental and suprasegmental cues in linguistic indexicality has previously been explored by Levon (<xref ref-type="bibr" rid="B35">2007</xref>), who manipulated sibilant duration and pitch range in a read passage. Levon found that pitch range only affected judgements of the speaker on an effeminate-masculine scale when the sibilants were short, and that sibilant duration only affected such judgements when pitch range was narrow. The current study takes a different approach to the interaction of segmental and suprasegmental cues. Rather than asking participants to judge characteristics of the speaker on the basis of phonetic cues, it exploits the fact that the realization of the <sc>SQUARE</sc> vowel, the use of uptalk, and the alignment differences in statement and question pitch rises are typically co-indexical of younger speakers,<xref ref-type="fn" rid="n2">2</xref> and examines whether as a consequence variation in the segmental cue will affect the interpretation of the suprasegmental cues. That is, if a <sc>SQUARE</sc> diphthong in an utterance has an [i&#600;] realization and signals a speaker who is more likely to produce uptalk, then a subsequent rising intonation is more likely to be interpreted as uptalk. In addition, if the type of speaker signalled by the [i&#600;] realization of <sc>SQUARE</sc> is also the type of speaker who makes a phonetic distinction between question and statement rises, then a subsequent early rise should be more likely to signal a question and a late rise a statement, relative to a situation where the diphthong has a more conservative [e&#600;] realization.</p>
<p>The two speech features indicated above&#8212;the realization of the <sc>SQUARE</sc> diphthong as [i&#600;] or [e&#600;] and the realization of a rise with an early or late onset&#8212;were manipulated on utterances with declarative word order which were then used in a forced-choice task, where participants had to select between two sentence types&#8212;question and statement. An example sentence is given in (1).</p>
<list list-type="gloss">
<list-item>
<list list-type="wordfirst">
<list-item><p>(1)</p></list-item>
</list>
</list-item>
<list-item>
<list list-type="sentence-gloss">
<list-item>
<list list-type="final-sentence">
<list-item><p>John&#8217;s mother cared for stray animals.</p></list-item>
</list>
</list-item>
</list>
</list-item>
</list>
<p>The word that has the <sc>SQUARE</sc> diphthong in this example is <italic>cared</italic>. All relevant words in the test utterances were words which would have an [e&#600;] realization of the <sc>SQUARE</sc> diphthong in the speech of more conservative speakers, and for which there was no minimal pair word that differed only in having the <sc>NEAR</sc> vowel. This vowel was manipulated (see below) so that it had either the conservative [e&#600;] realization or the innovative [i&#600;] realization.</p>
<p>The rise was on the final nuclear accented word, <italic>animals</italic>, and was manipulated (see below) so that it started either at the end of the accented syllable (i.e., the first syllable) or at the beginning of the final syllable, which was also the last syllable in the utterance. All final words were three syllables long. These rise alignments are compatible with those found in the production research described above.</p>
<p>A further aspect of the experiment reported below that differs from the previous forced-choice studies of sentence type is that in addition to collecting response choice and reaction time, the experiment tracks mouse movements made by participants as they make their selection. The mouse-tracking technique records the (x, y) pixel coordinates of the trajectory of cursor movements on a computer screen as participants use the computer mouse to move the cursor from its starting point to a decision target. Studies using this technique have shown that details of the mouse trajectory, such as its curvature and complexity, co-vary with a range of cognitive processes involved for instance in auditory lexical decision (<xref ref-type="bibr" rid="B43">Spivey et al., 2005</xref>), visual lexical decision (<xref ref-type="bibr" rid="B2">Barca &amp; Pezzulo, 2012</xref>, <xref ref-type="bibr" rid="B3">2015</xref>), decisions as to the truth or falsity of negated sentences (<xref ref-type="bibr" rid="B11">Dale &amp; Duran, 2011</xref>), memory strength (<xref ref-type="bibr" rid="B38">Papesh &amp; Goldinger, 2012</xref>), the automatic activation of phonological information during visual word processing (<xref ref-type="bibr" rid="B1">Barca et al., 2016</xref>), and social categorization (<xref ref-type="bibr" rid="B17">Freeman, 2014</xref>; <xref ref-type="bibr" rid="B19">Freeman et al., 2008</xref>; <xref ref-type="bibr" rid="B21">Freeman et al., 2011</xref>). Analyzing mouse trajectories can provide additional insights into decision processes that cannot be measured simply from the outcome decision, such as the attraction strength of competing responses over time. In addition, it has been demonstrated that movement trajectories reflect the confidence of a participant&#8217;s response, with more direct trajectories to the target correlating with higher confidence scores reported by the participants after their decisions (<xref ref-type="bibr" rid="B38">Papesh &amp; Goldinger, 2012</xref>).<xref ref-type="fn" rid="n3">3</xref></p>
<p>The experiment tests the following hypotheses, which arise from consideration of the literature reviewed above. Firstly, there will be more &#8216;question&#8217; responses (which will also be faster and more direct) following early rises than following late rises. Secondly, &#8216;statement&#8217; responses will be more likely (and faster and more direct) in the context of an [i&#600;] realization of the <sc>SQUARE</sc> diphthong. This would be a new finding, and would indicate that a sociophonetic cue at the segmental level influences the interpretation of a potentially ambiguous cue at the suprasegmental level (i.e., whether a rise indicates a question or uptalk). Thirdly, an [i&#600;] realization of <sc>SQUARE</sc>, indicating a younger speaker, will show greater compatibility with the use of early and late rises to indicate questions vs. statements respectively. This again would be a new finding, but one which is compatible with a &#8216;gestalt-like understanding of indexicality&#8217; (<xref ref-type="bibr" rid="B35">Levon, 2007: 546</xref>).</p>
</sec>
<sec>
<title>2 Experiment</title>
<sec sec-type="methods">
<title>2.1 Method</title>
<sec>
<title>2.1.1 Materials</title>
<p>The experimental materials consisted of 20 short utterances (see Appendix), each of which contained a word with a <sc>SQUARE</sc> vowel, followed on average 5.2 syllables later (<italic>SD</italic> 1.77, range 2&#8211;8 syllables) by the word bearing the nuclear accent. The nuclear accented word was in all cases three syllables long and had initial stress. In addition, there were 40 filler items and 12 practice items. The fillers consisted of 10 questions involving inversion (e.g., &#8220;Are they leaving tomorrow?&#8221;), 10 wh-questions (&#8220;What is the best way of getting red wine stains out of clothes?&#8221;), and 20 statements with final falls (&#8220;They were disappointed that the concert ended so early.&#8221;). The practice items consisted of 6 statements with final falls, 2 wh-questions, 2 inversion questions, and 2 declarative utterances with final rises. So that participants would not be influenced by the realizations of any <sc>SQUARE</sc> or <sc>NEAR</sc> vowels outside of the test utterances, none of the filler or practice items contained either of these diphthongs.</p>
<p>A 25-year old female native speaker of NZE was recorded reading all 20 test items both as questions and as uptalk sentences. She was not asked specifically to produce a particular variant of the <sc>SQUARE</sc> vowel in the test utterances, since the intention was to resynthesize the vowel using average vowel formants from three further female speakers who consistently distinguished <sc>SQUARE</sc> and <sc>NEAR</sc> (see below). In the same recording session she also read out all the fillers and practice items. From the set of recorded test items, 10 question recordings and 10 uptalk recordings were selected as the source utterances for manipulations. Pitch manipulation was carried out using PSOLA implemented in Praat (<xref ref-type="bibr" rid="B8">Boersma &amp; Weenink, 2012</xref>). Pitch values for the final rise were based on those of the source utterance. The average starting value for the pitch rise across the 20 test utterances was 184 Hz (<italic>SD</italic> 6.9 Hz). The average pitch at this point was marginally higher for utterances sourced from questions (185 Hz, <italic>SD</italic> 8.7 Hz) than for those sourced from statements (183 Hz, <italic>SD</italic> 4.8 Hz). This difference was not significant by <italic>t</italic>-test (<italic>p</italic> = 0.64). The pitch level at the end of the final rise was similarly based on that of the source utterance and was at an average of 423 Hz (<italic>SD</italic> 24.0 Hz). The final pitch value of utterances sourced from questions (422 Hz, <italic>SD</italic> 19 Hz) was marginally lower than that of utterances sourced from statements (425 Hz, <italic>SD</italic> 29 Hz). Again, this difference was not significant (<italic>p</italic> = 0.80). The fact that these beginning and end values did not differ is further confirmation that in this variety the distinction between question and uptalk rises does not equate to a difference in pitch height (unlike, say, Australian English).</p>
<p>In the early rise condition, the pitch level of the start of the rise was kept constant across the accented syllable, after which it rose linearly to the end of the word, i.e., at the end of the third syllable. In the late rise condition, the pitch level of the start of the rise was maintained across both the accented syllable and the following unaccented syllable, and then rose linearly and sharply across the final syllable to the end of the word. These pitch shapes are based on the previous production studies of NZE intonation noted in the Introduction. The mean duration of the early rise was 349 ms (<italic>SD</italic> 106 ms). It was slightly longer for utterances taken from a statement source utterance (350 ms, <italic>SD</italic> 116 ms) than for those taken from questions (348 ms, <italic>SD</italic> 101 ms), a difference that was not significant (<italic>p</italic> = 0.96). The mean duration of the late rise was 174 ms (<italic>SD</italic> 84 ms), and was longer for utterances sourced from questions (182 ms, <italic>SD</italic> 84 ms) than for those sourced from statements (165 ms, <italic>SD</italic> 87 ms). This difference was also not significant (<italic>p</italic> = 0.66). That these duration values are so similar within each set is probably a reflection of the fact that all nuclei were realized across three-syllable words with stress on the first syllable. Examples of early and late rise stimuli are shown in Figure <xref ref-type="fig" rid="F1">1</xref> for the utterance &#8220;John&#8217;s mother cared for stray animals.&#8221;</p>
<fig id="F1">
<label>Figure 1</label>
<caption>
<p>Examples of early rise (top) and late rise (bottom) utterances used in the experiment, showing waveforms, pitch contours, and word-level TextGrids. The pitch range of the display is 75Hz to 500Hz. For each such pair, two examples of each stimulus were used, with [i&#600;] and [e&#600;] variants of the <sc>SQUARE</sc> vowel, in this case in the word &#8216;cared.&#8217; Examples of the four stimuli based on the sentence illustrated in this figure can be accessed at DOI: <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://doi.org/10.5334/labphon.92.s1">https://doi.org/10.5334/labphon.92.s1</ext-link></p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="/article/id/6198/file/74775/"/>
</fig>
<p>The formants of the beginning portion of the <sc>SQUARE</sc> diphthong were manipulated and resynthesized using linear predictive coding in Praat, via a purpose-designed script written by the author, in order to produce variants of this diphthong with (relatively) closer and more open starting points. As indicated above, these variants will be referred to as the [i&#600;] and [e&#600;] variants of <sc>SQUARE</sc>. The [i&#600;] variant is the innovative variant and would be in the <sc>NEAR</sc> vowel space of more conservative speakers. The script requested the first and second formant values to be used in the resynthesis and then prompted the user to specify, via mouse-clicks on the Praat display of speech wave and spectrogram, the beginnings and ends of the area to be manipulated. The script imposed logistic functions to create smooth transitions for the formants from the source wave before the manipulation area into the manipulation area and back again from the manipulation area to the source wave. The script was used to modify F1 and F2 of the first target of the diphthong, stretching over the first third of the diphthong, with the final target of the diphthong left as in the original recording and varying naturally depending on the following sound. The F1 and F2 values used for the first target of the resynthesized diphthongs were based on average formant values from minimal pair word-list recordings from a group of 3 female NZE speakers who distinguish <sc>NEAR</sc> and <sc>SQUARE</sc> in their speech. For <sc>NEAR</sc> the average F1 for these speakers was 331 Hz and F2 was 2711 Hz, while for <sc>SQUARE</sc> their average F1 and F2 were 657 Hz and 2170 Hz respectively. These values formed the basis for the resynthesis of the [i&#600;] and [e&#600;] variants of <sc>SQUARE</sc>.</p>
</sec>
<sec>
<title>2.1.2 Design</title>
<p>Two test lists were produced; in one the test sentences contained the more open conservative [e&#600;] version of the <sc>SQUARE</sc> diphthong, and in the other they contained the closer innovative [i&#600;] version. Participants were randomly allocated to one of these lists. In each list, the 20 test items occurred twice, once with the early rise, and once with the late rise. The two rise versions were allocated to two separate blocks, and each block had an equal number of early and late rise test items. Thus all participants heard each item with both the early- and late-rise intonation, but they only heard one version of the diphthong. So that not all repeated utterances were test utterances, half of the 40 fillers were repeated in the course of the experiment, giving 60 filler items in total, and a total experimental list (excluding practice items) of 100 utterances. The two instances of repeated fillers were assigned to separate blocks. The 60 filler items were evenly and pseudo-randomly distributed throughout the test list, such that sequences of more than two items of the same type were avoided. The position on the screen of the choice targets (&#8216;question&#8217; and &#8216;statement,&#8217; presented in capitals; see below) remained constant for each participant, but was switched for half the participants in each list.</p>
</sec>
<sec>
<title>2.1.3 Participants</title>
<p>Data from 36 native speakers of NZE (27 females) were included in the analysis below. Their age range was 18&#8211;32 (mean 22.4, <italic>SD</italic> 3.5). Three further participants were replaced, 2 because of equipment failure and 1 because he failed to select a response before the time-out of 5 seconds for a high proportion of test items (52.5%; no other participant had more than 7.5% of test data missing for that reason).</p>
</sec>
<sec>
<title>2.1.4 Procedure</title>
<p>The experiment was run in E-Prime 2.0 (<xref ref-type="bibr" rid="B39">Psychology Software Tools, 2012</xref>), using a script developed by the author. The screen resolution was 1920 &#215; 1080. Mouse positions were tracked every 10 milliseconds. The E-Prime script first presented an instruction screen that contained the text below. Note that the text points out that some utterances have declarative word order but rising intonation and thus have the potential of being either questions or statements, but does not draw attention to how these might differ from one another.</p>
<disp-quote>
<p>In this task, you will hear utterances that could be questions or statements. Your task is to decide whether each one is a question or a statement. Some of them will be easier than others because they start with question words like &#8216;when.&#8217; Others may only be marked by intonation. Note, though, that some intonation patterns, such as rising intonation, are often found on statements too, so you will still need to decide whether the utterance was intended as a statement or question.</p>
<p>For each utterance you will first see a START box at the bottom of the page. Click on START. You will then hear the utterance and at the top of the page you will see the words QUESTION and STATEMENT. Click on the word that corresponds to the utterance type for that utterance.</p>
</disp-quote>
<p>Once participants had indicated that they understood these instructions, the practice items were presented. First, a screen with three clickable areas was presented, as per the instructions above. The &#8216;start&#8217; point was centered at the bottom, taking up 11% of the screen&#8217;s width and 6% of its height. Response targets were top left and top right, each taking 25% of the screen&#8217;s width and 10% of its height (see Figure <xref ref-type="fig" rid="F2">2</xref> for a not-to-scale representation of the layout). The presentation of each audio stimulus over headphones commenced as soon as the participant clicked in the region marked by &#8216;start&#8217; and mouse coordinates were continuously recorded until the participant clicked in one of the two target regions (&#8216;question&#8217; and &#8216;statement&#8217;), or until 5 seconds elapsed, whichever was sooner. Immediately after this, the screen was blanked briefly and then the cycle began for the next stimulus.</p>
<fig id="F2">
<label>Figure 2</label>
<caption>
<p>Screen layout (not to scale) with example mouse trajectory, showing start point, choice target points, and AUC and MD measures (see text).</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="/article/id/6198/file/74776/"/>
</fig>
<p>At the conclusion of the practice set, participants were encouraged to seek any advice on procedural issues, and then the two main blocks of test and filler items were run without a further break.</p>
</sec>
</sec>
<sec>
<title>2.2 Mouse-tracking measures</title>
<p>The (x, y) coordinates from the mouse-tracking data can be analyzed in a number of ways. The most common analyses are measures of the displacement of the mouse trajectory from a straight-line response from the starting position to the target for the choice being made (see for instance <xref ref-type="bibr" rid="B17">Freeman, 2014</xref>; <xref ref-type="bibr" rid="B18">Freeman &amp; Ambady, 2010</xref>; <xref ref-type="bibr" rid="B29">Hehman et al., 2014</xref>; <xref ref-type="bibr" rid="B38">Papesh &amp; Goldinger, 2012</xref>). Two such measures are shown in Figure <xref ref-type="fig" rid="F2">2</xref>. One of these is Maximum Deviation (MD), which is the largest perpendicular distance from the straight-line trajectory to the actual trajectory. The other is the Area Under the Curve (AUC), which is the area defined by the actual trajectory and the straight-line trajectory. The greater the value of MD or AUC, the more the trajectory has deviated from the straight line and moved towards the alternative response. Negative values of AUC and MD can also exist, reflecting a path that goes below the idealized straight-line trajectory. In addition, (x, y) coordinate data can be analyzed for changes in direction, speed, or acceleration, in either the horizontal or vertical dimension, or both. Changes in direction can either be simple reversals (x-flips, y-flips, or both) or more complex measures such as sample entropy (for details of these and other measures, see <xref ref-type="bibr" rid="B29">Hehman et al., 2014</xref>).</p>
<p>MD is the measure selected for analysis in the current paper. It has been claimed to index &#8220;the partial, simultaneous activation of a competing representation of the opposite category&#8221; (<xref ref-type="bibr" rid="B17">Freeman, 2014: 87</xref>). In a test of recognition of visually-presented words as old (they had previously been seen in a training set) or new (they had not), stronger memories were associated with fast and linear mouse-tracking responses, while weaker memories had tracks that were slower and curvilinear (<xref ref-type="bibr" rid="B38">Papesh &amp; Goldinger, 2012</xref>). The same researchers argued that &#8220;movement trajectories revealed underlying response confidence&#8221; (p. 906). Competition effects were found by Spivey et al. (<xref ref-type="bibr" rid="B43">2005</xref>) in a task where participants listened to a word and had to match this to one of two pictures, one of the object corresponding to the word and one of a different object. They found less curvature of the mouse-track towards the alternative response when the word corresponding to the competing picture was a phonologically unrelated word (e.g., <italic>jacket</italic>) than when it was a member of the same phonological cohort (e.g., <italic>candle</italic>, for the auditory word <italic>candy</italic>). The MD measure, then, should provide additional information about the competition between question and statement responses in the context of the experimental manipulations. It should add an index of confidence to data involving response choice and response times, and it might reveal subtleties in participants&#8217; response behaviours that do not show up in the other measures.</p>
</sec>
<sec>
<title>2.3 Statistical analysis</title>
<p>In addition to the participant exclusions described above, individual data that exceeded the 5 second time out were excluded from analysis. This amounted to 2.0% of the test data (29 responses). In addition, because our interest is in the impact of the rise alignment on decisions, responses that were made earlier than the point at which the rise started were also excluded. This was a further 2.6% (38 responses). A total of 1373 mouse tracks remained for the test items.</p>
<p>Statistical analysis of the response choices, response times, and mouse-tracking data for test items was by means of mixed effects models in R, using the <italic>lme4</italic> package (<xref ref-type="bibr" rid="B4">Bates et al., 2015</xref>). Logistic models (using <italic>glmer</italic>) were applied to binary data such as response choices, and linear models (using <italic>lmer</italic>) to continuous data such as response times and MD. The statistical significance of including a factor or interaction in a model was assessed using the <italic>mixed</italic> command from the <italic>afex</italic> package (<xref ref-type="bibr" rid="B42">Singmann et al., 2015</xref>). Fixed effects included Vowel ([i&#600;] or [e&#600;]), Rise (early or late), Source (whether the stimulus was derived from an original question recording or uptalk statement recording&#8212;see &#8216;Preparation of test stimuli&#8217; above), the Serial Position of the stimulus in the experiment, and, where appropriate, the Choice (question or statement) made by the participant. Source was included as a test of whether there are other aspects of the utterances beyond Vowel and Rise that might affect responses. This seemed a sensible addition given that half of the source recordings were questions and half were uptalk statements. Items and participants were included as random effects, together with random slopes by participant for the Serial Position of a stimulus in the experiment. These random slopes were included to account for between-participant variation in how response behaviour (e.g., speeding up in making responses) changes over the course of the experiment.</p>
</sec>
<sec>
<title>2.4 Predictions</title>
<p>The hypotheses set out in the Introduction lead to the following predictions for the forced choice binary selection between question and statement responses:</p>
<list list-type="order">
<list-item><p>There will be a significant effect of rise alignment on response selection, with more question responses for early rises than for late rises. Note that this augments the previous perceptual result reported by Zwartz and Warren (<xref ref-type="bibr" rid="B50">2003</xref>), which was based on a longer terminal sequence (a two-word sequence of 6 syllables, compared with a single 3-syllable word in the current study), which might be expected to result in a clearer contrast between early and late rises. In addition, the selection of the question response will be more rapid and involve a more direct mouse movement to the target in the early rise condition, while statement responses will be made more rapidly and with more direct mouse trajectories in the late rise condition.</p></list-item>
<list-item><p>There will be a significant effect of the realization of <sc>SQUARE</sc> on response selection. Since the realization of <sc>SQUARE</sc> as [i&#600;] is more likely in the speech of younger speakers, who are also more likely to produce uptalk, statement responses are predicted to be more likely (and faster and more direct) following an [i&#600;] realization of <sc>SQUARE</sc> than after an [e&#600;] realization.</p></list-item>
<list-item><p>There will be a significant interaction of rise alignment and <sc>SQUARE</sc> realization on response selection. Since speakers who merge <sc>NEAR</sc> and <sc>SQUARE</sc> onto an [i&#600;] pronunciation are members of the same social group (younger speakers) for whom there is evidence that question and statement rises are becoming distinguished through an earlier rise on questions than on statements, it should follow that an [i&#600;] pronunciation of <sc>SQUARE</sc> will signal a speaker who is likely to have early question rises. Therefore when an early rise is heard in combination with an [i&#600;] pronunciation of <sc>SQUARE</sc>, question responses will be more likely and will be made both quickly and with a direct mouse trajectory. Conversely, if an [i&#600;] pronunciation of <sc>SQUARE</sc> is followed by a late rise, then this will be more likely to result in a statement response. However, since the [e&#600;] realization of <sc>SQUARE</sc> is the more conservative variant, not only (as per prediction 2) will uptalk be unexpected after [e&#600;], but also the difference between early and late rises will be less reliable as a cue to question and statement respectively.</p></list-item>
</list>
</sec>
</sec>
<sec>
<title>3 Results</title>
<sec>
<title>3.1 Response choice</title>
<p>The logistic regression model for response choice (question or statement) as the dependent variable included the random effect structure outlined above, together with Vowel, Rise, Source, and Serial Position as fixed effects, as well as the interactions of Rise with Vowel (to test whether an [i&#600;] realization of the <sc>SQUARE</sc> diphthong would make it more likely that an early rise would be interpreted as marking a question) and of Rise with Source (to test whether there are other properties of the original question and uptalk utterances that operated in conjunction with rise alignment to signal the intended utterance). The model produced a significant effect of Rise (&#967;<sup>2</sup>(1) = 43.47, <italic>p</italic> &lt; 0.0001). The factors Vowel, Source, and Serial Position all failed to produce significant simple effects, and there were no significant interactions. The effect of Rise is shown in Figure <xref ref-type="fig" rid="F3">3</xref>. Prediction 1 above is therefore supported by the finding of significantly more question responses after early rises than after late rises, but the lack of an effect of Vowel means that the response choice data do not support prediction 2. In addition, prediction 3 is not supported, since there are no differences in the proportion of question responses that depend on the interaction of the vowel with the alignment of the rise.</p>
<fig id="F3">
<label>Figure 3</label>
<caption>
<p>Effect of Rise alignment on the selection of question (vs. statement) responses.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="/article/id/6198/file/74777/"/>
</fig>
<p>A further observation from this initial analysis is that overall there were more question responses than statement responses (73% across all test items and conditions). One likely reason for the low number of statement (uptalk) responses is that uptalk occurs more typically in narrative structures with connected utterances than in isolated sentences such as those presented in the experiment. This low overall count of statement responses will prove to be important in later analyses of subgroups of data.</p>
</sec>
<sec>
<title>3.2 Response times</title>
<p>An initial inspection of response times (RTs) indicated that, as is commonly found, they were not normally distributed. A comparison of a selection of typical transformations indicated that the logarithm of RTs produced the best fit to a normal distribution (<italic>r</italic> = 0.999, compared with <italic>r</italic> = 0.983 for untransformed RTs). The statistical tests for response times were therefore conducted on log-transformed RTs. For clarity of presentation, however, raw RTs will be graphed.</p>
<p>In addition to the factors examined in the analysis of response choices, the RT analysis included Choice (i.e., whether the participant clicked on the question or statement response box). The analysis therefore considered the simple effects of Rise, Vowel, Source and Choice, as well as their interactions, and Serial Position. There was no simple effect (or interaction) related to whether the source of the manipulated stimuli was a question or a statement. Significant interactions were found between Choice and Rise (&#967;<sup>2</sup>(1) = 16.52, <italic>p</italic> &lt; 0.0001) and between Choice and Vowel (&#967;<sup>2</sup>(1) = 4.31, <italic>p</italic> &lt; 0.05). There was also a significant simple effect of Serial Position (&#967;<sup>2</sup>(1) = 11.61, <italic>p</italic> &lt; 0.001) &#8211; participants made their response selection more quickly as the experiment progressed.</p>
<p>The interaction of Choice and Rise (Figure <xref ref-type="fig" rid="F4">4</xref>) partially supports prediction 1. That is, after early rises, question responses are faster than statement responses. However, after late rises the speed of statement responses is not any different to that of question responses. This suggests that although participants take an early rise as an indicator of a question, they do not show any preference to interpret a late rise as indicating a statement. Note, however, that this is not entirely surprising, since the trend over apparent time reported in the Introduction is for question rises to move to earlier alignment, meaning that older speakers (who our participants will still be listening to) will have late rises for questions as well as for statements, although rises on statements will be rather rare, since these older speakers are less likely to use uptalk.</p>
<fig id="F4">
<label>Figure 4</label>
<caption>
<p>Effect of Rise alignment on response times (RTs) in the selection of question and statement responses.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="/article/id/6198/file/74778/"/>
</fig>
<p>The two-way interaction of Choice and Vowel (Figure <xref ref-type="fig" rid="F5">5</xref>) provides support for prediction 2. The interaction is due to an increase in latencies for statement responses following stimuli containing the more conservative [e&#600;] realization of the <sc>SQUARE</sc> diphthong, compared to the other factor combinations shown in the figure. The finding that statement decisions take longer than question decisions after the [e&#600;] realization is compatible with the observation that as the conservative variant [e&#600;] signals speakers who are less likely to produce statements with rising intonation.</p>
<fig id="F5">
<label>Figure 5</label>
<caption>
<p>Effect of Vowel on response times (RTs) in the selection of question and statement responses.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="/article/id/6198/file/74779/"/>
</fig>
</sec>
<sec>
<title>3.3 Maximum Deviation</title>
<p>It has been pointed out (e.g., <xref ref-type="bibr" rid="B18">Freeman &amp; Ambady, 2010</xref>; <xref ref-type="bibr" rid="B29">Hehman et al., 2014</xref>) that if the mean trajectory for an experimental condition shows a moderate deviation from a straight line, and therefore also a moderate average MD value, then it might result from averaging two trajectory distributions&#8212;one that reflects marked attractions to the alternative response before resolving onto the selected response, and one that consists of more-or-less direct movement to the selected response. In at least one study using mouse-tracking there has been an explicit prediction of such a pattern. Dale and Duran (<xref ref-type="bibr" rid="B11">2011</xref>) asked participants to judge whether statements were true. One of the main experimental parameters of interest was whether the statement sentence contained a negation, since it has long been attested that the presence of a negation slows down readers during verification tasks. The researchers predicted that verification trials with high MD values, reflecting a strong but ultimately resisted temptation to respond that the sentence is not true, would be found when the sentence contained a negation. Their results supported this prediction. In another study, in which faces were categorized for sex, Freeman (<xref ref-type="bibr" rid="B17">2014</xref>) found that the abrupt trajectory reversals (i.e., changes of mind) that are reflected in high MD values were more likely to occur with more ambiguous stimuli. In a second experiment these reversals were most likely when atypical faces (male faces with long hair or females with short hair) were seen in a normative context (i.e., where the majority of the male faces had short hair and the females long hair), but were also more likely for typical faces in counter-normative contexts (where the majority of male faces were seen with long hair, for instance) than in normative contexts.</p>
<p>The interpretation of MD data therefore crucially depends on whether mouse trajectories belong to a unimodal or bimodal distribution. While bimodality would mean that the usual statistical assumptions concerning normal distributions are challenged for the dataset as a whole, the presence of a bimodal distribution can itself be informative. As can be seen from Figure <xref ref-type="fig" rid="F6">6</xref>, the trajectories for test items in the current experiment are bimodally distributed. Hartigan&#8217;s dip statistic for unimodality (<xref ref-type="bibr" rid="B25">Hartigan &amp; Hartigan, 1985</xref>) confirms that the distribution is significantly bimodal (D = 0.032, <italic>p</italic> &lt; 0.0001). (For validation of the use of this statistic in the context of mouse-tracking data, see <xref ref-type="bibr" rid="B20">Freeman &amp; Dale, 2013</xref>.) The vertical line indicates the empirically-derived cutoff value between the two distributions. Note that this is lower than the 0.9 value reported by Freeman (<xref ref-type="bibr" rid="B17">2014</xref>) for his study of sex categorization of faces. As explained earlier, negative MD values reflect trajectories that go below the idealized straight line from start to finish.</p>
<fig id="F6">
<label>Figure 6</label>
<caption>
<p>Distribution of Maximum Deviation values. (Inset is for question responses only. See text).</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="/article/id/6198/file/74780/"/>
</fig>
<p>To explore factors that might influence the bimodal distribution evidenced in Figure <xref ref-type="fig" rid="F6">6</xref>, responses were categorized into low and high MD groups, and this classification was used as the dependent variable in further regression analysis. This analysis will be reported for the question responses only, since an initial model using all responses failed to converge once the low number of statement responses (see Figure <xref ref-type="fig" rid="F3">3</xref>) was further subdivided into low and high MD groups. As is apparent from the inset in Figure <xref ref-type="fig" rid="F6">6</xref>, the question responses showed a similar bimodal distribution to the overall pattern. The dip test returned the same result as that for the complete set (D = 0.032, <italic>p</italic> &lt; 0.0001); the empirically-derived cutoff value used for the categorization was slightly higher (at 0.781). The high MD group had a mean MD value of 1.18 (<italic>SD</italic> 0.20) and the low MD group had a mean MD value of 0.17 (<italic>SD</italic> 0.28). An analysis of RTs for question responses with MD group as one of the predictors returned an unsurprising result, with the more direct low MD tracks significantly faster than the high MD tracks (&#967;<sup>2</sup>(1) = 46.46, <italic>p</italic> &lt; 0.0001).</p>
<p>Logistic regression analysis of the question response set with MD group as the dependent variable and the fixed effects of Vowel, Rise, Source, and Serial Position, as well as the interactions of Rise with Vowel and with Source, produced a single significant effect, namely that of Rise. Stimuli with late rises were significantly more likely to exhibit a trajectory reversal (early: 0.283, late 0.399; &#967;<sup>2</sup>(1) = 21.82, <italic>p</italic> &lt; 0.0001). In other words, as anticipated by prediction 1, participants were significantly more likely to track towards the statement response box before making a question response in the context of late rises than in the context of early rises. If greater likelihood of a trajectory reversal is symptomatic of ambiguous stimuli, as claimed by Freeman (<xref ref-type="bibr" rid="B17">2014</xref>) in connection with his results for sex categorization of faces, then the late rise stimuli are more ambiguous than the early rise stimuli. This is further confirmation of the asymmetry in the signal value of the two rise types that was reflected in the response time data reported in connection with Figure <xref ref-type="fig" rid="F4">4</xref>.</p>
<p>However, an alternative explanation of this result is that in the case of an early rise, the information that indicates a question, i.e., the rise, becomes available earlier, and as a consequence there is less opportunity for the alternative response to compete. In the late condition, on the other hand, increased attraction to the statement response might be a result of the rise information becoming available later, with the information before the rise remaining compatible with the statement response. That is, the finding that the move to the question response is more direct with early rises (and potentially also the overall effect in response choice shown in Figure <xref ref-type="fig" rid="F3">3</xref>) may not be due to a functional distinction between early and late rises, but might simply be a consequence of when the high pitch information becomes available in the task being performed by participants for this experiment.</p>
<p>Such an explanation would seem to be discounted by a closer analysis of the MD measure in the question responses. Of particular interest is whether there are measurable influences on MD of factors other than rise alignment within the high MD distribution, i.e., in the set where the competition effects of the alternative response are evident. A further regression model was therefore performed on the high MD group, with MD as the dependent variable. The factors explored were Vowel, Rise, Source, and Serial Position, as well as the interactions of Rise with Vowel and with Source. Source had no significant impact, neither as a simple effect nor in interaction, indicating that trajectory movements towards the competing responses were not influenced by uncontrolled properties distinguishing the original sets of statement and question utterances from which the test items were derived. There was however a significant interaction of Vowel and Rise (&#967;<sup>2</sup>(1) = 8.79, <italic>p</italic> &lt; 0.005), as shown in Figure <xref ref-type="fig" rid="F7">7</xref>. Stimuli with early rises exhibited less attraction to the alternative statement response (i.e., had smaller MD values) when the rise followed an [i&#600;] vowel than when it followed an [e&#600;] vowel. In addition, Figure <xref ref-type="fig" rid="F7">7</xref> indicates that after [i&#600;], question responses to early rises show less attraction to the statement response than question responses to late rises. These interaction effects involving Vowel and Rise were anticipated by prediction 3.</p>
<fig id="F7">
<label>Figure 7</label>
<caption>
<p>Maximum Deviation by Rise and Vowel, for question responses in the high MD set.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="/article/id/6198/file/74781/"/>
</fig>
</sec>
</sec>
<sec>
<title>4 Discussion</title>
<p>The experiment reported in this paper was designed to test whether the nature of a segmental phonetic variable influences listeners&#8217; interpretation of prosodic variation. The experiment exploited the co-variation, predominantly among younger speakers, of a closer articulation of the <sc>SQUARE</sc> diphthong (approximating the <sc>NEAR</sc> vowel, i.e., [i&#600;]) and final statement rises (uptalk) in NZE, and in particular the recent finding that earlier rises may provide a possible means of distinguishing questions from uptalk statements (<xref ref-type="bibr" rid="B45">Warren, 2005</xref>; <xref ref-type="bibr" rid="B47">Warren &amp; Daly, 2005</xref>; <xref ref-type="bibr" rid="B48">Warren &amp; Fletcher, 2016</xref>). Therefore two aspects of prosodic variability are involved&#8212;the variable use of final rises to signal either questions or statements in the same speech community, and the variability in the alignment of the final rise that is trending towards a marker that distinguishes these sentence types. Such variability becomes informative once it can be readily disentangled, so that there is greater clarity over whether a final rise indicates a question or a statement. The current study aimed to see whether the segmental phonetic cue in the test utterances would indicate whether or not the speaker is from a group of speakers that is likely firstly to produce uptalk and secondly to distinguish uptalk from question rises through alignment difference in the rise starting point.</p>
<p>The predictions set out in section 2.4 received good support. In line with previous perceptual research (<xref ref-type="bibr" rid="B52">Zwartz &amp; Warren, 2003</xref>), the experiment provided further evidence that early rises are more clearly associated with questions and late rises with statements. This was reflected in the analysis of the response choice data, which showed a clear effect of rise type on the proportion of question responses. Question responses after early rises were also faster and showed a more direct mouse trajectory than question responses after late rises, and trajectory reversals were more likely after late rises, indicating increased competition from the statement response in this condition.</p>
<p>In addition, the manipulation of the <sc>SQUARE</sc> vowel successfully shifted performance in the forced-choice response task. The slower statement responses to items containing an [e&#600;] version of this vowel than to items containing an [i&#600;] version is an indication that although uptalk rises from more conservative speakers are not ruled out, they are less expected. The mouse-tracking data for question responses made after a trajectory reversal show that there is less competition from the statement response for an early rise after the [i&#600;] realization of <sc>SQUARE</sc> than for the same rise after the [e&#600;] realization. The trajectories also show that in the context of the [i&#600;] realization there is less competition from the statement response when the rise is early than when it is late. These findings of significant results in the mouse-tracking data, particularly in the absence of significant differences in overall response choice, support the value of the technique in adding a more qualitative aspect to the interpretation of participants&#8217; responses. In particular, these trajectory reversal data linked to the MD measure allow us to say more about the competition effects that exist between alternative response choices.</p>
<p>The results of this study indicate that variability in intonational rise alignment in NZE is meaningful, and is more closely tied to certain speaker groups than to others. Listener performance in the experiment reflects the parallel development in this variety of a segment-level merger and changes in the use and shape of an intonation pattern. The vowel merger has been able to happen largely because the consequences are not extreme&#8212;although letters to newspaper editors might suggest otherwise (e.g., complaints about hearing of people &#8220;crossing on the Cook Strait &#8216;Fairy&#8217; and flying &#8216;Ear&#8217; New Zealand&#8221; [<xref ref-type="bibr" rid="B9">Bravery, 2001</xref>]), confusion between words with <sc>NEAR</sc> or <sc>SQUARE</sc> tokens is not frequent, partly because the vowel sounds are relatively rare (17th and 18th most frequent out of 20 English vowels, <xref ref-type="bibr" rid="B22">Gimson, 1963</xref>), but also because there are few relevant minimal pairs, and because utterance contexts usually serve to disambiguate, just as they do for other homophones. Nevertheless, the merger is still strongly age-graded, and so the use of the [i&#600;] or [e&#600;] variant of the <sc>SQUARE</sc> vowel has potential as a marker of social grouping.</p>
<p>On the other hand, the prosodic variability considered here has the potential to cause confusion between two radically different sentence types&#8212;questions and statements. This is reflected again in complaints from the public, but also in the frequent definition of uptalk using terms such as &#8216;question-like intonation on a declarative utterance.&#8217; Such a definition is problematic on a number of counts. The notion of &#8216;question&#8217; is ambiguous between function and form, and the belief that there is any single type of &#8216;question intonation&#8217; is misplaced, as is the assumption that questions have to have rising intonation. Even if we accept that &#8216;question-like intonation&#8217; stands for the type of rising intonation often found on yes-no questions, there is lack of agreement about whether that type of rising intonation is identical or even similar to uptalk. Some of this disagreement may result from the presence of different styles and uses of uptalk in different English varieties (<xref ref-type="bibr" rid="B46">Warren, 2016</xref>). It is worth remembering also that different varieties may have reached different stages in the development and use of uptalk, with some varieties (e.g., NZE) beginning to show novel distinctions, while others (such as British English, <xref ref-type="bibr" rid="B41">Shobbrook &amp; House, 2003</xref>) do not.</p>
<p>Ladd (<xref ref-type="bibr" rid="B33">2008: 126</xref>) acknowledges that uptalk and question rises may differ, but that &#8220;the differences are subtle, and arguably gradient,&#8221; and that it may be plausible to &#8220;analyse HRT statements as having the same phonological representations as high-rising question contours.&#8221; The results of the current study show that the &#8216;arguably gradient&#8217; distinction between early and late rise in NZE does have the potential to signal a distinction in meaning, and that uptalk and question contours may have different phonological representations. As pointed out by House (<xref ref-type="bibr" rid="B30">2006</xref>), the development of a new phonological distinction is a natural resolution of a potentially confusing situation in which one pattern can have multiple functions. Note though that other factors may make such a split unlikely to happen precipitously. As Guy et al. (<xref ref-type="bibr" rid="B24">1986</xref>) pointed out in their discussion of Australian English, even without a phonetic difference between uptalk rises and question rises, ambiguity is unlikely, as contexts will clarify the intended meaning. For instance, true questions usually include some anaphoric reference to a previous utterance, and will usually be at the end of a turn, while uptalk tends to provide new information, with the speaker typically continuing to hold the floor. &#8220;If these clear contextual and textual differences were not sufficient to disambiguate between the two meanings, we might expect structural change to occur: For example, phonetic differentiation of the contours or disuse of the contour for one of the meanings&#8221; (<xref ref-type="bibr" rid="B24">Guy et al., 1986: 27</xref>). It may be that recent phonetic differences between uptalk and question intonation that have been noted for Australian and New Zealand English indicate that a level of potential confusion has been reached that requires such differentiation.</p>
</sec>
<sec sec-type="supplementary-material">
<title>Additional Files</title>
<p>The additional file for this article can be found as follows:</p>
<supplementary-material id="S1" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://doi.org/10.5334/labphon.92.s1">
<label>File 1</label>
<caption>
<p>Examples of the four stimuli based on the sentence illustrated Figure 1. DOI: <uri>https://doi.org/10.5334/labphon.92.s1</uri></p>
</caption>
</supplementary-material>
<supplementary-material id="S1" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://doi.org/10.5334/labphon.92.s2">
<label>File 2</label>
<caption>
<p>List of test items used in the experiment. DOI: <uri>https://doi.org/10.5334/labphon.92.s2</uri></p>
</caption>
</supplementary-material>
</sec>
</body>
<back>
<fn-group>
<fn id="n1"><p>The lexical-set labels <sc>SQUARE</sc> and <sc>NEAR</sc> are used in this paper, following Wells (<xref ref-type="bibr" rid="B51">1982</xref>), to refer to the relevant centering diphthongs.</p></fn>
<fn id="n2"><p>All stimuli were prepared using source utterances from the same female speaker, so the experiment reported here cannot explicitly test the differences between female and male speakers in the realization of either of the phonetic variables.</p></fn>
<fn id="n3"><p>Although an eye-tracking paradigm might have been able to reveal similar patterns of behaviour, the mouse-tracking technique has the advantage of being considerably cheaper in terms both of equipment and of analysis costs.</p></fn>
</fn-group>
<ack>
<title>Acknowledgements</title>
<p>I am grateful to Chigusa Kurumada and to three anonymous reviewers for their comments on earlier versions of this article. My thanks also go to participants at <italic>Variation and Language Processing 2</italic> and at the <italic>15th Australasian International Conference on Speech Science and Technology</italic> for feedback on presentations of the research reported here.</p>
</ack>
<sec>
<title>Competing Interests</title>
<p>The author has no competing interests to declare.</p>
</sec>
<ref-list>
<ref id="B1">
<label>1</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Barca</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Benedetti</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Pezzulo</surname>
<given-names>G.</given-names>
</name>
</person-group>
<article-title>The effects of phonological similarity on the semantic categorisation of pictorial and lexical stimuli: Evidence from continuous behavioural measures</article-title>
<source>Journal of Cognitive Psychology</source>
<year iso-8601-date="2016">2016</year>
<volume>28</volume>
<issue>2</issue>
<fpage>159</fpage>
<lpage>170</lpage>
<pub-id pub-id-type="doi">10.1080/20445911.2015.1101117</pub-id>
</element-citation>
</ref>
<ref id="B2">
<label>2</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Barca</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Pezzulo</surname>
<given-names>G.</given-names>
</name>
</person-group>
<article-title>Unfolding visual lexical decision in time</article-title>
<source>PLoS ONE</source>
<year iso-8601-date="2012">2012</year>
<volume>7</volume>
<issue>4</issue>
<fpage>e35932</fpage>
<pub-id pub-id-type="doi">10.1371/journal.pone.0035932</pub-id>
</element-citation>
</ref>
<ref id="B3">
<label>3</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Barca</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Pezzulo</surname>
<given-names>G.</given-names>
</name>
</person-group>
<article-title>Tracking second thoughts: Continuous and discrete revision processes during visual lexical decision</article-title>
<source>PLoS ONE</source>
<year iso-8601-date="2015">2015</year>
<volume>10</volume>
<issue>2</issue>
<fpage>e0116193</fpage>
<pub-id pub-id-type="doi">10.1371/journal.pone.0116193</pub-id>
</element-citation>
</ref>
<ref id="B4">
<label>4</label>
<element-citation publication-type="webpage">
<person-group person-group-type="author">
<name>
<surname>Bates</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Maechler</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Bolker</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Walker</surname>
<given-names>S.</given-names>
</name>
</person-group>
<article-title>lme4 (Version 1.1&#8211;8)</article-title>
<year iso-8601-date="2015">2015</year>
<comment>Retrieved from <uri>http://CRAN.R-project.org/package=lme4</uri></comment>
</element-citation>
</ref>
<ref id="B5">
<label>5</label>
<element-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Bauer</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Warren</surname>
<given-names>P.</given-names>
</name>
</person-group>
<person-group person-group-type="editor">
<name>
<surname>Kortmann</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Schneider</surname>
<given-names>E. W.</given-names>
</name>
<name>
<surname>Burridge</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Mesthrie</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Upton</surname>
<given-names>C.</given-names>
</name>
</person-group>
<chapter-title>New Zealand English: Phonology</chapter-title>
<source>A handbook of varieties of English: A multimedia reference tool</source>
<year iso-8601-date="2004">2004</year>
<publisher-loc>Berlin</publisher-loc>
<publisher-name>Mouton de Gruyter</publisher-name>
<volume>1</volume>
<fpage>580</fpage>
<lpage>602</lpage>
</element-citation>
</ref>
<ref id="B6">
<label>6</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bell</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Johnson</surname>
<given-names>G.</given-names>
</name>
</person-group>
<article-title>Towards a sociolinguistics of style</article-title>
<source>University of Pennsylvania Working Papers in Linguistics</source>
<year iso-8601-date="1997">1997</year>
<volume>4</volume>
<fpage>1</fpage>
<lpage>21</lpage>
</element-citation>
</ref>
<ref id="B7">
<label>7</label>
<element-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Benton</surname>
<given-names>R. A.</given-names>
</name>
</person-group>
<source>Research into the English language difficulties of Maori school children 1963&#8211;1964</source>
<year iso-8601-date="1965">1965</year>
<publisher-loc>Wellington</publisher-loc>
<publisher-name>M&#257;ori Education Foundation</publisher-name>
</element-citation>
</ref>
<ref id="B8">
<label>8</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Boersma</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Weenink</surname>
<given-names>D.</given-names>
</name>
</person-group>
<article-title>Praat: Doing phonetics by computer (Version 5.3.06)</article-title>
<year iso-8601-date="2012">2012</year>
</element-citation>
</ref>
<ref id="B9">
<label>9</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bravery</surname>
<given-names>L.</given-names>
</name>
</person-group>
<article-title>Letter to the Editor</article-title>
<source>New Zealand Listener</source>
<year iso-8601-date="2001">2001</year>
<month>June</month>
<day>9th</day>
</element-citation>
</ref>
<ref id="B10">
<label>10</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Britain</surname>
<given-names>D.</given-names>
</name>
</person-group>
<article-title>Linguistic change in intonation: The use of high rising terminals in New Zealand English</article-title>
<source>Language Variation and Change</source>
<year iso-8601-date="1992">1992</year>
<volume>4</volume>
<fpage>77</fpage>
<lpage>104</lpage>
<pub-id pub-id-type="doi">10.1017/S0954394500000661</pub-id>
</element-citation>
</ref>
<ref id="B11">
<label>11</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dale</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Duran</surname>
<given-names>N. D.</given-names>
</name>
</person-group>
<article-title>The cognitive dynamics of negated sentence verification</article-title>
<source>Cognitive Science</source>
<year iso-8601-date="2011">2011</year>
<volume>35</volume>
<issue>5</issue>
<fpage>983</fpage>
<lpage>996</lpage>
<pub-id pub-id-type="doi">10.1111/j.1551-6709.2010.01164.x</pub-id>
</element-citation>
</ref>
<ref id="B12">
<label>12</label>
<element-citation publication-type="thesis">
<person-group person-group-type="author">
<name>
<surname>Dorrington</surname>
<given-names>N.</given-names>
</name>
</person-group>
<source>&#8216;Speaking up&#8217;: A comparative investigation into the onset of uptalk in General South African English (BA[Honours] thesis)</source>
<year iso-8601-date="2010">2010</year>
<publisher-name>Rhodes University</publisher-name>
</element-citation>
</ref>
<ref id="B13">
<label>13</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Drager</surname>
<given-names>K.</given-names>
</name>
</person-group>
<article-title>Speaker age and vowel perception</article-title>
<source>Language &amp; Speech</source>
<year iso-8601-date="2011">2011</year>
<volume>54</volume>
<issue>1</issue>
<fpage>99</fpage>
<lpage>121</lpage>
<pub-id pub-id-type="doi">10.1177/0023830910388017</pub-id>
</element-citation>
</ref>
<ref id="B14">
<label>14</label>
<element-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Fletcher</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Grabe</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Warren</surname>
<given-names>P.</given-names>
</name>
</person-group>
<person-group person-group-type="editor">
<name>
<surname>Jun</surname>
<given-names>S.-A.</given-names>
</name>
</person-group>
<chapter-title>Intonational variation in four dialects of English: The high rising tone</chapter-title>
<source>Prosodic typology: The phonology of intonation and phrasing</source>
<year iso-8601-date="2005">2005</year>
<publisher-loc>Oxford</publisher-loc>
<publisher-name>Oxford University Press</publisher-name>
<fpage>390</fpage>
<lpage>409</lpage>
<pub-id pub-id-type="doi">10.1093/acprof:oso/9780199249633.003.0014</pub-id>
</element-citation>
</ref>
<ref id="B15">
<label>15</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fletcher</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Harrington</surname>
<given-names>J.</given-names>
</name>
</person-group>
<article-title>High-rising terminals and fall-rise tunes in Australian English</article-title>
<source>Phonetica</source>
<year iso-8601-date="2001">2001</year>
<volume>58</volume>
<issue>4</issue>
<fpage>215</fpage>
<lpage>229</lpage>
<pub-id pub-id-type="doi">10.1159/000046176</pub-id>
</element-citation>
</ref>
<ref id="B16">
<label>16</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fletcher</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Loakes</surname>
<given-names>D.</given-names>
</name>
</person-group>
<article-title>Interpreting rising intonation in Australian English</article-title>
<source>Speech Prosody 2010</source>
<year iso-8601-date="2010">2010</year>
<fpage>1</fpage>
<lpage>4</lpage>
<comment>100124</comment>
</element-citation>
</ref>
<ref id="B17">
<label>17</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Freeman</surname>
<given-names>J. B.</given-names>
</name>
</person-group>
<article-title>Abrupt category shifts during real-time person perception</article-title>
<source>Psychonomic Bulletin &amp; Review</source>
<year iso-8601-date="2014">2014</year>
<volume>21</volume>
<issue>1</issue>
<fpage>85</fpage>
<lpage>92</lpage>
<pub-id pub-id-type="doi">10.3758/s13423-013-0470-8</pub-id>
</element-citation>
</ref>
<ref id="B18">
<label>18</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Freeman</surname>
<given-names>J. B.</given-names>
</name>
<name>
<surname>Ambady</surname>
<given-names>N.</given-names>
</name>
</person-group>
<article-title>MouseTracker: Software for studying real-time mental processing using a computer mouse-tracking method</article-title>
<source>Behavior Research Methods</source>
<year iso-8601-date="2010">2010</year>
<volume>42</volume>
<issue>1</issue>
<fpage>226</fpage>
<lpage>241</lpage>
<pub-id pub-id-type="doi">10.3758/BRM.42.1.226</pub-id>
</element-citation>
</ref>
<ref id="B19">
<label>19</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Freeman</surname>
<given-names>J. B.</given-names>
</name>
<name>
<surname>Ambady</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Rule</surname>
<given-names>N. O.</given-names>
</name>
<name>
<surname>Johnson</surname>
<given-names>K. L.</given-names>
</name>
</person-group>
<article-title>Will a category cue attract you? Motor output reveals dynamic competition across person construal</article-title>
<source>Journal of Experimental Psychology: General</source>
<year iso-8601-date="2008">2008</year>
<volume>137</volume>
<issue>4</issue>
<fpage>673</fpage>
<lpage>690</lpage>
<pub-id pub-id-type="doi">10.1037/a0013875</pub-id>
</element-citation>
</ref>
<ref id="B20">
<label>20</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Freeman</surname>
<given-names>J. B.</given-names>
</name>
<name>
<surname>Dale</surname>
<given-names>R.</given-names>
</name>
</person-group>
<article-title>Assessing bimodality to detect the presence of a dualcognitive process</article-title>
<source>Behavior Research Methods</source>
<year iso-8601-date="2013">2013</year>
<volume>45</volume>
<issue>1</issue>
<fpage>83</fpage>
<lpage>97</lpage>
<pub-id pub-id-type="doi">10.3758/s13428-012-0225-x</pub-id>
</element-citation>
</ref>
<ref id="B21">
<label>21</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Freeman</surname>
<given-names>J. B.</given-names>
</name>
<name>
<surname>Penner</surname>
<given-names>A. M.</given-names>
</name>
<name>
<surname>Saperstein</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Scheutz</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ambady</surname>
<given-names>N.</given-names>
</name>
</person-group>
<article-title>Looking the part: Social status cues shape race perception</article-title>
<source>PLoS ONE</source>
<year iso-8601-date="2011">2011</year>
<volume>6</volume>
<fpage>e25107</fpage>
<pub-id pub-id-type="doi">10.1371/journal.pone.0025107</pub-id>
</element-citation>
</ref>
<ref id="B22">
<label>22</label>
<element-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Gimson</surname>
<given-names>A. C.</given-names>
</name>
</person-group>
<source>An Introduction to the Pronunciation of English</source>
<year iso-8601-date="1963">1963</year>
<publisher-loc>London</publisher-loc>
<publisher-name>Edward Arnold</publisher-name>
</element-citation>
</ref>
<ref id="B23">
<label>23</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gordon</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Maclagan</surname>
<given-names>M.</given-names>
</name>
</person-group>
<article-title>Capturing a sound change: A real time study over 15 years of the NEAR/SQUARE diphthong merger in New Zealand English</article-title>
<source>Australian Journal of Linguistics</source>
<year iso-8601-date="2001">2001</year>
<volume>21</volume>
<issue>2</issue>
<fpage>215</fpage>
<lpage>238</lpage>
<pub-id pub-id-type="doi">10.1080/07268600120080578</pub-id>
</element-citation>
</ref>
<ref id="B24">
<label>24</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guy</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Horvath</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Vonwiller</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Daisley</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Rogers</surname>
<given-names>I.</given-names>
</name>
</person-group>
<article-title>An intonational change in progress in Australian English</article-title>
<source>Language in Society</source>
<year iso-8601-date="1986">1986</year>
<volume>15</volume>
<issue>1</issue>
<fpage>23</fpage>
<lpage>51</lpage>
<pub-id pub-id-type="doi">10.1017/S0047404500011635</pub-id>
</element-citation>
</ref>
<ref id="B25">
<label>25</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hartigan</surname>
<given-names>J. A.</given-names>
</name>
<name>
<surname>Hartigan</surname>
<given-names>P. M.</given-names>
</name>
</person-group>
<article-title>The dip test of unimodality</article-title>
<source>The Annals of Statistics</source>
<year iso-8601-date="1985">1985</year>
<volume>13</volume>
<fpage>70</fpage>
<lpage>84</lpage>
<pub-id pub-id-type="doi">10.1214/aos/1176346577</pub-id>
</element-citation>
</ref>
<ref id="B26">
<label>26</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hay</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Drager</surname>
<given-names>K.</given-names>
</name>
</person-group>
<article-title>Stuffed toys and speech perception</article-title>
<source>Linguistics</source>
<year iso-8601-date="2010">2010</year>
<volume>48</volume>
<issue>4</issue>
<fpage>865</fpage>
<lpage>892</lpage>
<pub-id pub-id-type="doi">10.1515/ling.2010.027</pub-id>
</element-citation>
</ref>
<ref id="B27">
<label>27</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hay</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Nolan</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Drager</surname>
<given-names>K.</given-names>
</name>
</person-group>
<article-title>From fush to feesh: Exemplar priming in speech perception</article-title>
<source>Linguistic Review</source>
<year iso-8601-date="2006a">2006a</year>
<volume>23</volume>
<issue>3</issue>
<fpage>351</fpage>
<lpage>379</lpage>
<pub-id pub-id-type="doi">10.1515/TLR.2006.014</pub-id>
</element-citation>
</ref>
<ref id="B28">
<label>28</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hay</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Warren</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Drager</surname>
<given-names>K.</given-names>
</name>
</person-group>
<article-title>Factors influencing speech perception in the context of a merger-in-progress</article-title>
<source>Journal of Phonetics</source>
<year iso-8601-date="2006b">2006b</year>
<volume>34</volume>
<issue>4</issue>
<fpage>458</fpage>
<lpage>484</lpage>
<pub-id pub-id-type="doi">10.1016/j.wocn.2005.10.001</pub-id>
</element-citation>
</ref>
<ref id="B29">
<label>29</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hehman</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Stolier</surname>
<given-names>R. M.</given-names>
</name>
<name>
<surname>Freeman</surname>
<given-names>J. B.</given-names>
</name>
</person-group>
<article-title>Advance mouse-tracking analytic techniques for enhancing psychological science</article-title>
<source>Group Processes and Intergroup Relations</source>
<year iso-8601-date="2014">2014</year>
<volume>18</volume>
<issue>3</issue>
<fpage>384</fpage>
<lpage>401</lpage>
<pub-id pub-id-type="doi">10.1177/1368430214538325</pub-id>
</element-citation>
</ref>
<ref id="B30">
<label>30</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>House</surname>
<given-names>J.</given-names>
</name>
</person-group>
<article-title>Constructing a context with intonation</article-title>
<source>Journal of Pragmatics</source>
<year iso-8601-date="2006">2006</year>
<volume>38</volume>
<issue>10</issue>
<fpage>1542</fpage>
<lpage>1558</lpage>
<pub-id pub-id-type="doi">10.1016/j.pragma.2005.07.005</pub-id>
</element-citation>
</ref>
<ref id="B31">
<label>31</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Johnson</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Strand</surname>
<given-names>E. A.</given-names>
</name>
<name>
<surname>D&#8217;Imperio</surname>
<given-names>M.</given-names>
</name>
</person-group>
<article-title>Auditory&#8211;visual integration of talker gender in vowel perception</article-title>
<source>Journal of Phonetics</source>
<year iso-8601-date="1999">1999</year>
<volume>27</volume>
<issue>4</issue>
<fpage>359</fpage>
<lpage>384</lpage>
<pub-id pub-id-type="doi">10.1006/jpho.1999.0100</pub-id>
</element-citation>
</ref>
<ref id="B32">
<label>32</label>
<element-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Ladd</surname>
<given-names>D. R.</given-names>
</name>
</person-group>
<source>Intonational Phonology</source>
<year iso-8601-date="1996">1996</year>
<publisher-loc>Cambridge</publisher-loc>
<publisher-name>Cambridge University Press</publisher-name>
</element-citation>
</ref>
<ref id="B33">
<label>33</label>
<element-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Ladd</surname>
<given-names>D. R.</given-names>
</name>
</person-group>
<source>Intonational Phonology</source>
<year iso-8601-date="2008">2008</year>
<edition>2nd ed.</edition>
<publisher-loc>Cambridge</publisher-loc>
<publisher-name>Cambridge University Press</publisher-name>
<pub-id pub-id-type="doi">10.1017/CBO9780511808814</pub-id>
</element-citation>
</ref>
<ref id="B34">
<label>34</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lakoff</surname>
<given-names>R.</given-names>
</name>
</person-group>
<article-title>Language and woman&#8217;s place</article-title>
<source>Language in Society</source>
<year iso-8601-date="1973">1973</year>
<volume>2</volume>
<fpage>45</fpage>
<lpage>79</lpage>
<pub-id pub-id-type="doi">10.1017/S0047404500000051</pub-id>
</element-citation>
</ref>
<ref id="B35">
<label>35</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Levon</surname>
<given-names>E.</given-names>
</name>
</person-group>
<article-title>Sexuality in context: Variation and the sociolinguistic perception of identity</article-title>
<source>Language in Society</source>
<year iso-8601-date="2007">2007</year>
<volume>36</volume>
<issue>4</issue>
<fpage>533</fpage>
<lpage>554</lpage>
<pub-id pub-id-type="doi">10.1017/S0047404507070431</pub-id>
</element-citation>
</ref>
<ref id="B36">
<label>36</label>
<element-citation publication-type="thesis">
<person-group person-group-type="author">
<name>
<surname>McGregor</surname>
<given-names>J.</given-names>
</name>
</person-group>
<source>High rising tunes in Australian English (Doctoral dissertation)</source>
<year iso-8601-date="2005">2005</year>
<publisher-loc>Sydney, Australia</publisher-loc>
<publisher-name>Macquarie University</publisher-name>
</element-citation>
</ref>
<ref id="B37">
<label>37</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Niedzielski</surname>
<given-names>N.</given-names>
</name>
</person-group>
<article-title>The effect of social information on the perception of sociolinguistic variables</article-title>
<source>Journal of Language and Social Psychology</source>
<year iso-8601-date="1999">1999</year>
<volume>18</volume>
<issue>1</issue>
<fpage>62</fpage>
<lpage>85</lpage>
<pub-id pub-id-type="doi">10.1177/0261927X99018001005</pub-id>
</element-citation>
</ref>
<ref id="B38">
<label>38</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Papesh</surname>
<given-names>M. H.</given-names>
</name>
<name>
<surname>Goldinger</surname>
<given-names>S. D.</given-names>
</name>
</person-group>
<article-title>Memory in motion: Movement dynamics reveal memory strength</article-title>
<source>Psychonomic Bulletin &amp; Review</source>
<year iso-8601-date="2012">2012</year>
<volume>19</volume>
<issue>5</issue>
<fpage>906</fpage>
<lpage>913</lpage>
<pub-id pub-id-type="doi">10.3758/s13423-012-0281-3</pub-id>
</element-citation>
</ref>
<ref id="B39">
<label>39</label>
<element-citation publication-type="webpage">
<person-group person-group-type="author">
<collab>Psychology Software Tools</collab>
</person-group>
<article-title>E-Prime (Version 2.0)</article-title>
<year iso-8601-date="2012">2012</year>
<comment>Retrieved from <uri>http://www.pstnet.com</uri></comment>
</element-citation>
</ref>
<ref id="B40">
<label>40</label>
<element-citation publication-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ritchart</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Arvaniti</surname>
<given-names>A.</given-names>
</name>
</person-group>
<person-group person-group-type="editor">
<name>
<surname>Campbell</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Gibbon</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Hirst</surname>
<given-names>D.</given-names>
</name>
</person-group>
<article-title>The form and use of uptalk in Southern Californian English</article-title>
<conf-name>Social and Linguistic Speech Prosody (Proceedings of the 7th international conference on Speech Prosody)</conf-name>
<year iso-8601-date="2014">2014</year>
<conf-loc>Dublin, Eire</conf-loc>
<fpage>331</fpage>
<lpage>335</lpage>
</element-citation>
</ref>
<ref id="B41">
<label>41</label>
<element-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Shobbrook</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>House</surname>
<given-names>J.</given-names>
</name>
</person-group>
<person-group person-group-type="editor">
<name>
<surname>Sol&#233;</surname>
<given-names>M. J.</given-names>
</name>
<name>
<surname>Recasens</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Romero</surname>
<given-names>J.</given-names>
</name>
</person-group>
<chapter-title>High rising tones in Southern British English</chapter-title>
<source>15th International Congress of Phonetic Sciences</source>
<year iso-8601-date="2003">2003</year>
<publisher-loc>Barcelona</publisher-loc>
<publisher-name>Universitat Autonoma de Barcelona</publisher-name>
<fpage>1273</fpage>
<lpage>1276</lpage>
</element-citation>
</ref>
<ref id="B42">
<label>42</label>
<element-citation publication-type="webpage">
<person-group person-group-type="author">
<name>
<surname>Singmann</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Bolker</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Westfall</surname>
<given-names>J.</given-names>
</name>
</person-group>
<article-title>afex: Analysis of Factorial Experiments (Version 0.13&#8211;145)</article-title>
<year iso-8601-date="2015">2015</year>
<comment>Retrieved from <uri>https://cran.r-project.org/web/packages/afex/index.html</uri></comment>
</element-citation>
</ref>
<ref id="B43">
<label>43</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Spivey</surname>
<given-names>M. J.</given-names>
</name>
<name>
<surname>Grosjean</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Knoblich</surname>
<given-names>G.</given-names>
</name>
</person-group>
<article-title>Continuous attraction toward phonological competitors</article-title>
<source>Proceedings of the National Academy of Sciences of the United States of America</source>
<year iso-8601-date="2005">2005</year>
<volume>102</volume>
<issue>29</issue>
<fpage>10393</fpage>
<lpage>10398</lpage>
<pub-id pub-id-type="doi">10.1073/pnas.0503903102</pub-id>
</element-citation>
</ref>
<ref id="B44">
<label>44</label>
<element-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Strand</surname>
<given-names>E. A.</given-names>
</name>
<name>
<surname>Johnson</surname>
<given-names>K.</given-names>
</name>
</person-group>
<person-group person-group-type="editor">
<name>
<surname>Gibbon</surname>
<given-names>D.</given-names>
</name>
</person-group>
<chapter-title>Gradient and visual speaker normalization in the perception of fricatives</chapter-title>
<source>Natural Language Processing and Speech Technology</source>
<year iso-8601-date="1996">1996</year>
<publisher-loc>Berlin</publisher-loc>
<publisher-name>Mouton de Gruyter</publisher-name>
<fpage>14</fpage>
<lpage>26</lpage>
<pub-id pub-id-type="doi">10.1515/9783110821895-003</pub-id>
</element-citation>
</ref>
<ref id="B45">
<label>45</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Warren</surname>
<given-names>P.</given-names>
</name>
</person-group>
<article-title>Patterns of late rising in New Zealand English: Intonation variation or intonational change?</article-title>
<source>Language Variation and Change</source>
<year iso-8601-date="2005">2005</year>
<volume>17</volume>
<fpage>209</fpage>
<lpage>230</lpage>
<pub-id pub-id-type="doi">10.1017/S095439450505009X</pub-id>
</element-citation>
</ref>
<ref id="B46">
<label>46</label>
<element-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Warren</surname>
<given-names>P.</given-names>
</name>
</person-group>
<source>Uptalk: The phenomenon of rising intonation</source>
<year iso-8601-date="2016">2016</year>
<publisher-loc>Cambridge</publisher-loc>
<publisher-name>Cambridge University Press</publisher-name>
<pub-id pub-id-type="doi">10.1017/CBO9781316403570</pub-id>
</element-citation>
</ref>
<ref id="B47">
<label>47</label>
<element-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Warren</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Daly</surname>
<given-names>N.</given-names>
</name>
</person-group>
<person-group person-group-type="editor">
<name>
<surname>Bell</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Harlow</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Starks</surname>
<given-names>D.</given-names>
</name>
</person-group>
<chapter-title>Characterising New Zealand intonation: Broad and narrow analysis</chapter-title>
<source>Languages of New Zealand</source>
<year iso-8601-date="2005">2005</year>
<publisher-loc>Wellington</publisher-loc>
<publisher-name>Victoria University Press</publisher-name>
</element-citation>
</ref>
<ref id="B48">
<label>48</label>
<element-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Warren</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Fletcher</surname>
<given-names>J.</given-names>
</name>
</person-group>
<person-group person-group-type="editor">
<name>
<surname>Barnes</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Brugos</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Shattuck-Hufnagel</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Veilleux</surname>
<given-names>N.</given-names>
</name>
</person-group>
<chapter-title>Phonetic differences between uptalk and question rises in two Antipodean English varieties</chapter-title>
<source>Speech Prosody 2016</source>
<year iso-8601-date="2016">2016</year>
<publisher-loc>Boston, MA</publisher-loc>
<fpage>148</fpage>
<lpage>152</lpage>
<pub-id pub-id-type="doi">10.21437/SpeechProsody.2016-31</pub-id>
</element-citation>
</ref>
<ref id="B49">
<label>49</label>
<element-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Warren</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Hay</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Thomas</surname>
<given-names>B.</given-names>
</name>
</person-group>
<person-group person-group-type="editor">
<name>
<surname>Cole</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hualde</surname>
<given-names>J. I.</given-names>
</name>
</person-group>
<chapter-title>The loci of sound change effects in recognition and perception</chapter-title>
<source>Laboratory Phonology</source>
<year iso-8601-date="2007">2007</year>
<publisher-loc>Berlin</publisher-loc>
<publisher-name>Mouton de Gruyter</publisher-name>
<volume>9</volume>
<fpage>87</fpage>
<lpage>112</lpage>
</element-citation>
</ref>
<ref id="B50">
<label>50</label>
<element-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Warren</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Rae</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Hay</surname>
<given-names>J.</given-names>
</name>
</person-group>
<person-group person-group-type="editor">
<name>
<surname>Sol&#233;</surname>
<given-names>M. J.</given-names>
</name>
<name>
<surname>Recasens</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Romero</surname>
<given-names>J.</given-names>
</name>
</person-group>
<chapter-title>Word recognition and sound merger: The case of the front-centering diphthongs in NZ English</chapter-title>
<source>15th International Congress of Phonetic Sciences</source>
<year iso-8601-date="2003">2003</year>
<publisher-loc>Barcelona</publisher-loc>
<publisher-name>Universitat Autonoma de Barcelona</publisher-name>
<fpage>2989</fpage>
<lpage>2992</lpage>
</element-citation>
</ref>
<ref id="B51">
<label>51</label>
<element-citation publication-type="book">
<person-group person-group-type="author">
<name>
<surname>Wells</surname>
<given-names>J. C.</given-names>
</name>
</person-group>
<source>Accents of English</source>
<year iso-8601-date="1982">1982</year>
<publisher-loc>Cambridge, England</publisher-loc>
<publisher-name>Cambridge University Press</publisher-name>
<pub-id pub-id-type="doi">10.1017/CBO9780511611759</pub-id>
</element-citation>
</ref>
<ref id="B52">
<label>52</label>
<element-citation publication-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zwartz</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Warren</surname>
<given-names>P.</given-names>
</name>
</person-group>
<article-title>This is a statement? Lateness of rise as a factor in listener interpretation of HRTs</article-title>
<source>Wellington Working Papers in Linguistics</source>
<year iso-8601-date="2003">2003</year>
<volume>15</volume>
<fpage>51</fpage>
<lpage>62</lpage>
</element-citation>
</ref>
</ref-list>
</back>
</article>