<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20120330//EN" "http://jats.nlm.nih.gov/publishing/1.2/JATS-journalpublishing1.dtd">
<!--<?xml-stylesheet type="text/xsl" href="article.xsl"?>-->
<article article-type="research-article" dtd-version="1.2" xml:lang="en" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance">
<front>
<journal-meta>
<journal-id journal-id-type="issn">1868-6354</journal-id>
<journal-title-group>
<journal-title>Laboratory Phonology: Journal of the Association for Laboratory Phonology</journal-title>
</journal-title-group>
<issn pub-type="epub">1868-6354</issn>
<publisher>
<publisher-name>Open Library of Humanities</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.16995/labphon.7943</article-id>
<article-categories>
<subj-group>
<subject>Journal article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Proto-Lexicon Size and Phonotactic Knowledge are Linked in Non-M&#257;ori Speaking New Zealand Adults</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Panther</surname>
<given-names>Forrest</given-names>
</name>
<email>forrest.panther@canterbury.ac.nz</email>
<xref ref-type="aff" rid="aff-1">1</xref>
<xref ref-type="corresp" rid="cor-1">*</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Mattingley</surname>
<given-names>Wakayo</given-names>
</name>
<email>wakayo.mattingley@canterbury.ac.nz</email>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Todd</surname>
<given-names>Simon</given-names>
</name>
<email>sjtodd@ucsb.edu</email>
<xref ref-type="aff" rid="aff-2">2</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Hay</surname>
<given-names>Jennifer</given-names>
</name>
<email>jen.hay@canterbury.ac.nz</email>
<xref ref-type="aff" rid="aff-1">1</xref>
<xref ref-type="aff" rid="aff-3">3</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>King</surname>
<given-names>Jeanette</given-names>
</name>
<email>j.king@canterbury.ac.nz</email>
<xref ref-type="aff" rid="aff-1">1</xref>
<xref ref-type="aff" rid="aff-4">4</xref>
</contrib>
</contrib-group>
<aff id="aff-1"><label>1</label>New Zealand Institute of Language, Brain and Behaviour, NZILBB, University of Canterbury, New Zealand</aff>
<aff id="aff-2"><label>2</label>Department of Linguistics, University of California, Santa Barbara, USA</aff>
<aff id="aff-3"><label>3</label>Department of Linguistics, University of Canterbury, New Zealand</aff>
<aff id="aff-4"><label>4</label>Aotahi: School of M&#257;ori and Indigenous Studies, University of Canterbury, New Zealand</aff>
<author-notes>
<corresp id="cor-1"><label>*</label>Corresponding author.</corresp>
</author-notes>
<pub-date publication-format="electronic" date-type="pub" iso-8601-date="2023-02-01">
<day>01</day>
<month>02</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>14</volume>
<issue>1</issue>
<fpage>1</fpage>
<lpage>27</lpage>
<permissions>
<copyright-statement>Copyright: &#x00A9; 2023 The Author(s)</copyright-statement>
<copyright-year>2023</copyright-year>
<license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/4.0/">
<license-p>This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International License (CC-BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. See <uri xlink:href="http://creativecommons.org/licenses/by/4.0/">http://creativecommons.org/licenses/by/4.0/</uri>.</license-p>
</license>
</permissions>
<self-uri xlink:href="http://www.journal-labphon.org/articles/10.16995/labphon.7943/"/>
<abstract>
<p>Most people in New Zealand are exposed to the M&#257;ori language on a regular basis, but do not speak it. It has recently been claimed that this exposure leads them to create a large proto-lexicon, consisting of implicit memories of words and word parts, without semantic knowledge. This yields sophisticated phonotactic knowledge (<xref ref-type="bibr" rid="B42">Oh, Todd, Beckner, Hay, King, &amp; Needle, 2020</xref>). This claim was supported by two tasks in which Non-M&#257;ori-Speaking New Zealanders: (i) Distinguished real words from phonotactically matched non-words, suggesting lexical knowledge; (ii) Gave wellformedness ratings of non-words almost indistinguishable from those of fluent M&#257;ori speakers, demonstrating phonotactic knowledge.</p>
<p>Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>) ran these tasks on separate participants. While they hypothesised that phonotactic and lexical knowledge derived from the proto-lexicon, they did not establish a direct link between them. We replicate the two tasks, with improved stimuli, on the same set of participants. We find a statistically significant link between the tasks: Participants with a larger proto-lexicon (evidenced by performance in the Word Identification Task) show greater sensitivity to phonotactics in the Wellformedness Rating Task. This extends the previously reported results, increasing the evidence that exposure to a language you do not speak can lead to large-scale implicit knowledge about that language.</p>
</abstract>
</article-meta>
</front>
<body>
<sec>
<title>1 Introduction</title>
<p>In this paper, we present the results of a two-task experiment that explored the knowledge that Non-M&#257;ori Speaking New Zealanders (henceforth NMSs) have of M&#257;ori lexemes and phonotactics. This experiment extends the results of a previous study consisting of two experiments (<xref ref-type="bibr" rid="B42">Oh et al., 2020</xref>, described in section 1.3 below). The results of our experiment, as with the previous paper, show that NMSs have extensive knowledge of M&#257;ori words, and well-developed intuitions about M&#257;ori phonotactics. We go beyond the previous research in demonstrating a statistical association between individual participants&#8217; performance in these tasks, thus providing evidence in support of the hypothesis that these two forms of knowledge are causally linked.</p>
<p>First, we provide evidence in support of an adult &#8216;proto-lexicon,&#8217; analogous to the proto-lexicon that infants develop for their first language. This proto-lexicon contains knowledge of wordforms that need not be associated with detailed morpho-syntactic or semantic content. The existence of this proto-lexicon among non-speakers of a language has important implications for understanding the development of multilingualism, particularly in terms of identifying the stages of development from monolingualism to bilingualism, and the extent to which latent long-term exposure results in comprehensive linguistic knowledge of a non-dominant language.</p>
<p>Second, we provide further supporting evidence in favour of the hypothesis that non-M&#257;ori speakers have extensive phonotactic knowledge. As with Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>), we show that non-M&#257;ori speakers are sensitive to phonotactics when rating the wellformedness of non-words.</p>
<p>Finally, we go beyond Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>) by demonstrating that the proto-lexion and phonotactic knowledge are linked in the same speakers. We interpret this as providing support for the assumption that the phonotactic knowledge is generated from the proto-lexicon. Participants with a larger proto-lexicon (as evidenced by their ability to distinguish real words from similiar non-words) have more sophisticated phonotactic knowledge, and are thus more sensitive to phonotactics when rating the wellformedness of non-words.</p>
<sec>
<title>1.1 Language situation in New Zealand</title>
<p>M&#257;ori is the indigenous language of New Zealand. It is a Polynesian language in the Austronesian language family. In the 2018 New Zealand Census, 185,955 people self-reported as M&#257;ori speakers, representing about 3.9% of the total population of New Zealand (<xref ref-type="bibr" rid="B56">Statistics New Zealand, 2020</xref>).</p>
<p>M&#257;ori has a very small phoneme inventory, with ten consonant phonemes /p, t, k, m, n, &#331;, w, f, r, h/ and five vowels /i, e, a, o, u/. The five vowels also have long forms, which are indicated orthographically with a macron over the vowel. Because of this small inventory, when the M&#257;ori alphabet was developed by missionaries in the 1820s (<xref ref-type="bibr" rid="B45">Parkinson, 2016</xref>) the sound to spelling relationship was able to be fairly transparent, with the consonants represented by the letters &lt;p, t, k, m, n, ng, w, wh, r, h&gt;, and the vowels represented by &lt;i, e, a, o, u&gt;. Thus, written words can be assumed to access a similar basis of knowledge representation as spoken words. This assumption is supported by psychological research that shows that reading involves phonological awareness (<xref ref-type="bibr" rid="B70">Ziegler &amp; Goswami, 2005</xref>). In the experiments presented in this paper, we exploit this assumption by using written stimuli to probe phonological knowledge.</p>
<p>M&#257;ori vocabulary, as well as the use of M&#257;ori language, is pervasive in media, educational and cultural contexts in New Zealand. M&#257;ori expressions, including greetings, are largely and increasingly normalised. There are television stations and radio programs that broadcast in M&#257;ori, and in public buildings there is normally at least some M&#257;ori used or represented in writing. There are also public initiatives relating to the M&#257;ori language, including M&#257;ori Language Week. However, M&#257;ori is not a dominant language, and fluency is relatively rare. Consequently, people living in New Zealand receive consistent, low-level exposure to M&#257;ori, without ever learning to understand or speak it. The average non-M&#257;ori-speaking New Zealander is estimated to have an active M&#257;ori vocabulary of fewer than 100 words, as they can accurately match fewer than 100 words to their meaning in a multi-choice setting (<xref ref-type="bibr" rid="B34">Macalister, 2004</xref>).</p>
<p>This situation of a reasonable level of exposure without depth of vocabulary knowledge raises questions about the extent of the knowledge of M&#257;ori that NMSs have. How much phonological and lexical knowledge does this exposure result in? It is possible that the development of M&#257;ori knowledge from background exposure may reflect processes that occur the earliest stages of language acquisition. These are outlined in the next section.</p>
</sec>
<sec>
<title>1.2 Acquiring a lexicon</title>
<sec>
<title>1.2.1 Statistical word segmentation</title>
<p>Literature on language acquisition has identified the importance of phonology in both first and second acquisition (<xref ref-type="bibr" rid="B15">Curtin &amp; Hufnagle, 2009</xref>; <xref ref-type="bibr" rid="B52">Saffran, Werker, &amp; Werner, 2007</xref>). The literature on the first language acquisition of phonology is primarily focused on the acquisition of phonemes (e.g., <xref ref-type="bibr" rid="B33">Kuhl, 2004</xref>; <xref ref-type="bibr" rid="B66">Wang, Seidl, &amp; Cristia, 2021</xref>) and prosodic structure (<xref ref-type="bibr" rid="B17">Demuth, 2009</xref>). Other literature identifies the early acquisition of the ability to perceive recurring patterns in word forms (<xref ref-type="bibr" rid="B9">Chambers, Onishi, &amp; Fisher, 2003</xref>; <xref ref-type="bibr" rid="B12">Coady &amp; Aslin, 2004</xref>), even before children learn how to speak (<xref ref-type="bibr" rid="B48">Pierrehumbert, 2003</xref>). Before infants develop productive linguistic skills, they develop intricate speech perception abilities (<xref ref-type="bibr" rid="B14">Curtin &amp; Archer, 2015</xref>). These perceptual skills underpin and aid in the development of later productive linguistic skills. Research has shown that by the age of six months, infants are already perceptually attuned to the phonology and prosody of their native language (<xref ref-type="bibr" rid="B47">Perszyk &amp; Waxman, 2019</xref>; <xref ref-type="bibr" rid="B27">Johnson, Seidl, &amp; Tyler, 2014</xref>; <xref ref-type="bibr" rid="B55">Shi, Werker, &amp; Morgan, 1999</xref>; <xref ref-type="bibr" rid="B54">Shi, Morgan, &amp; Allopenna, 1998</xref>). The effect of this is that even before infants begin to learn to use language, they have well-developed skills around the phonological structure of the language they are learning.</p>
<p>One key early step in acquiring a lexicon is identifying the boundaries between words. Infants younger than one use multiple cues to parse words in online speech, including word frequency (<xref ref-type="bibr" rid="B6">Bortfeld, Morgan, Golinkoff, &amp; Rathbun, 2005</xref>), position in the utterance (<xref ref-type="bibr" rid="B27">Johnson et al., 2014</xref>), syllable transitional probabilities (<xref ref-type="bibr" rid="B51">Saffran, Aslin, &amp; Newport, 1996</xref>), segmental configuration (<xref ref-type="bibr" rid="B2">Archer &amp; Curtin, 2016</xref>, <xref ref-type="bibr" rid="B1">2011</xref>; <xref ref-type="bibr" rid="B35">MacKenzie, Curtin, &amp; Graham, 2012</xref>), and probabilistic phonotactics (<xref ref-type="bibr" rid="B2">Archer &amp; Curtin, 2016</xref>; <xref ref-type="bibr" rid="B37">Mattys &amp; Jusczyk, 2001</xref>; <xref ref-type="bibr" rid="B69">Zamuner, 2009</xref>). On this latter point, infants use statistical information about the probabilities of segmental sequences to identify word boundaries. Statistical information about sound sequences is therefore important from the very early stages of language acquisition.</p>
<p>Laboratory experiments reveal that adults also use these statistical segmentation strategies to start to identify word boundaries in artifical languages that they are exposed to in an experimental setting (<xref ref-type="bibr" rid="B46">Pe&#241;a, Bonatti, Nespor, &amp; Mehler, 2002</xref>; <xref ref-type="bibr" rid="B43">Onnis, Monaghan, Richmond, &amp; Chater, 2005</xref>; <xref ref-type="bibr" rid="B40">Newport &amp; Aslin, 2000</xref>). Thus, it appears human listeners across the lifespan relatively automatically orient to a speech stream, and will unconsciously track statistical properties of that speech stream, in order to begin one of the earliest stages of language acquisition: Word segmentation.</p>
</sec>
<sec>
<title>1.2.2 The proto-lexicon</title>
<p>Once the speech segmentation process starts to identify candidate words, these words are stored in a <italic>proto-lexicon</italic>. A proto-lexicon is a mental lexicon consisting of sound sequences that are recognised and stored in long-term memory through language exposure, but it is not necessarily endowed with the semantic or morpho-syntactic properties associated with a fully-developed lexicon (<xref ref-type="bibr" rid="B26">Johnson, 2016</xref>; <xref ref-type="bibr" rid="B41">Ngon, Martin, Dupoux, Cabrol, Dutat, &amp; Peperkamp, 2013</xref>, <xref ref-type="bibr" rid="B36">Martin, Peperkamp, &amp; Dupoux, 2013</xref>; <xref ref-type="bibr" rid="B23">Hall&#233; &amp; de Boysson-Bardies, 1996</xref>). Evidence for the existence and nature of the infant proto-lexicon emerges from experimental research into the lexical knowledge of infants (see <xref ref-type="bibr" rid="B28">Junge, 2017</xref>; <xref ref-type="bibr" rid="B60">Swingley, 2009</xref>; <xref ref-type="bibr" rid="B26">Johnson, 2016</xref> for reviews of this topic). Research has found that infants of six to nine months are able to recognise the meanings of certain high frequency words (<xref ref-type="bibr" rid="B5">Bergelson &amp; Swingley, 2012</xref>, see <xref ref-type="bibr" rid="B44">Parise &amp; Csibra, 2012</xref>, for similar results), but infants of around 11 months recognise word forms where they do not necessarily know the meaning of the word (<xref ref-type="bibr" rid="B62">Vihman, Nakai, DePaolis, &amp; Hall&#233;, 2004</xref>; see also <xref ref-type="bibr" rid="B60">Swingley, 2009, pp. 3619&#8211;3620, 3625&#8211;3626</xref>; <xref ref-type="bibr" rid="B29">Jusczyk, 2000</xref>; <xref ref-type="bibr" rid="B59">Swingley, 2005</xref>).</p>
<p>Importantly, the proto-lexicon is receptive, in that it is formed through language exposure, and is utilised in speech recognition but not yet in productive language use. Research emphasises the probabilistic nature of the proto-lexicon (<xref ref-type="bibr" rid="B41">Ngon et al., 2013</xref>): It is formed when infants parse speech streams into identifiable units, using the cues described in section 1.2.1. These units will often be words, but can also include highly frequent sound sequences that are not actually words in the language being learned (<xref ref-type="bibr" rid="B41">Ngon et al., 2013</xref>).</p>
<p>While most literature on the proto-lexicon focuses on infant language acquisition, there is also some evidence for proto-lexical knowledge in adults. For example in one language learning task, adults were exposed to an artificial language for 10 hours. An experiment conducted three years later revealed that participants could still recognise some high frequency words from that language (<xref ref-type="bibr" rid="B19">Frank, Tenenbaum, &amp; Gibson, 2013</xref>). There is also evidence from former speakers of a language, who have lost their overt language knowledge, where they still show some evidence of implict knowledge of word-forms, thereby showing properties of a proto-lexicon (<xref ref-type="bibr" rid="B16">de Bot &amp; Stoessel, 2000</xref>; <xref ref-type="bibr" rid="B61">van der Hoeven &amp; de Bot, 2012</xref>; <xref ref-type="bibr" rid="B24">Hansen, Umeda, &amp; McKinney, 2002</xref>). In fact, other research shows that non-speakers of Norwegian were able to distinguish words from non-words in Norwegian after a brief period of familiarisation (<xref ref-type="bibr" rid="B32">Kittleson, Aguilar, Tokerud, Plante, &amp; Asbj&#248;rnsen, 2010</xref>). It seems likely, then, that adults who are exposed to a language in a real-life situation will start to automatically segment forms, and store them in a proto-lexicon.</p>
</sec>
<sec>
<title>1.2.3 Phonotactic knowledge and wellformedness</title>
<p>Related to, but separate from, the literature on word-segmentation is a literature within phonology, examining the degree to which phonological knowledge is gradient. This literature developed in response to the idea that phonology is based exclusively on categorical patterns that distinguish legal words from illegal words in a language. It shows that speakers of a language have fine-grained intuitions about degrees of wellformedness, showing that phonological knowledge is gradient, and is related to statistical patterns in the the language.</p>
<p>The main task used in the literature on phonotactic knowledge is ratings of non-words, and the overwhelming finding is that fluent speakers of a language provide ratings that are very highly correlated with lexical statistics calculated over word-types (not tokens) (<xref ref-type="bibr" rid="B22">Frisch, Large, Zawaydeh, &amp; Pisoni, 2001</xref>; <xref ref-type="bibr" rid="B25">Hay, Pierrehumbert, &amp; Beckman, 2004</xref>; <xref ref-type="bibr" rid="B50">Richtsmeier, 2011</xref>). In fact, approaches to phonotactics generally model the phonotactics of a language based on word-types in the lexicon of that language (<xref ref-type="bibr" rid="B13">Coleman &amp; Pierrehumbert, 1997</xref>; <xref ref-type="bibr" rid="B64">Vitevitch &amp; Luce, 1999</xref>; <xref ref-type="bibr" rid="B21">Frisch, Large, &amp; Pisoni, 2000</xref>; <xref ref-type="bibr" rid="B3">Bailey &amp; Hahn, 2001</xref>; <xref ref-type="bibr" rid="B25">Hay et al., 2004</xref>; <xref ref-type="bibr" rid="B65">Vitevitch &amp; Luce, 2004</xref>). Thus, there is good evidence that once a lexicon is established, a speaker is able to generalise over the types in the lexicon in order to create strong intuitions about phonotactic wellformedness.</p>
<p>However, it is important to note that while word-segmentation strategies and phonotactic wellformedness intuitions are both probabilistic and related to each other, they draw on different types of specific knowledge. Information about phonotactic wellformedness cannot be straightforwardly extracted from a running speech stream if it relies on types in a lexicon. The statistics used for segmentation, by contrast, are tracked across tokens, and do not presuppose any knowledge of words (e.g., <xref ref-type="bibr" rid="B8">Cairns, Shillcock, Chater, &amp; Levy, 1997</xref>). These &#8216;bottom-up&#8217; approaches build phonotactic models directly from running speech, rather than word-types. It is important to note that while such approaches account for word segmentation, they do not account for wellformedness judgments as effectively as models built over a lexicon. Indeed, Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>) compared this &#8216;bottom-up&#8217; approach to a lexicon-based approach to modelling M&#257;ori phonotactics. They found that a phonotactic model based on lexical forms accounted for phonotactic judgments of M&#257;ori and non-M&#257;ori speakers better than a &#8216;bottom-up&#8217; approach.</p>
<p>There is less work looking at phonotactic knowledge in children, although work from different tasks certainly suggests that word-based phonotactic knowledge is available and used. For example Edwards, Beckman, and Munson (<xref ref-type="bibr" rid="B18">2004</xref>) report that three to nine year old children produce non-words faster and more accurately when they are more phonotactically well-formed, noting that knowledge of word phonotactics is a key part of acquisition: &#8220;An increase in vocabulary size does not simply mean that the child knows more words, but also that the child is able to make more and more robust phonological generalizations&#8221; (p. 434). Similarly, Storkel and Rogers (<xref ref-type="bibr" rid="B58">2000</xref>) report that 10- and 13-year-old children can more easily learn high probability non-words than low probability non-words.</p>
<p>Even infants have been shown to be sensitive to phonotactics in ways that are suggestive of early stages of lexical acquisition. For example, infants show sensitivity to legal vs illegal words in their language from a young age (<xref ref-type="bibr" rid="B57">Steber &amp; Rossi, 2020</xref>). Furthermore, there is evidence that sensitivity to gradient patterns exists as early as nine months: Jusczyk, Luce, and Charles-Luce (<xref ref-type="bibr" rid="B30">1994</xref>) found that nine month olds listened longer to high probability non-words than to ones containing low-probability phonotactic sequences.</p>
<p>Thus, while work on children has not specifically investigated the size and nature of the lexicon underpinning their phonotactic knowledge, the fact that it is evidenced so young would support the conjecture that explicit word knowledge is not a requirement. The word forms in the proto-lexicon seem to be sufficient to develop sophisticated probabilistic phonotactic knowledge.</p>
<p>In all, work on acquisition supports a trajectory in which learners use probabilistic information to segment words from the word stream, form a proto-lexicon, and then generate phonotactic knowledge from the lexicon. Of course this is not strictly sequential, and any knowledge in one of these domains can accelerate learning in the other.</p>
<p>Another relevant area of research is research into the phonotactic knowledge of bilinguals. Messer, Leseman, Boom, and Mayo (<xref ref-type="bibr" rid="B38">2010</xref>) compared the phonotactic knowledge of Turkish-dominant Turkish-Dutch bilingual preschoolers with Dutch monolingual preschoolers by conducting a wellformedness rating task with nonce words. They found that the bilingual participants had a stronger effect of phonotactics in Turkish in comparison with Dutch, while the monolingual participants outperformed the bilingual participants in Dutch. Similar results were found in research on infant and adult bilinguals in Catalan and Spanish, in which their response to phonotactically legal and illegal nonce words was affected by their dominant language (<xref ref-type="bibr" rid="B53">Sebasti&#225;n-Gall&#233;s &amp; Bosch, 2002</xref>). Indeed, research finds that even highly proficient L2 speakers show transfer effects from L1, both in related language like English and German (<xref ref-type="bibr" rid="B67">Weber &amp; Cutler, 2006</xref>), and very different languages like English and Cantonese (<xref ref-type="bibr" rid="B68">Yip, 2020</xref>). The effect of this is that listeners unconsciously use certain phonotactic cues from their L1 in order to segment speech in L2. Importantly, even if not to the same degree as monolinguals, the research generally supports phonotactic sensitivity in bilinguals to both their dominant and their non-dominant language (<xref ref-type="bibr" rid="B20">Frisch &amp; Brea-Spahn, 2010</xref>).</p>
<p>Our work follows on from a previous study which considers the type of learning that may occur in adults who are exposed to a language regularly, but who are not overtly trying to learn that language. Do they learn to spot word boundaries and form a proto-lexicon? And what level of phonotactic knowledge might they generate?</p>
</sec>
</sec>
<sec>
<title>1.3 Previous work</title>
<p>Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>) reported that non-Maori speaking adults living in New Zealand have extensive phonotactic knowledge of M&#257;ori, which stems from a surprisingly large proto-lexicon obtained through their environmental exposure to the language. This claim was on the basis of two online experiments:</p>
<list list-type="order">
<list-item><p><bold><italic>Word Identification Task:</italic></bold> Participants were presented with 150 pairs of words and non-words. They were asked to rate on a Likert scale how confident they were that each stimulus is a real word.</p></list-item>
<list-item><p><bold><italic>Wellformedness Rating Task:</italic></bold> Participants were presented with 240&#8211;320 M&#257;ori-like non-words. They were asked to rate on a Likert scale how M&#257;ori-like each non-word was.</p></list-item>
</list>
<p>Along with non-M&#257;ori speaking New Zealanders, fluent M&#257;ori speakers and Americans with no knowledge of M&#257;ori also participated in the experiments.</p>
<p>The study reported the following findings:</p>
<list list-type="order">
<list-item><p>In the Word Identification Task, across five frequency categories (from high frequency words to low frequency words), NMSs rated words more highly than non-words.</p></list-item>
<list-item><p>NMSs&#8217; ratings of stimuli in both the Word Identification Task and the Wellformedness Rating Task were positively correlated with the phonotactic wellformedness of the stimulus.</p></list-item>
<list-item><p>In the Wellformedness Rating Task, NMSs significantly outperformed Non-M&#257;ori speakers in the United States, and their results were nearly identical to the results of the M&#257;ori speaking participants.</p></list-item>
<list-item><p>The phonotactic knowledge of the NMSs is best modelled under the assumption that lexical items are decomposed into parts that occur with statistical regularity in the language, called <italic>morphs</italic>. Moreover, NMSs&#8217; knowledge can be best modelled by a proto-lexicon consisting of approximately 1500 relatively common morphs.</p></list-item>
</list>
<p>Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>) conclude that there is good evidence that non-M&#257;ori speakers in New Zealand have a proto-lexicon of approximately 1500 morphs. This proto-lexicon is responsible for their ability to distinguish real words from non-words (under the assumption the real words are in the proto-lexicon), and also for their fine-grained phonotactic knowledge (which arises as a statistical generalization over the forms in the proto-lexicon).</p>
<p>The two experiments in Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>) were conducted on separate groups of participants. This means that the results of these experiments could not be explicitly connected. While Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>) hypothesise that the phonotactic knowledge stems from proto-lexical knowledge, they do not show a relationship between the two tasks. If there was a causal relationship, then participants with a greater ability to distinguish real words from non-words would also show increased phonotactic knowledge. The relationship between lexical and phonotactic knowledge is crucial to understanding the source of these effects, and consequently the inability to connect these experiments is a key limitation in the previous study. This paper conducts these two tasks on the same set of participants, allowing us to test whether there is a statistical link between responses in the two tasks.</p>
<p>In rerunning the experiments on a new set of participants, we also use an improved set of stimuli, and thus attempt to replicate the key findings of Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>) with this new set of stimuli that address some potential limitations:</p>
<disp-quote>
<p><italic>Construction of Non-Words</italic>: In the Word Identification Task reported in Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>) there was an attempt to match real words with non-words of the same length and approximate phonotactic score. However the degree to which the non-words really &#8216;match&#8217; is entirely dependent on the phonotactic model. If there is any way in which the real words actually systematically differ from the non-words, then participants could use this difference to perform the task, rather than word knowledge itself. In our experiment we use word pairs where each non-word is manually created to be maximally matched to its real-word pair. This enables us to confirm that the result replicates with an entirely different means of non-word construction.</p>
<p><italic>Morphologically Complex Stimuli</italic>: The original stimuli in the Wellformedness Rating Task varied in length, and were manipulated to reflect different phonotactic probabilities. Potential morphological parses were not considered, but the analysis of Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>) suggests not only that the proto-lexicon contain morphs (rather than unanalysed morphologically-complex words), but also that the participants applied statistical knowledge to analyse some of the stimuli as morphologically complex. Our study uses shorter stimuli, which are carefully examined from a variety of perspectives, to minimise the likelihood of perceived morphological complexity (see the Supplementary Materials).</p>
<p><italic>Orthographic Cues</italic>: The results of Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>) showed a small effect by the presence of a macron in the stimuli. Non-M&#257;ori speakers showed a tendency to rate stimuli containing macrons as more M&#257;ori-like. Macrons are a very salient feature of M&#257;ori orthography for NMSs, and consequently the presence of a macron impacted the results of the experiments. Our experiment does not include any stimuli with macrons.</p>
</disp-quote>
<p>Our experiments thus use a different and much cleaner set of stimuli to replicate the key results of Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>), and to seek a link between their two tasks. This would add considerable weight to the claim that a M&#257;ori proto-lexicon leads to phonotactic knowledge.</p>
</sec>
</sec>
<sec sec-type="methods">
<title>2 Methods</title>
<sec>
<title>2.1 Experimental tasks</title>
<p>We replicated the online <italic>Word Identification</italic> and <italic>Wellformedness Rating</italic> tasks that were conducted by Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>), to verify the existence of the proto-lexicon and its relationship to phonological knowledge with updated stimulus sets. All experimental protocols in the online experiment were approved by the Human Research Ethics Committee at the University of Canterbury. This section contains a description of the generation of the experimental stimuli, factors used in the statistical analysis of the experimental results, experimental procedure, participant details, and the procedure of the statistical analyses. A more comprehensive description of the methodology of this experiment is in the Supplementary Materials.</p>
</sec>
<sec>
<title>2.2 Stimuli and materials</title>
<p>The full set of stimulus materials consists of 1042 items: 521 M&#257;ori word and M&#257;ori-like non-word pairs in total, which span 21 different phonotactic shapes. All stimuli were two to three syllables long, and none of them contained long vowels. This set provided the stimuli for both the Word Identification Task, and the Wellformedness Rating Task.</p>
<p>Unlike in Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>), non-words were manually obtained through minimal manipulation of real-word counterparts. We began by identifying candidates from all two to three syllable words consisting of at least three phonemes that are attested in the dictionary. Words that had an English lookalike, and words that potentially involved reduplication were removed from consideration. We then altered these words by manipulating up to three phonemes. A successful paired non-word meets the following criteria: (i) The word and non-word have the same number of syllables; (ii) The word and non-word have the same number of phonemes; (iii) The word and non-word have consonants and vowels in the same positions; (iv) The word and non-word have extremely similar phonotactic scores, as defined below. The number of phonemes ranges from three to six phonemes. Steps were taken to ensure that neither the real words or the non-words would be recognised as morphologically complex.</p>
<p>Each word in the stimuli was grouped into one of three <italic>frequency</italic> bins. A frequency score was generated for each word stimulus based on its raw count from the concatenation of the MAONZE corpus (<xref ref-type="bibr" rid="B31">King, Maclagan, Harlow, Keegan, &amp; Watson, 2010</xref>) and the M&#257;ori Broadcast Corpus (MBC) (<xref ref-type="bibr" rid="B7">Boyce, 2006</xref>), totalling approximately 1.35M words of transcribed speech. For the purpose of this experiment, a non-word was grouped into the same frequency bin as its paired word. In all, there were 64 pairs in the high-frequency category (100+ per million; 135+ times in corpus), 101 pairs in the mid-frequency category (6-99 per million; 7-134 times in corpus), and 107 in the low-frequency category (1&#8211;5 per million; 1&#8211;6 times in corpus), for a total of 272 words. The stimuli included an additional 249 pairs that were unattested in the corpus.</p>
<p>A <italic>phonotactic score</italic> was calculated using length-normalised log-probabilities from trigram-based n-gram language models trained on morph types obtained from words in the Te Aka Dictionary using the SRI Language Modeling Toolkit (SRILM). We obtained a list of all headwords in Te Aka Dictionary (<xref ref-type="bibr" rid="B39">Moorfield, n.d.</xref>) as well as their inflected (i.e. passive voice) forms, excluding affixes and proper nouns. For the purposes of constructing the phonotactic model, we ignored all vowel length distinctions, because this sort of model showed the highest correlation with human ratings in the experiments reported by Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>).</p>
<p>Each stimulus in the dataset was assigned the number of phonological neighbours it has, in order to estimate the <italic>neighbourhood density</italic> of each stimulus. For the purpose of this study, we estimated this through a &#8220;neighbourhood occupancy rate&#8221; measure. This corresponds to a measure of the number of variations on a stimulus with an edit distance of one that correspond to a real M&#257;ori word.</p>
<p>For the Word Identification Task, the set of stimuli consisted of 272 word-nonword pairs for a total of 544 stimuli. &#8216;Unattested&#8217; items (249 pairs) were not included in this task. The frequency values &#8216;high&#8217; (<italic>n<sub>pairs</sub></italic> = 64), &#8216;mid&#8217; (<italic>n<sub>pairs</sub></italic> = 101), and &#8216;low&#8217; (<italic>n<sub>pairs</sub></italic> = 107) were used as bins. From this set of stimuli, for each participant, 20 pairs were randomly sampled without replacement from each bin, for a total of 120 stimuli. Each participant in the experiment received a different random sample. For each participant, no stimuli were shared between the Wellformedness Rating Task and the Word Identification Task.</p>
<p>For the Wellformedness Rating Task, the set of stimuli consisted of 521 M&#257;ori-like non-words. This dataset consisted of all non-words in the dataset. This includes non-words paired with unattested words (and thus excluded from the Word Identification Task). These non-words were grouped into three roughly equally-sized bins based on their phonotactic score: High (<italic>n</italic> = 175), medium (<italic>n</italic> = 174), low (<italic>n</italic> = 172). From this set, for each participant, 40 stimuli were sampled without replacement from each bin, for a total of 120 stimuli. Each participant in the experiment received a different random sample.</p>
<p>In order to investigate the correlation between performance in the Word Identification Task and Wellformedness Task, performance in the Word Identification Task was scored for each participant. We used this overall measure of participant accuracy as a predictor for our analysis of the Wellformedness Task. We estimate Word Identification Task accuracy for a participant by subtracting their mean confidence rating for non-words from their mean confidence rating for words:</p>
<disp-quote>
<p><italic>Word Identification Task Score</italic> = <italic>mean</italic>(<italic>word rating</italic>) <italic>&#8211; mean</italic>(<italic>nonword rating</italic>)</p>
</disp-quote>
<p>Therefore, higher positive Word Identification Task scores indicate that the participant is more likely to reliably separate words from non-words, than participants who have lower scores.</p>
<sec>
<title>2.2.1 Post-experiment questionnaire</title>
<p>A post-experiment questionnaire containing 26 questions was used in the experiment, which was almost identical to that of Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>). The questions were mostly about participants&#8217; sociolinguistic background information. In addition, the participants were asked about their levels of M&#257;ori proficiency and exposure as well as their knowledge of M&#257;ori. Because we wanted participants who had lived most of their life in New Zealand we added a question about whether they had lived overseas for more than one year and another question about whether they had studied linguistics at university. A couple of questions about participants&#8217; social attitude towards the M&#257;ori language were also added. The questions used in this questionnaire are in the Supplementary Materials.</p>
</sec>
</sec>
<sec>
<title>2.3 Procedure</title>
<p>The experiment was conducted online using a custom in-browser interface (<xref ref-type="bibr" rid="B10">Chan, 2018</xref>).</p>
<p>The experiment consisted of the Word Identification Task and Wellformedness Rating Task. Each task contained 120 trials, for a total of 240 trials. The order of these tasks was counterbalanced. Separate task instructions were presented prior to each task. The entire procedure took less than 30 minutes.</p>
<p>The procedure for each trial was as follows:</p>
<disp-quote>
<p><italic>Word Identification Task</italic>: Participants were told to judge stimuli without looking up the word in a dictionary, or asking anyone else for help. A stimulus was presented orthographically in the middle of the screen, which was either a real word or a M&#257;ori-like non-word. Participants were asked to rate how confident they are that each item they saw was an actual M&#257;ori word, using a scale ranging from 1 (&#8216;Confident that it is NOT a M&#257;ori word&#8217;) to 5 (&#8216;Confident that it IS a M&#257;ori word&#8217;). After the participant clicked one of the options and &#8216;Next,&#8217; the next stimulus was presented.</p>
<p><italic>Wellformedness Rating Task</italic>: Participants were instructed that they would see a made-up word in the middle of the screen. Participants were asked to rate how good this word would be as a M&#257;ori word, using a scale ranging from 1 (&#8216;Non M&#257;ori-like non-word&#8217;) to 5 (&#8216;Highly M&#257;ori-like non-word&#8217;). Rating examples were given with a highly M&#257;ori-like non-word and non M&#257;ori-like non-word which were not actual stimulus items. After the participant clicked one of the options and &#8216;Next,&#8217; the next stimulus was presented.</p>
</disp-quote>
</sec>
<sec>
<title>2.4 Participants</title>
<p>A total of 221 online participants were recruited using paid Facebook advertisements. Participants could choose to be paid a $10 online gift voucher. We set several criteria for participants results to be used for data analysis. These are summarised as follows:</p>
<list list-type="bullet">
<list-item><p>Be a native speaker of New Zealand English and 18 years or older.</p></list-item>
<list-item><p>Not have lived outside New Zealand, for any period of longer than a year, since they were aged seven.</p></list-item>
<list-item><p>Never have studied linguistics at a university.</p></list-item>
<list-item><p>Not be able to hold a basic conversion in M&#257;ori.</p></list-item>
</list>
<p>Following screening on the above criteria, and further outlier removal (see the Supplementary Materials for details), there were 187 participants in total. Four participants were removed from analysis for the Word Identification Task and seven participants were removed from analysis of the Wellformedness Rating Task, due to low variation in the distribution of their responses.</p>
<p>A majority of participants were female (145, 77.5%) and lived on the North Island of New Zealand (131, 70.1%). They generally rated themselves very poorly for ability to speak or understand M&#257;ori. The survey included two questions that asked participants to rate themselves on a zero to five scale for their ability to speak M&#257;ori, and a zero to five scale for their ability to understand M&#257;ori, with zero indicating &#8216;not at all,&#8217; and five indicating &#8216;very well.&#8217; These measures were added up into a 10-point score to represent the M&#257;ori language ability of each participant. 149 (81.4%) participants scored at most two on this 10-point scale, and due to data filtering the maximum score was four, which 14 (7.7%) received. Although participants generally rated themselves as unable to speak or understand M&#257;ori, they indicated that they are exposed to M&#257;ori regularly. The survey also included two questions that asked participants to rate themselves on a zero to five scale for how often they were exposed to M&#257;ori through the media in everyday life, and how often they were exposed to it through social interactions. In this case, zero represented &#8216;less than once a year&#8217; and five represented &#8216;multiple times a day.&#8217; The two measures were again added up into a single 10-point score, this time representing the degree of a participant&#8217;s exposure to M&#257;ori in everyday life. 155 (82.9%) participants scored six or greater, representing exposure to M&#257;ori of some form on at least a monthly or weekly basis; thus, most participants generally received consistent exposure to M&#257;ori.</p>
</sec>
<sec>
<title>2.5 Statistical analyses</title>
<p>Consistent with Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>), we analysed the data with (logit) ordinal regression using the ordinal package (<xref ref-type="bibr" rid="B11">Christensen, 2019</xref>) in R (<xref ref-type="bibr" rid="B49">R Core Team, 2021</xref>). Using stepwise regression, models are compared to each other to see which fits the best, based on the Akaike Information Criterion (AIC). Random effects structure was kept as maximal as possible while maintaining the ability to successfully create the model (<xref ref-type="bibr" rid="B4">Barr, Levy, Scheepers, &amp; Tily, 2013</xref>). For the Word Identification Task, the dependent variable was confidence rating on a Likert scale. For the Wellformedness Rating Task, the dependent variable was wellformedness rating on a Likert scale. The random effects were grouped by participant (corresponding to participant ID) and word.</p>
<p>As part of the stepwise regression approach, multiple effects were modelled initially, and non-significant effects were removed. For the Word Identification Task, the fixed effects considered were <italic>phonotactic score, neighbourhood density, number of phonemes, stimulus type</italic> (words, non-words) and <italic>frequency bin</italic> (high, mid, low). Phonotactic score and stimulus type were our test predictors, while phonological neighbourhood density, number of phonemes, and frequency bin were control predictors. Phonotactic score, neighbourhood density, and number of phonemes were centred. We started with a model containing five-way interactions. We then pruned the model by removing non-significant factors, the full details of which are in the Supplementary Materials.</p>
<p>For the Wellformedness Rating Task, the fixed effects considered were <italic>phonotactic score, neighbour density</italic>, and <italic>Word Identification Task score</italic>. All factors were centered. A similar stepwise regression approach was used for modelling this task as with the Word Identification Task.</p>
</sec>
</sec>
<sec>
<title>3 Results</title>
<sec>
<title>3.1 Word identification task</title>
<p>In order to factor the contrasting distribution of high frequency stimuli vs. mid &amp; low frequency stimuli into the mixed effects analysis, we used Helmert contrasts to compare low vs. mid bins, and low &amp; mid vs. high bins. In the model, the dependent variable was the confidence rating. Phonotactic score, neighbourhood density, number of phonemes, stimulus type and frequency bin were used as fixed effects. Phonotactic score, stimulus type, and frequency bin were entered into the model as a three-way interaction. Participant and stimulus were random effects, with the fixed effects as a random slope for participant. A summary of fixed effects in this model is in <xref ref-type="table" rid="T1">Table 1</xref>; for full details including threshold values, see the Supplementary Materials.</p>
<table-wrap id="T1">
<label>Table 1</label>
<caption>
<p>Fixed effects of the ordinal regression model results for Word Identification Task. Asterisked <italic>p</italic> values are statistically significant. The estimates are rounded to two significant figures.</p>
</caption>
<table>
<thead>
<tr>
<td align="left" valign="top"><bold>Effects</bold></td>
<td align="left" valign="top"><bold>Estimate</bold></td>
<td align="left" valign="top"><bold>Standard Error</bold></td>
<td align="left" valign="top"><bold><italic>z</italic> Value</bold></td>
<td align="left" valign="top"><bold><italic>p</italic> Value</bold></td>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Phonotactic Score (norm.)</td>
<td align="left" valign="top">3.23</td>
<td align="left" valign="top">0.88</td>
<td align="left" valign="top">3.69</td>
<td align="left" valign="top">&lt;0.001*</td>
</tr>
<tr>
<td align="left" valign="top">Stimulus Type &#8211; Word</td>
<td align="left" valign="top">1.0</td>
<td align="left" valign="top">0.10</td>
<td align="left" valign="top">9.60</td>
<td align="left" valign="top">&lt;0.001*</td>
</tr>
<tr>
<td align="left" valign="top">Frequency Bin (low vs. mid)</td>
<td align="left" valign="top">0.06</td>
<td align="left" valign="top">0.08</td>
<td align="left" valign="top">0.82</td>
<td align="left" valign="top">0.41</td>
</tr>
<tr>
<td align="left" valign="top">Frequency Bin (low/mid vs. high)</td>
<td align="left" valign="top">0.06</td>
<td align="left" valign="top">0.05</td>
<td align="left" valign="top">1.09</td>
<td align="left" valign="top">0.28</td>
</tr>
<tr>
<td align="left" valign="top">Number of Phonemes (norm.)</td>
<td align="left" valign="top">0.44</td>
<td align="left" valign="top">0.12</td>
<td align="left" valign="top">3.63</td>
<td align="left" valign="top">&lt;0.001*</td>
</tr>
<tr>
<td align="left" valign="top">Neighbourhood Density (norm.)</td>
<td align="left" valign="top">5.82</td>
<td align="left" valign="top">1.47</td>
<td align="left" valign="top">3.96</td>
<td align="left" valign="top">&lt;0.001*</td>
</tr>
<tr>
<td align="left" valign="top">Ph. Score (norm.) * Stim. Type &#8211; Wd.</td>
<td align="left" valign="top">&#8211;1.38</td>
<td align="left" valign="top">0.95</td>
<td align="left" valign="top">&#8211;1.46</td>
<td align="left" valign="top">0.14</td>
</tr>
<tr>
<td align="left" valign="top">Ph. Score (norm.) * Fq. Bin (low vs. mid)</td>
<td align="left" valign="top">0.07</td>
<td align="left" valign="top">0.75</td>
<td align="left" valign="top">0.08</td>
<td align="left" valign="top">0.93</td>
</tr>
<tr>
<td align="left" valign="top">Ph. Score (norm.) * Fq. Bin (low/mid vs. high)</td>
<td align="left" valign="top">0.35</td>
<td align="left" valign="top">0.51</td>
<td align="left" valign="top">0.70</td>
<td align="left" valign="top">0.49</td>
</tr>
<tr>
<td align="left" valign="top">Stim. Type &#8211; Wd. * Fq. Bin (low vs. mid)</td>
<td align="left" valign="top">0.18</td>
<td align="left" valign="top">0.11</td>
<td align="left" valign="top">1.63</td>
<td align="left" valign="top">0.1</td>
</tr>
<tr>
<td align="left" valign="top">Stim. Type &#8211; Wd. * Fq. Bin (low/mid vs. high)</td>
<td align="left" valign="top">0.52</td>
<td align="left" valign="top">0.08</td>
<td align="left" valign="top">6.88</td>
<td align="left" valign="top">&lt;0.001*</td>
</tr>
<tr>
<td align="left" valign="top">Ph. Score (n.) * Stim. Type &#8211; Wd. * Fq. Bin (l. vs. m.)</td>
<td align="left" valign="top">&#8211;0.21</td>
<td align="left" valign="top">1.05</td>
<td align="left" valign="top">&#8211;0.2</td>
<td align="left" valign="top">0.84</td>
</tr>
<tr>
<td align="left" valign="top">Ph. Score (n.) * Stim. Type &#8211; Wd. * Fq. Bin (l./m. vs. h.)</td>
<td align="left" valign="top">&#8211;1.56</td>
<td align="left" valign="top">0.73</td>
<td align="left" valign="top">&#8211;2.14</td>
<td align="left" valign="top">0.04*</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>This model shows a significant effect of phonotactic score, stimulus type, number of phonemes, and neighbourhood density. It also shows a significant effect in the three-way interaction between phonotactic score, stimulus type, and frequency bin. <xref ref-type="fig" rid="F1">Figure 1</xref> shows this three-way interaction. First, this plot shows that for all stimulus type and bin categories, apart from high frequency words, there is a positive relationship between predicted rating and phonotactic score. Second, in all frequency bins there is a higher predicted rating for words than non-words, with a positive correlation between the frequency bin and the difference in rating between words and non-words. Third, high frequency words appear to show a negative, rather than positive, relationship with phonotactic score. In relation to this latter point, the wide confidence band of the high frequency word category indicates that the negative slope is not likely to be significant. Instead, it indicates that phonotactic score is not a significant predictor for the confidence rating of high frequency words: participants were able to generally recognise the high frequency words used in this experiment, regardless of their wellformedness. This interaction is highly consistent with the idea of a proto-lexicon and can be interpreted as follows. Since participants are exposed to high frequency words so often, they are extremely likely to have entries for such words in their proto-lexicons and may even be consciously aware of them, yielding high confidence scores regardless of phonotactics. In other categories (i.e. for nonwords, and for lower frequency words), participants may or may not have an entry for a given stimulus in their proto-lexicon, or may be unaware of it, leading them to use the wellformedness of the stimulus as a proxy for whether it is likely to be a word or not. In summary, as with Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>), we find that participants generally responded to the phonotactic probability of the stimuli, even though they were not explicitly tasked to in this task. In addition to showing sensitivity to phonotactics, they are also able to distinguish words from non-words. However, in Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>), the relationship between frequency and confidence rating was mixed. In these results, there is a much clearer effect.</p>
<fig id="F1">
<label>Figure 1</label>
<caption>
<p>The relationship between predicted confidence rating and phonotactic score, stimulus type, and frequency bin in the Word Identification Task model.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-7943-g1.png"/>
</fig>
<p>The model also shows a significant positive effect of number of phonemes and neighbourhood density. Neighbourhood density is likely a predictor for a higher confidence rating because stimuli with a denser phonological neighbourhood are more likely to resemble words known to the participants, and therefore be more easily confused with them. The positive effect of the number of phonemes is more difficult to interpret. We speculate that it relates to a possible effect of complexity in relation to length: stimuli that have more phonemes are more complex, and so may appear to be more word-like to participants if the participant is attempting to guess if a stimulus is a real word or not. This is only conjecture, however; we did not have a pre-existing hypothesis relating to this effect of number of phonemes. Consequently, we do not place much weight on this result.</p>
</sec>
<sec>
<title>3.2 Wellformedness Rating Task</title>
<p>The best model for the Wellformedness Rating Task is presented in <xref ref-type="table" rid="T2">Table 2</xref>. In the model, the dependent variable was the wellformedness rating. Phonotactic score, Word Identification Task score, and neighborhood density were used as fixed effects. Phonotactic score was entered into the model with a two-way interaction with Word Identification Task score and neighbourhood density. Participant and stimulus were random effects, with phonotactic score + neighbourhood density, and Word Identification Task score as random slopes for these random effects respectively.</p>
<table-wrap id="T2">
<label>Table 2</label>
<caption>
<p>Fixed effects of the ordinal regression model results for Wellformedness Rating Task. Asterisked <italic>p</italic> values are statistically significant. All values are normalised. The estimates are rounded to two significant figures.</p>
</caption>
<table>
<thead>
<tr>
<td align="left" valign="top"><bold>Effects</bold></td>
<td align="left" valign="top"><bold>Estimate</bold></td>
<td align="left" valign="top"><bold>Standard Error</bold></td>
<td align="left" valign="top"><bold><italic>z</italic> Value</bold></td>
<td align="left" valign="top"><bold><italic>p</italic> Value</bold></td>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Phonotactic Score (norm.)</td>
<td align="left" valign="top">6.84</td>
<td align="left" valign="top">0.467</td>
<td align="left" valign="top">14.6</td>
<td align="left" valign="top">&lt;0.001*</td>
</tr>
<tr>
<td align="left" valign="top">Word Identification Task Score</td>
<td align="left" valign="top">&#8211;0.07</td>
<td align="left" valign="top">0.259</td>
<td align="left" valign="top">&#8211;0.3</td>
<td align="left" valign="top">0.793</td>
</tr>
<tr>
<td align="left" valign="top">Neighbourhood Density (norm.)</td>
<td align="left" valign="top">&#8211;1.24</td>
<td align="left" valign="top">0.830</td>
<td align="left" valign="top">&#8211;1.5</td>
<td align="left" valign="top">0.136</td>
</tr>
<tr>
<td align="left" valign="top">Phon. Score (norm.): Word Ident. Task Score</td>
<td align="left" valign="top">2.22</td>
<td align="left" valign="top">0.828</td>
<td align="left" valign="top">2.7</td>
<td align="left" valign="top">&lt;0.01*</td>
</tr>
<tr>
<td align="left" valign="top">Phon. Score (norm.): Neigh. Density (norm.)</td>
<td align="left" valign="top">&#8211;17.77</td>
<td align="left" valign="top">7.666</td>
<td align="left" valign="top">&#8211;2.3</td>
<td align="left" valign="top">0.02*</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The model shows a strong effect of phonotactic score, and both interactions are significant: Phonotactic score and Word Identification Task Score, and phonotactic score and the neighbourhood density. <xref ref-type="fig" rid="F2">Figure 2</xref> shows the mean rating for each stimulus in the Wellformedness Rating Task plotted against their phonotactic scores. The data shows that phonotactic score of stimuli corresponds positively with the wellformedness rating. The slope of the mean ratings is similar to the NMSs&#8217; ratings in Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>), and is steeper than the slope of the American ratings in that experiment. Consequently these results replicate Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>).</p>
<fig id="F2">
<label>Figure 2</label>
<caption>
<p>The relationship between phonotactic score and mean wellformedness rating of stimuli in the Wellformedness Rating Task in the raw experimental results. The dotted vertical bars indicate boundaries between phonotactic score bins (low-medium-high). The shaded area indicates the 95% confidence interval of the fitted line.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-7943-g2.png"/>
</fig>
<p>The phonotactic score of the stimuli interacts with two other factors in this model. Crucially, given our main research question, is that participant score in the Word Identification Task interacts with phonotactic score in predicting the wellformedness ratings. <xref ref-type="fig" rid="F3">Figure 3</xref> shows this interaction from the model.</p>
<fig id="F3">
<label>Figure 3</label>
<caption>
<p>Performance of participants in the Wellformedness Rating Task based on their accuracy in the Word Identification Task (Word Identification Task score). The red line is ratings by participants with a Word Identification Task score of one. The teal line is ratings by participants with a Word Identification Task score of zero. The shaded areas indicate the 95% confidence interval for the predicted mean ratings.</p>
</caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="labphon-14-7943-g3.png"/>
</fig>
<p><xref ref-type="fig" rid="F3">Figure 3</xref> shows the relationship between phonotactic score and mean rating for participants who performed well in the Word Identification Task (red) and those that performed less well (teal). The red line has a steeper slope, showing greater sensitivity to phonotactics. We can see that the better performing participants from the Wellformedness Rating Task (in red) generally gave more low ratings to forms with low phonotactic scores, and more high ratings to forms with high phonotactic scores. This significant interaction, then, shows the hypothesised link between the two tasks.</p>
<p>Finally, the second interaction involves the neighbourhood density of the stimulus. Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>) did not test for such an effect, but this was more important for our study, as our non-words were shorter, more closely modelled on real words, and thus more likely to have neighbours. The significant interaction indicates that non-words with a greater neighbourhood density do not benefit from higher phonotactic scores as strongly as non-words with a less dense neighbourhood. At first glance this appears somewhat counter intuitive, as two factors that should make a non-word seem more word-like combine to predict lower wellformedness ratings. We speculate that this relates to the nature of the task. Participants are told that the words are not words of M&#257;ori, and are asked to rate &#8220;how good the word would be as a M&#257;ori word.&#8221; If it actually activates a real word, this may make it a poor candidate as a real word, as this could lead to potential lexical confusion. Thus, if the question was reworded to ask how much the word resembles existing words, this interaction may well disappear or reverse. This conjecture would need testing, and as this was not a predicted effect, we do not place too much weight on the result at this point.</p>
<p>The key results of the experiment are that: (a) Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>) is replicated with a new set of stimuli &#8211; non-M&#257;ori speakers remain highly sensitive to M&#257;ori phonotactics in this task, and (b) our research question is answered positively. Participants with a larger apparent proto-lexicon (as assessed through the Word Identification Task) are more sensitive to phonotactics (as assessed through the Wellformedness Rating Task).</p>
</sec>
</sec>
<sec>
<title>4 Discussion</title>
<p>Using a completely different set of stimuli and participants, we replicate the key findings of Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>), namely that: (a) Non-M&#257;ori speakers in New Zealand can discriminate words from non-words; and (b) they are highly sensitive to M&#257;ori phonotactics. We provide an important demonstration that the key findings of Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>) replicate with a set of stimuli that is much more carefully controlled. In the previous study, for example, the morphological complexity of the stimuli was not controlled, because it only emerged as a relevant factor after data analysis began, but the conclusions make critical reference to morphological structure. Furthermore, the phonotactics of words and non-words were also not optimally matched in the previous work. A major contribution of this work is to demonstrate that, despite shortcomings in the stimuli of Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>), their major findings are robust.</p>
<p>In terms of the discrimination of words and non-words, the effect that we find is robust, but it strongly correlates with word frequency. In the high frequency bin, words are reliably discriminated by participants in a way that depends less on phonotactics. In the mid and low frequency bins, there is a stronger effect of phonotactics. This effect is likely to do with the reduced role of lexical knowledge at lower frequencies: Lower frequency words are less likely to be recognised by the participants, and consequently participants rely on probabilistic phonotactics to identify words at lower frequencies. Note that while this distribution differs from Oh et al.&#8217;s findings (<xref ref-type="bibr" rid="B42">2020, p. 3</xref>), it is consistent with their modelling procedures, which show that the phonotactic knowledge is best understood as stemming from a proto-lexicon of relatively frequent lexemes.</p>
<p>Crucially, our study conducted both tasks on the same set of participants. We were therefore able go significantly beyond Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>), by demonstrating a link between (proto)-lexical knowledge and phonotactic knowledge. Participants who can better distinguish real and non-words are also more sensitive to phonotactic probability in their ratings of a set of unrelated non-words.</p>
<p>We hypothesised that this would be so if the phonotactic knowledge was a direct consequence of the proto-lexicon, and the fact that we have found this correlation provides further support for that interpretation. Of course correlation is not complete proof of causation, and indeed, if the phonotactics was not causally linked to the proto-lexicon, it is not at all implausible to imagine that both were affected by a shared third factor, such as degree of exposure, or language learning aptitude. However the correlation between these tasks certainly adds to the weight of evidence that the phonotactic knowledge stems from lexical knowledge. Crucial here, too, is that fact that Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>) attempted multiple models of the phonotactic knowledge, including models that used sequences of phonemes as extracted from running speech, without knowledge of word boundaries. However the models that best explained the knowledge were over types in the lexicon &#8211; models that presuppose there is an inventory of words or morphs over which phonotactic generalizations can be made. We thus have multiple sources of converging evidence.</p>
<p>Our Word Identification Task shows that the ability of participants to identify monomorphemic words positively correlates with the frequency of the word. Oh et al&#8217;s (<xref ref-type="bibr" rid="B42">2020</xref>) modelling shows the phonotactic knowledge is best explained by assuming knowledge of approximately 1500 common morphs, and our results show that speakers that can identify more morphs have more sophisticated phonotactic knowledge. This interpretation also retains a key assumption of the general literature on wellformedness rating tasks, namely that the phonotactic knowledge is a generalization over types in the lexicon (<xref ref-type="bibr" rid="B22">Frisch et al., 2001</xref>; <xref ref-type="bibr" rid="B25">Hay et al., 2004</xref>). Consequently, with evidence in the literature relating to the ability of individuals to segment words from the speech stream (<xref ref-type="bibr" rid="B46">Pe&#241;a et al., 2002</xref>; <xref ref-type="bibr" rid="B43">Onnis et al., 2005</xref>; <xref ref-type="bibr" rid="B40">Newport &amp; Aslin, 2000</xref>), the evidence points to a set of forms that have been segmented and stored in a proto-lexicon (<xref ref-type="bibr" rid="B19">Frank et al., 2013</xref>).</p>
<p>Our experiments support the interpretation that phonotactic knowledge appears to be derived from a proto-lexicon, but it is important to note that we do not make any further claims about the particular representation of that knowledge. It would be possible for such generalizations to be built on the fly, for example, via activation of words containing relevant sequences. It would also be possible for phonotactic knowledge to derive from generalizations in the lexicon, but for the lexical knowledge and phonotactic knowledge to be represented quite separately in the mind (<xref ref-type="bibr" rid="B63">Vitevitch &amp; Luce, 1998</xref>, <xref ref-type="bibr" rid="B64">1999</xref>). Some evidence for this account comes from reaction time data, showing that non-words tend to show facilitation effects from high phonotactic similiarity, whereas words tend to be slowed down by competition effects (<xref ref-type="bibr" rid="B64">Vitevitch &amp; Luce, 1999</xref>). There is evidence of an effect of neighbourhood density in both the Word Identification Task and the Wellformedness Rating Task. In the Word Identification Task, neighbourhood density has an effect on ratings, but the modelling procedure did not show any significant interaction with words over nonwords, i.e. neighbourhood density had a positive effect on confidence ratings of stimuli generally, rather than just words. The Wellformedness Rating task shows a negative interaction between neighbourhood density and phonotactic score, which, as stated in section 3.2, we interpret in light of the nature of the task. Consequently our results do not support or refute this approach, given that we are not examining reaction times, and we did not design our stimuli to detect differences between phonotactic and neighbourhood effects; we included neighbourhood density as a control. Further experiments with non-M&#257;ori-speakers which are explicitly designed to probe questions relating to representation would be interesting. However the main conclusion of the current paper is that lexical knowledge feeds into phonotactic knowledge, and we do not claim anything further about how these types of knowledge are represented.</p>
<p>Our results, together with those reported in Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>), suggest that adult New Zealanders who do not speak M&#257;ori show the initial stages of language learning, similar to that identified in infants. The NMS&#8217;s sensitivity to phonotactics calculated over lexical forms provides a direct analogue to work on first language acquisition (<xref ref-type="bibr" rid="B30">Jusczyk et al., 1994</xref>), providing further evidence that such phonotactic knowledge does not require overt word knowledge. Just as with work on infants (<xref ref-type="bibr" rid="B41">Ngon et al., 2013</xref>), these processes are unlikely to be strictly sequential, but as each type of knowledge increases it will feed back into other stages and accelerate further learning.</p>
<p>Even further than this, the performance of the participants in our research is comparable to that of bilinguals, especially in their wellformedness judgments. In section 1.2.3, we presented evidence in the literature that bilinguals are sensitive to the phonotactics of both of their languages, even if they show differences from monolinguals. In this experiment, we show that participants were highly sensitive to M&#257;ori phonotactics. While we did not compare fluent and non-fluent speaker judgments in this study, Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>) found that the judgments of non-M&#257;ori speakers were almost indistinguishable from that of fluent speakers. This result is significant, because it shows that the participants have highly-developed intuitions comparable to that of bilinguals, despite not being obviously bilingual. The reason for this is undoubtedly related to long-term exposure to M&#257;ori.</p>
<p>One thing that we cannot know for certain, is how much exposure is needed to build up this kind of knowledge, nor whether adult exposure is sufficient, or whether exposure during childhood is also necessary. Our participants had all lived in New Zealand since they were seven-years-old, without any significant breaks, so this data-set does not allow us to disentangle this question. This is something we will be exploring in future work with participants with different types of exposure.</p>
<p>Nor can we be completely certain about the balance of different types of input in creating this knowledge. While we have focused on the fact that New Zealanders regularly hear spoken or sung M&#257;ori, it is also true that there is a reasonable amount of written M&#257;ori in the physical environment, in the form of street and town names, business names, email signatures, phrases in written reports, and so on. Unlike infants, our participants have an extra source of information in that they are literate. As written M&#257;ori contains word breaks (though not morpheme breaks), this may have somewhat contributed to identification of certain words. However, we consider it very unlikely that orthographical exposure is the key driver of our results, as New Zealanders who don&#8217;t speak M&#257;ori are unlikely to visually attend to substantial excerpts of written M&#257;ori, whereas they can&#8217;t help being exposed to substantial excerpts of spoken or sung M&#257;ori.</p>
<p>Significant further understanding of this phenomenon would come from work on exposure at different life stages, untangling the role of different types of exposure, and work on implicit learning of languages with less transparent orthography, and with different degrees of morphological transparency and complexity.</p>
</sec>
<sec>
<title>5 Conclusion</title>
<p>Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>) reported that non-M&#257;ori-speaking adult New Zealanders had a M&#257;ori proto-lexicon, and phonotactic sensitivity very similar to fluent speakers of M&#257;ori. Our work uses improved stimuli and a new set of participants, and draws the same conclusion. We go beyond Oh et al. (<xref ref-type="bibr" rid="B42">2020</xref>) by demonstrating a statistical link between their two experimental tasks. Participants who are best able to distinguish real M&#257;ori words from closely matched non-words are more sensitive to phonotactics in the Wellformedness Rating Task. This reinforces the interpretation that phonotactic knowledge emerges as a generalization over types in a lexicon &#8211; the more types you have, the more sophisticated the knowledge can be. It also suggests that this lexicon need not be overt: A proto-lexicon &#8211; a set of stored forms, which may not be associated with semantic knowledge or even overt awareness &#8211; can form a basis for phonotactic knowledge.</p>
</sec>
<sec>
<title>Supplementary materials</title>
<p>The supplementary materials of this paper are found at: <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://github.com/FPanther/Replication22">https://github.com/FPanther/Replication22</ext-link></p>
</sec>
</body>
<back>
<ack>
<title>Acknowledgements</title>
<p>This work was made possible by the use of the RCC facilities at the University of Canterbury. We acknowledge Chun-Liang Chan for the original development of the software underpinning our experiments.</p>
<p>We thank Yoon Mi Oh for providing guidance on running the online experimentation. We thank Robert Fromont for his technical support and help during the deployment of the online experimentation. We also thank Peter Keegan. This research has benefited from feedback from colleagues at the NZILBB.</p>
<p>This research was supported by the Marsden Fund (UOC1502 and UOC1908).</p>
</ack>
<sec>
<title>Competing interests</title>
<p>The authors have no competing interests to declare.</p>
</sec>
<ref-list>
<ref id="B1"><label>1</label><mixed-citation publication-type="journal"><string-name><surname>Archer</surname>, <given-names>S. L.</given-names></string-name>, &amp; <string-name><surname>Curtin</surname>, <given-names>S.</given-names></string-name> (<year>2011</year>). <article-title>Perceiving onset clusters in infancy</article-title>. <source>Infant behavior and development</source>, <volume>34</volume>(<issue>4</issue>), <fpage>534</fpage>&#8211;<lpage>540</lpage>. DOI: <pub-id pub-id-type="doi">10.1016/j.infbeh.2011.07.001</pub-id></mixed-citation></ref>
<ref id="B2"><label>2</label><mixed-citation publication-type="journal"><string-name><surname>Archer</surname>, <given-names>S. L.</given-names></string-name>, &amp; <string-name><surname>Curtin</surname>, <given-names>S.</given-names></string-name> (<year>2016</year>). <article-title>Nine-month-olds use frequency of onset clusters to segment novel words</article-title>. <source>Journal of experimental child psychology</source>, <volume>148</volume>, <fpage>131</fpage>&#8211;<lpage>141</lpage>. DOI: <pub-id pub-id-type="doi">10.1016/j.jecp.2016.04.004</pub-id></mixed-citation></ref>
<ref id="B3"><label>3</label><mixed-citation publication-type="journal"><string-name><surname>Bailey</surname>, <given-names>T. M.</given-names></string-name>, &amp; <string-name><surname>Hahn</surname>, <given-names>U.</given-names></string-name> (<year>2001</year>). <article-title>Determinants of wordlikeness: Phonotactics or lexical neighborhoods?</article-title> <source>Journal of Memory and Language</source>, <volume>44</volume>(<issue>4</issue>), <fpage>568</fpage>&#8211;<lpage>591</lpage>. DOI: <pub-id pub-id-type="doi">10.1006/jmla.2000.2756</pub-id></mixed-citation></ref>
<ref id="B4"><label>4</label><mixed-citation publication-type="journal"><string-name><surname>Barr</surname>, <given-names>D. J.</given-names></string-name>, <string-name><surname>Levy</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Scheepers</surname>, <given-names>C.</given-names></string-name>, &amp; <string-name><surname>Tily</surname>, <given-names>H. J.</given-names></string-name> (<year>2013</year>). <article-title>Random effects structure for confirmatory hypothesis testing: Keep it maximal</article-title>. <source>Journal of memory and language</source>, <volume>68</volume>(<issue>3</issue>), <fpage>255</fpage>&#8211;<lpage>278</lpage>. DOI: <pub-id pub-id-type="doi">10.1016/j.jml.2012.11.001</pub-id></mixed-citation></ref>
<ref id="B5"><label>5</label><mixed-citation publication-type="journal"><string-name><surname>Bergelson</surname>, <given-names>E.</given-names></string-name>, &amp; <string-name><surname>Swingley</surname>, <given-names>D.</given-names></string-name> (<year>2012</year>). <article-title>At 6&#8211;9 months, human infants know the meanings of many common nouns</article-title>. <source>Proceedings of the National Academy of Sciences</source>, <volume>109</volume>(<issue>9</issue>), <fpage>3253</fpage>&#8211;<lpage>3258</lpage>. DOI: <pub-id pub-id-type="doi">10.1073/pnas.1113380109</pub-id></mixed-citation></ref>
<ref id="B6"><label>6</label><mixed-citation publication-type="journal"><string-name><surname>Bortfeld</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Morgan</surname>, <given-names>J. L.</given-names></string-name>, <string-name><surname>Golinkoff</surname>, <given-names>R. M.</given-names></string-name>, &amp; <string-name><surname>Rathbun</surname>, <given-names>K.</given-names></string-name> (<year>2005</year>). <article-title>Mommy and me: Familiar names help launch babies into speech-stream segmentation</article-title>. <source>Psychological science</source>, <volume>16</volume>(<issue>4</issue>), <fpage>298</fpage>&#8211;<lpage>304</lpage>. DOI: <pub-id pub-id-type="doi">10.1111/j.0956-7976.2005.01531.x</pub-id></mixed-citation></ref>
<ref id="B7"><label>7</label><mixed-citation publication-type="thesis"><string-name><surname>Boyce</surname>, <given-names>M. T.</given-names></string-name> (<year>2006</year>). <source>A corpus of modern spoken M&#257;ori</source>. (Unpublished doctoral dissertation). <publisher-name>Victoria University of Wellington</publisher-name>.</mixed-citation></ref>
<ref id="B8"><label>8</label><mixed-citation publication-type="journal"><string-name><surname>Cairns</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Shillcock</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Chater</surname>, <given-names>N.</given-names></string-name>, &amp; <string-name><surname>Levy</surname>, <given-names>J.</given-names></string-name> (<year>1997</year>). <article-title>Bootstrapping word boundaries: A bottom-up corpus-based approach to speech segmentation</article-title>. <source>Cognitive Psychology</source>, <volume>33</volume>(<issue>2</issue>), <fpage>111</fpage>&#8211;<lpage>153</lpage>. DOI: <pub-id pub-id-type="doi">10.1006/cogp.1997.0649</pub-id></mixed-citation></ref>
<ref id="B9"><label>9</label><mixed-citation publication-type="journal"><string-name><surname>Chambers</surname>, <given-names>K. E.</given-names></string-name>, <string-name><surname>Onishi</surname>, <given-names>K. H.</given-names></string-name>, &amp; <string-name><surname>Fisher</surname>, <given-names>C.</given-names></string-name> (<year>2003</year>). <article-title>Infants learn phonotactic regularities from brief auditory experience</article-title>. <source>Cognition</source>, <volume>87</volume>(<issue>2</issue>), <fpage>B69</fpage>&#8211;<lpage>B77</lpage>. DOI: <pub-id pub-id-type="doi">10.1016/s0010-0277(02)00233-0</pub-id></mixed-citation></ref>
<ref id="B10"><label>10</label><mixed-citation publication-type="book"><string-name><surname>Chan</surname>, <given-names>C. L.</given-names></string-name> (<year>2018</year>). <source>Speech In Noise 2</source>. <publisher-name>Northwestern University</publisher-name>. (Computer program)</mixed-citation></ref>
<ref id="B11"><label>11</label><mixed-citation publication-type="webpage"><string-name><surname>Christensen</surname>, <given-names>R. H. B.</given-names></string-name> (<year>2019</year>). <source>ordinal&#8212;regression models for ordinal data</source>. (R package version 2019.12-10. <uri>https://CRAN.R-project.org/package=ordinal</uri>)</mixed-citation></ref>
<ref id="B12"><label>12</label><mixed-citation publication-type="journal"><string-name><surname>Coady</surname>, <given-names>J. A.</given-names></string-name>, &amp; <string-name><surname>Aslin</surname>, <given-names>R. N.</given-names></string-name> (<year>2004</year>). <article-title>Young children&#8217;s sensitivity to probabilistic phonotactics in the developing lexicon</article-title>. <source>Journal of experimental child psychology</source>, <volume>89</volume>(<issue>3</issue>), <fpage>183</fpage>&#8211;<lpage>213</lpage>. DOI: <pub-id pub-id-type="doi">10.1016/j.jecp.2004.07.004</pub-id></mixed-citation></ref>
<ref id="B13"><label>13</label><mixed-citation publication-type="webpage"><string-name><surname>Coleman</surname>, <given-names>J.</given-names></string-name>, &amp; <string-name><surname>Pierrehumbert</surname>, <given-names>J.</given-names></string-name> (<year>1997</year>). <chapter-title>Stochastic phonological grammars and acceptability</chapter-title>. In <source>Computational phonology: Third meeting of the acl special interest group in computational phonology</source>. Retrieved from <uri>https://aclanthology.org/W97-1107</uri></mixed-citation></ref>
<ref id="B14"><label>14</label><mixed-citation publication-type="book"><string-name><surname>Curtin</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name><surname>Archer</surname>, <given-names>S. L.</given-names></string-name> (<year>2015</year>). <chapter-title>Speech perception</chapter-title>. In <string-name><given-names>E. L.</given-names> <surname>Bavin</surname></string-name> &amp; <string-name><given-names>L. R.</given-names> <surname>Naigles</surname></string-name> (Eds.), <source>The Cambridge Handbook of Child Language</source> (<edition>2nd</edition> ed., p. <fpage>137</fpage>&#8211;<lpage>158</lpage>). <publisher-name>Cambridge University Press</publisher-name>. DOI: <pub-id pub-id-type="doi">10.1017/CBO9781316095829.007</pub-id></mixed-citation></ref>
<ref id="B15"><label>15</label><mixed-citation publication-type="book"><string-name><surname>Curtin</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name><surname>Hufnagle</surname>, <given-names>D.</given-names></string-name> (<year>2009</year>). <chapter-title>Speech perception</chapter-title>. In <string-name><given-names>E.</given-names> <surname>Bavin</surname></string-name> (Ed.), <source>The Cambridge Handbook of Child Language</source> (<edition>1st</edition> ed., pp. <fpage>107</fpage>&#8211;<lpage>127</lpage>). <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>. DOI: <pub-id pub-id-type="doi">10.1017/CBO9781316095829.007</pub-id></mixed-citation></ref>
<ref id="B16"><label>16</label><mixed-citation publication-type="journal"><string-name><surname>de Bot</surname>, <given-names>K.</given-names></string-name>, &amp; <string-name><surname>Stoessel</surname>, <given-names>S.</given-names></string-name> (<year>2000</year>). <article-title>In search of yesterday&#8217;s words: Reactivating a long-forgotten language</article-title>. <source>Applied Linguistics</source>, <volume>21</volume>(<issue>3</issue>), <fpage>333</fpage>&#8211;<lpage>353</lpage>. DOI: <pub-id pub-id-type="doi">10.1093/applin/21.3.333</pub-id></mixed-citation></ref>
<ref id="B17"><label>17</label><mixed-citation publication-type="book"><string-name><surname>Demuth</surname>, <given-names>K.</given-names></string-name> (<year>2009</year>). <chapter-title>The prosody of syllables, words and morphemes</chapter-title>. In <source>Cambridge Handbook of Child Language</source> (<edition>1st</edition> ed., pp. <fpage>183</fpage>&#8211;<lpage>98</lpage>). <publisher-name>Cambridge University Press</publisher-name>. DOI: <pub-id pub-id-type="doi">10.1017/CBO9780511576164.011</pub-id></mixed-citation></ref>
<ref id="B18"><label>18</label><mixed-citation publication-type="journal"><string-name><surname>Edwards</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Beckman</surname>, <given-names>M. E.</given-names></string-name>, &amp; <string-name><surname>Munson</surname>, <given-names>B.</given-names></string-name> (<year>2004</year>). <article-title>The interaction between vocabulary size and phonotactic probability effects on children&#8217;s production accuracy and fluency in nonword repetition</article-title>. DOI: <pub-id pub-id-type="doi">10.1044/1092-4388(2004/034)</pub-id></mixed-citation></ref>
<ref id="B19"><label>19</label><mixed-citation publication-type="journal"><string-name><surname>Frank</surname>, <given-names>M. C.</given-names></string-name>, <string-name><surname>Tenenbaum</surname>, <given-names>J. B.</given-names></string-name>, &amp; <string-name><surname>Gibson</surname>, <given-names>E.</given-names></string-name> (<year>2013</year>, 01). <article-title>Learning and long-term retention of large-scale artificial languages</article-title>. <source>PLOS ONE</source>, <volume>8</volume>, <fpage>1</fpage>&#8211;<lpage>6</lpage>. DOI: <pub-id pub-id-type="doi">10.1371/journal.pone.0052500</pub-id></mixed-citation></ref>
<ref id="B20"><label>20</label><mixed-citation publication-type="journal"><string-name><surname>Frisch</surname>, <given-names>S. A.</given-names></string-name>, &amp; <string-name><surname>Brea-Spahn</surname>, <given-names>M. R.</given-names></string-name> (<year>2010</year>). <article-title>Metalinguistic judgments of phonotactics by monolinguals and bilinguals</article-title>. <source>Laboratory Phonology</source>, <volume>1</volume>(<issue>2</issue>), <fpage>345</fpage>&#8211;<lpage>360</lpage>. DOI: <pub-id pub-id-type="doi">10.1515/labphon.2010.018</pub-id></mixed-citation></ref>
<ref id="B21"><label>21</label><mixed-citation publication-type="journal"><string-name><surname>Frisch</surname>, <given-names>S. A.</given-names></string-name>, <string-name><surname>Large</surname>, <given-names>N. R.</given-names></string-name>, &amp; <string-name><surname>Pisoni</surname>, <given-names>D. B.</given-names></string-name> (<year>2000</year>). <article-title>Perception of wordlikeness: Effects of segment probability and length on the processing of nonwords</article-title>. <source>Journal of memory and language</source>, <volume>42</volume>(<issue>4</issue>), <fpage>481</fpage>&#8211;<lpage>496</lpage>. DOI: <pub-id pub-id-type="doi">10.1006/jmla.1999.2692</pub-id></mixed-citation></ref>
<ref id="B22"><label>22</label><mixed-citation publication-type="book"><string-name><surname>Frisch</surname>, <given-names>S. A.</given-names></string-name>, <string-name><surname>Large</surname>, <given-names>N. R.</given-names></string-name>, <string-name><surname>Zawaydeh</surname>, <given-names>B.</given-names></string-name>, &amp; <string-name><surname>Pisoni</surname>, <given-names>D. B.</given-names></string-name> (<year>2001</year>). <chapter-title>Emergent phonotactic generalizations in english and arabic</chapter-title>. In <string-name><given-names>J.</given-names> <surname>Bybee</surname></string-name> &amp; <string-name><given-names>P.</given-names> <surname>Hopper</surname></string-name> (Eds.), <source>Frequency and the emergence of linguistic structure</source>. <publisher-loc>Amsterdam/Philadelphia</publisher-loc>: <publisher-name>John Benjamins</publisher-name>. DOI: <pub-id pub-id-type="doi">10.1075/tsl.45.09fri</pub-id></mixed-citation></ref>
<ref id="B23"><label>23</label><mixed-citation publication-type="journal"><string-name><surname>Hall&#233;</surname>, <given-names>P. A.</given-names></string-name>, &amp; <string-name><surname>de Boysson-Bardies</surname>, <given-names>B.</given-names></string-name> (<year>1996</year>). <article-title>The format of representation of recognized words in infants&#8217; early receptive lexicon</article-title>. <source>Infant Behavior and Development</source>, <volume>19</volume>(<issue>4</issue>), <fpage>463</fpage>&#8211;<lpage>481</lpage>. DOI: <pub-id pub-id-type="doi">10.1016/S0163-6383(96)90007-7</pub-id></mixed-citation></ref>
<ref id="B24"><label>24</label><mixed-citation publication-type="journal"><string-name><surname>Hansen</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Umeda</surname>, <given-names>Y.</given-names></string-name>, &amp; <string-name><surname>McKinney</surname>, <given-names>M.</given-names></string-name> (<year>2002</year>). <article-title>Savings in the relearning of second language vocabulary: The effects of time and proficiency</article-title>. <source>Language Learning</source>, <volume>52</volume>(<issue>4</issue>), <fpage>653</fpage>&#8211;<lpage>678</lpage>. DOI: <pub-id pub-id-type="doi">10.1111/1467-9922.00200</pub-id></mixed-citation></ref>
<ref id="B25"><label>25</label><mixed-citation publication-type="book"><string-name><surname>Hay</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Pierrehumbert</surname>, <given-names>J.</given-names></string-name>, &amp; <string-name><surname>Beckman</surname>, <given-names>M.</given-names></string-name> (<year>2004</year>). <chapter-title>Speech perception, well-formedness and the statistics of the lexicon</chapter-title>. In <string-name><given-names>J.</given-names> <surname>Local</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Ogden</surname></string-name>, &amp; <string-name><given-names>R.</given-names> <surname>Temple</surname></string-name> (Eds.), <source>Papers in laboratory phonology vi</source>. <publisher-name>Cambridge University Press</publisher-name>. DOI: <pub-id pub-id-type="doi">10.1017/CBO9780511486425.004</pub-id></mixed-citation></ref>
<ref id="B26"><label>26</label><mixed-citation publication-type="journal"><string-name><surname>Johnson</surname>, <given-names>E. K.</given-names></string-name> (<year>2016</year>). <article-title>Constructing a proto-lexicon: An integrative view of infant language development</article-title>. <source>Annual Review of Linguistics</source>, <volume>2</volume>, <fpage>391</fpage>&#8211;<lpage>412</lpage>. DOI: <pub-id pub-id-type="doi">10.1146/annurev-linguistics-011415-040616</pub-id></mixed-citation></ref>
<ref id="B27"><label>27</label><mixed-citation publication-type="journal"><string-name><surname>Johnson</surname>, <given-names>E. K.</given-names></string-name>, <string-name><surname>Seidl</surname>, <given-names>A.</given-names></string-name>, &amp; <string-name><surname>Tyler</surname>, <given-names>M. D.</given-names></string-name> (<year>2014</year>). <article-title>The edge factor in early word segmentation: utterance-level prosody enables word form extraction by 6-month-olds</article-title>. <source>PloS one</source>, <volume>9</volume>(<issue>1</issue>), <elocation-id>e83546</elocation-id>. DOI: <pub-id pub-id-type="doi">10.1371/journal.pone.0083546</pub-id></mixed-citation></ref>
<ref id="B28"><label>28</label><mixed-citation publication-type="book"><string-name><surname>Junge</surname>, <given-names>C.</given-names></string-name> (<year>2017</year>). <chapter-title>The proto-lexicon: segmenting word-like units from speech</chapter-title>. In <string-name><given-names>G.</given-names> <surname>Westermann</surname></string-name> &amp; <string-name><given-names>N.</given-names> <surname>Mani</surname></string-name> (Eds.), <source>Early word learning</source> (pp. <fpage>15</fpage>&#8211;<lpage>29</lpage>). <publisher-name>Routledge</publisher-name>. DOI: <pub-id pub-id-type="doi">10.4324/9781315730974-2</pub-id></mixed-citation></ref>
<ref id="B29"><label>29</label><mixed-citation publication-type="book"><string-name><surname>Jusczyk</surname>, <given-names>P. W.</given-names></string-name> (<year>2000</year>). <source>The Discovery of Spoken Language</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>. DOI: <pub-id pub-id-type="doi">10.7551/mitpress/2447.001.0001</pub-id></mixed-citation></ref>
<ref id="B30"><label>30</label><mixed-citation publication-type="journal"><string-name><surname>Jusczyk</surname>, <given-names>P. W.</given-names></string-name>, <string-name><surname>Luce</surname>, <given-names>P. A.</given-names></string-name>, &amp; Charles-Luce, J. (<year>1994</year>). <article-title>Infants&#8242; sensitivity to phonotactic patterns in the native language</article-title>. <source>Journal of Memory and Language</source>, <volume>33</volume>(<issue>5</issue>), <fpage>630</fpage>&#8211;<lpage>645</lpage>. DOI: <pub-id pub-id-type="doi">10.1006/jmla.1994.1030</pub-id></mixed-citation></ref>
<ref id="B31"><label>31</label><mixed-citation publication-type="journal"><string-name><surname>King</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Maclagan</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Harlow</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Keegan</surname>, <given-names>P.</given-names></string-name>, &amp; <string-name><surname>Watson</surname>, <given-names>C.</given-names></string-name> (<year>2010</year>). <article-title>The MAONZE corpus: Establishing a corpus of M&#257;ori speech</article-title>. <source>New Zealand Studies in Applied Linguistics</source>, <volume>16</volume>(<issue>2</issue>), <fpage>1</fpage>&#8211;<lpage>16</lpage>.</mixed-citation></ref>
<ref id="B32"><label>32</label><mixed-citation publication-type="journal"><string-name><surname>Kittleson</surname>, <given-names>M. M.</given-names></string-name>, <string-name><surname>Aguilar</surname>, <given-names>J. M.</given-names></string-name>, <string-name><surname>Tokerud</surname>, <given-names>G. L.</given-names></string-name>, <string-name><surname>Plante</surname>, <given-names>E.</given-names></string-name>, &amp; <string-name><surname>Asbj&#248;rnsen</surname>, <given-names>A. E.</given-names></string-name> (<year>2010</year>). <article-title>Implicit language learning: Adults&#8217; ability to segment words in norwegian</article-title>. <source>Bilingualism: Language and Cognition</source>, <volume>13</volume>(<issue>4</issue>), <fpage>513</fpage>&#8211;<lpage>523</lpage>. DOI: <pub-id pub-id-type="doi">10.1017/S1366728910000039</pub-id></mixed-citation></ref>
<ref id="B33"><label>33</label><mixed-citation publication-type="journal"><string-name><surname>Kuhl</surname>, <given-names>P. K.</given-names></string-name> (<year>2004</year>). <article-title>Early language acquisition: cracking the speech code</article-title>. <source>Nature reviews neuroscience</source>, <volume>5</volume>(<issue>11</issue>), <fpage>831</fpage>&#8211;<lpage>843</lpage>. DOI: <pub-id pub-id-type="doi">10.1038/nrn1533</pub-id></mixed-citation></ref>
<ref id="B34"><label>34</label><mixed-citation publication-type="journal"><string-name><surname>Macalister</surname>, <given-names>J.</given-names></string-name> (<year>2004</year>). <article-title>A survey of M&#257;ori word knowledge</article-title>. <source>English in Aotearoa</source>, <volume>52</volume>, <fpage>69</fpage>&#8211;<lpage>73</lpage>.</mixed-citation></ref>
<ref id="B35"><label>35</label><mixed-citation publication-type="journal"><string-name><surname>MacKenzie</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Curtin</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name><surname>Graham</surname>, <given-names>S. A.</given-names></string-name> (<year>2012</year>). <article-title>12-month-olds&#8217; phonotactic knowledge guides their word&#8211;object mappings</article-title>. <source>Child development</source>, <volume>83</volume>(<issue>4</issue>), <fpage>1129</fpage>&#8211;<lpage>1136</lpage>. DOI: <pub-id pub-id-type="doi">10.1111/j.1467-8624.2012.01764.x</pub-id></mixed-citation></ref>
<ref id="B36"><label>36</label><mixed-citation publication-type="journal"><string-name><surname>Martin</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Peperkamp</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name><surname>Dupoux</surname>, <given-names>E.</given-names></string-name> (<year>2013</year>). <article-title>Learning phonemes with a proto-lexicon</article-title>. <source>Cognitive science</source>, <volume>37</volume>(<issue>1</issue>), <fpage>103</fpage>&#8211;<lpage>124</lpage>. DOI: <pub-id pub-id-type="doi">10.1111/j.1551-6709.2012.01267.x</pub-id></mixed-citation></ref>
<ref id="B37"><label>37</label><mixed-citation publication-type="journal"><string-name><surname>Mattys</surname>, <given-names>S. L.</given-names></string-name>, &amp; <string-name><surname>Jusczyk</surname>, <given-names>P. W.</given-names></string-name> (<year>2001</year>). <article-title>Phonotactic cues for segmentation of fluent speech by infants</article-title>. <source>Cognition</source>, <volume>78</volume>(<issue>2</issue>), <fpage>91</fpage>&#8211;<lpage>121</lpage>. DOI: <pub-id pub-id-type="doi">10.1016/S0010-0277(00)00109-8</pub-id></mixed-citation></ref>
<ref id="B38"><label>38</label><mixed-citation publication-type="journal"><string-name><surname>Messer</surname>, <given-names>M. H.</given-names></string-name>, <string-name><surname>Leseman</surname>, <given-names>P. P.</given-names></string-name>, <string-name><surname>Boom</surname>, <given-names>J.</given-names></string-name>, &amp; <string-name><surname>Mayo</surname>, <given-names>A. Y.</given-names></string-name> (<year>2010</year>). <article-title>Phonotactic probability effect in nonword recall and its relationship with vocabulary in monolingual and bilingual preschoolers</article-title>. <source>Journal of Experimental Child Psychology</source>, <volume>105</volume>(<issue>4</issue>), <fpage>306</fpage>&#8211;<lpage>323</lpage>. DOI: <pub-id pub-id-type="doi">10.1016/j.jecp.2009.12.006</pub-id></mixed-citation></ref>
<ref id="B39"><label>39</label><mixed-citation publication-type="webpage"><string-name><surname>Moorfield</surname>, <given-names>J.</given-names></string-name> (n.d.). <source>Te Aka Online M&#257;ori Dictionary</source>. Retrieved from <uri>https://maoridictionary.co.nz/</uri></mixed-citation></ref>
<ref id="B40"><label>40</label><mixed-citation publication-type="book"><string-name><surname>Newport</surname>, <given-names>E. L.</given-names></string-name>, &amp; <string-name><surname>Aslin</surname>, <given-names>R. N.</given-names></string-name> (<year>2000</year>). <chapter-title>Innately constrained learning: Blending old and new approaches to language acquisition</chapter-title>. In <source>Proceedings of the 24th annual boston university conference on language development</source> (Vol. <volume>1</volume>, pp. <fpage>1</fpage>&#8211;<lpage>21</lpage>).</mixed-citation></ref>
<ref id="B41"><label>41</label><mixed-citation publication-type="journal"><string-name><surname>Ngon</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Martin</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Dupoux</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Cabrol</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Dutat</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Peperkamp</surname>, <given-names>S.</given-names></string-name> (<year>2013</year>). <article-title>(non) words,(non) words,(non) words: evidence for a protolexicon during the first year of life</article-title>. <source>Developmental Science</source>, <volume>16</volume>(<issue>1</issue>), <fpage>24</fpage>&#8211;<lpage>34</lpage>. DOI: <pub-id pub-id-type="doi">10.1111/j.1467-7687.2012.01189.x</pub-id></mixed-citation></ref>
<ref id="B42"><label>42</label><mixed-citation publication-type="journal"><string-name><surname>Oh</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Todd</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Beckner</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Hay</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>King</surname>, <given-names>J.</given-names></string-name>, &amp; <string-name><surname>Needle</surname>, <given-names>J.</given-names></string-name> (<year>2020</year>). <article-title>Non-M&#257;ori-speaking New Zealanders have a M&#257;ori proto-lexicon</article-title>. <source>Scientific reports</source>, <volume>10</volume>(<issue>1</issue>), <fpage>1</fpage>&#8211;<lpage>9</lpage>. DOI: <pub-id pub-id-type="doi">10.1038/s41598-020-78810-4</pub-id></mixed-citation></ref>
<ref id="B43"><label>43</label><mixed-citation publication-type="journal"><string-name><surname>Onnis</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Monaghan</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Richmond</surname>, <given-names>K.</given-names></string-name>, &amp; <string-name><surname>Chater</surname>, <given-names>N.</given-names></string-name> (<year>2005</year>). <article-title>Phonology impacts segmentation in online speech processing</article-title>. <source>Journal of Memory and Language</source>, <volume>53</volume>(<issue>2</issue>), <fpage>225</fpage>&#8211;<lpage>237</lpage>. DOI: <pub-id pub-id-type="doi">10.1016/j.jml.2005.02.011</pub-id></mixed-citation></ref>
<ref id="B44"><label>44</label><mixed-citation publication-type="journal"><string-name><surname>Parise</surname>, <given-names>E.</given-names></string-name>, &amp; <string-name><surname>Csibra</surname>, <given-names>G.</given-names></string-name> (<year>2012</year>). <article-title>Electrophysiological evidence for the understanding of maternal speech by 9-month-old infants</article-title>. <source>Psychological Science</source>, <volume>23</volume>(<issue>7</issue>), <fpage>728</fpage>&#8211;<lpage>733</lpage>. DOI: <pub-id pub-id-type="doi">10.1177/0956797612438734</pub-id></mixed-citation></ref>
<ref id="B45"><label>45</label><mixed-citation publication-type="webpage"><string-name><surname>Parkinson</surname>, <given-names>P.</given-names></string-name> (<year>2016</year>). <article-title>The M&#257;ori grammars and vocabularies of Thomas Kendall and John Gare Butler</article-title>. <source>Asia-Pacific Linguistics</source>, <volume>26</volume>, <fpage>1</fpage>&#8211;<lpage>163</lpage>. Retrieved from <uri>http://hdl.handle.net/1885/104299</uri></mixed-citation></ref>
<ref id="B46"><label>46</label><mixed-citation publication-type="journal"><string-name><surname>Pe&#241;a</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Bonatti</surname>, <given-names>L. L.</given-names></string-name>, <string-name><surname>Nespor</surname>, <given-names>M.</given-names></string-name>, &amp; <string-name><surname>Mehler</surname>, <given-names>J.</given-names></string-name> (<year>2002</year>). <article-title>Signal-driven computations in speech processing</article-title>. <source>Science</source>, <volume>298</volume>(<issue>5593</issue>), <fpage>604</fpage>&#8211;<lpage>607</lpage>. DOI: <pub-id pub-id-type="doi">10.1126/science.1072901</pub-id></mixed-citation></ref>
<ref id="B47"><label>47</label><mixed-citation publication-type="journal"><string-name><surname>Perszyk</surname>, <given-names>D. R.</given-names></string-name>, &amp; <string-name><surname>Waxman</surname>, <given-names>S. R.</given-names></string-name> (<year>2019</year>). <article-title>Infants&#8217; advances in speech perception shape their earliest links between language and cognition</article-title>. <source>Scientific reports</source>, <volume>9</volume>(<issue>1</issue>), <fpage>1</fpage>&#8211;<lpage>6</lpage>. DOI: <pub-id pub-id-type="doi">10.1038/s41598-019-39511-9</pub-id></mixed-citation></ref>
<ref id="B48"><label>48</label><mixed-citation publication-type="journal"><string-name><surname>Pierrehumbert</surname>, <given-names>J. B.</given-names></string-name> (<year>2003</year>). <article-title>Phonetic diversity, statistical learning, and acquisition of phonology</article-title>. <source>Language and speech</source>, <volume>46</volume>(<issue>2&#8211;3</issue>), <fpage>115</fpage>&#8211;<lpage>154</lpage>. DOI: <pub-id pub-id-type="doi">10.1177/00238309030460020501</pub-id></mixed-citation></ref>
<ref id="B49"><label>49</label><mixed-citation publication-type="webpage"><collab>R Core Team</collab>. (<year>2021</year>). <chapter-title>R: A language and environment for statistical computing [Computer software manual]</chapter-title>. <publisher-loc>Vienna, Austria</publisher-loc>. Retrieved from <uri>https://www.R-project.org/</uri></mixed-citation></ref>
<ref id="B50"><label>50</label><mixed-citation publication-type="journal"><string-name><surname>Richtsmeier</surname>, <given-names>P. T.</given-names></string-name> (<year>2011</year>). <article-title>Word-types not word-tokens, facilitate extraction of phonotactic sequences by adults</article-title>. <source>Laboratory Phonology</source>, <volume>2</volume>, <fpage>157</fpage>&#8211;<lpage>183</lpage>. DOI: <pub-id pub-id-type="doi">10.1515/labphon.2011.005</pub-id></mixed-citation></ref>
<ref id="B51"><label>51</label><mixed-citation publication-type="journal"><string-name><surname>Saffran</surname>, <given-names>J. R.</given-names></string-name>, <string-name><surname>Aslin</surname>, <given-names>R. N.</given-names></string-name>, &amp; <string-name><surname>Newport</surname>, <given-names>E. L.</given-names></string-name> (<year>1996</year>). <article-title>Statistical learning by 8-month-old infants</article-title>. <source>Science</source>, <volume>274</volume>(<issue>5294</issue>), <fpage>1926</fpage>&#8211;<lpage>1928</lpage>. DOI: <pub-id pub-id-type="doi">10.1126/science.274.5294.1926</pub-id></mixed-citation></ref>
<ref id="B52"><label>52</label><mixed-citation publication-type="journal"><string-name><surname>Saffran</surname>, <given-names>J. R.</given-names></string-name>, <string-name><surname>Werker</surname>, <given-names>J. F.</given-names></string-name>, &amp; <string-name><surname>Werner</surname>, <given-names>L. A.</given-names></string-name> (<year>2007</year>). <article-title>The infant&#8217;s auditory world: Hearing, speech, and the beginnings of language</article-title>. <source>Handbook of child psychology</source>, <fpage>2</fpage>. DOI: <pub-id pub-id-type="doi">10.1002/9780470147658.chpsy0202</pub-id></mixed-citation></ref>
<ref id="B53"><label>53</label><mixed-citation publication-type="journal"><string-name><surname>Sebasti&#225;n-Gall&#233;s</surname>, <given-names>N.</given-names></string-name>, &amp; <string-name><surname>Bosch</surname>, <given-names>L.</given-names></string-name> (<year>2002</year>). <article-title>Building phonotactic knowledge in bilinguals: role of early exposure</article-title>. <source>Journal of Experimental Psychology: Human Perception and Performance</source>, <volume>28</volume>(<issue>4</issue>), <fpage>974</fpage>. DOI: <pub-id pub-id-type="doi">10.1037/0096-1523.28.4.974</pub-id></mixed-citation></ref>
<ref id="B54"><label>54</label><mixed-citation publication-type="journal"><string-name><surname>Shi</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Morgan</surname>, <given-names>J. L.</given-names></string-name>, &amp; <string-name><surname>Allopenna</surname>, <given-names>P.</given-names></string-name> (<year>1998</year>). <article-title>Phonological and acoustic bases for earliest grammatical category assignment: A cross-linguistic perspective</article-title>. <source>Journal of child language</source>, <volume>25</volume>(<issue>1</issue>), <fpage>169</fpage>&#8211;<lpage>201</lpage>. DOI: <pub-id pub-id-type="doi">10.1017/S0305000997003395</pub-id></mixed-citation></ref>
<ref id="B55"><label>55</label><mixed-citation publication-type="journal"><string-name><surname>Shi</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Werker</surname>, <given-names>J. F.</given-names></string-name>, &amp; <string-name><surname>Morgan</surname>, <given-names>J. L.</given-names></string-name> (<year>1999</year>). <article-title>Newborn infants&#8217; sensitivity to perceptual cues to lexical and grammatical words</article-title>. <source>Cognition</source>, <volume>72</volume>(<issue>2</issue>), <fpage>B11</fpage>&#8211;<lpage>B21</lpage>. DOI: <pub-id pub-id-type="doi">10.1016/S0010-0277(99)00047-5</pub-id></mixed-citation></ref>
<ref id="B56"><label>56</label><mixed-citation publication-type="webpage"><collab>Statistics New Zealand</collab>. (<year>2020</year>). <source>2018 census totals by topic &#8211; national highlights</source>. Retrieved 2021-06-24, from <uri>https://www.stats.govt.nz/information-releases/2018-census-totals-by-topic-national-highlights-updated</uri></mixed-citation></ref>
<ref id="B57"><label>57</label><mixed-citation publication-type="journal"><string-name><surname>Steber</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name><surname>Rossi</surname>, <given-names>S.</given-names></string-name> (<year>2020</year>). <article-title>So young, yet so mature? electrophysiological and vascular correlates of phonotactic processing in 18-month-olds</article-title>. <source>Developmental cognitive neuroscience</source>, <volume>43</volume>, <fpage>100784</fpage>. DOI: <pub-id pub-id-type="doi">10.1016/j.dcn.2020.100784</pub-id></mixed-citation></ref>
<ref id="B58"><label>58</label><mixed-citation publication-type="journal"><string-name><surname>Storkel</surname>, <given-names>H. L.</given-names></string-name>, &amp; <string-name><surname>Rogers</surname>, <given-names>M. A.</given-names></string-name> (<year>2000</year>). <article-title>The effect of probabilistic phonotactics on lexical acquisition</article-title>. <source>clinical linguistics &amp; phonetics</source>, <volume>14</volume>(<issue>6</issue>), <fpage>407</fpage>&#8211;<lpage>425</lpage>. DOI: <pub-id pub-id-type="doi">10.1080/026992000415859</pub-id></mixed-citation></ref>
<ref id="B59"><label>59</label><mixed-citation publication-type="journal"><string-name><surname>Swingley</surname>, <given-names>D.</given-names></string-name> (<year>2005</year>). <article-title>Statistical clustering and the contents of the infant vocabulary</article-title>. <source>Cognitive psychology</source>, <volume>50</volume>(<issue>1</issue>), <fpage>86</fpage>&#8211;<lpage>132</lpage>. DOI: <pub-id pub-id-type="doi">10.1016/j.cogpsych.2004.06.001</pub-id></mixed-citation></ref>
<ref id="B60"><label>60</label><mixed-citation publication-type="journal"><string-name><surname>Swingley</surname>, <given-names>D.</given-names></string-name> (<year>2009</year>). <article-title>Contributions of infant word learning to language development</article-title>. <source>Philosophical Transactions of the Royal Society B: Biological Sciences</source>, <volume>364</volume>(<issue>1536</issue>), <fpage>3617</fpage>&#8211;<lpage>3632</lpage>. DOI: <pub-id pub-id-type="doi">10.1098/rstb.2009.0107</pub-id></mixed-citation></ref>
<ref id="B61"><label>61</label><mixed-citation publication-type="journal"><string-name><surname>van der Hoeven</surname>, <given-names>N.</given-names></string-name>, &amp; <string-name><surname>de Bot</surname>, <given-names>K.</given-names></string-name> (<year>2012</year>). <article-title>Relearning in the elderly: age-related effects on the size of savings</article-title>. <source>Language Learning</source>, <volume>62</volume>(<issue>1</issue>), <fpage>42</fpage>&#8211;<lpage>67</lpage>. DOI: <pub-id pub-id-type="doi">10.1111/j.1467-9922.2011.00689.x</pub-id></mixed-citation></ref>
<ref id="B62"><label>62</label><mixed-citation publication-type="journal"><string-name><surname>Vihman</surname>, <given-names>M. M.</given-names></string-name>, <string-name><surname>Nakai</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>DePaolis</surname>, <given-names>R. A.</given-names></string-name>, &amp; <string-name><surname>Hall&#233;</surname>, <given-names>P.</given-names></string-name> (<year>2004</year>). <article-title>The role of accentual pattern in early lexical representation</article-title>. <source>Journal of Memory and Language</source>, <volume>50</volume>(<issue>3</issue>), <fpage>336</fpage>&#8211;<lpage>353</lpage>. DOI: <pub-id pub-id-type="doi">10.1016/j.jml.2003.11.004</pub-id></mixed-citation></ref>
<ref id="B63"><label>63</label><mixed-citation publication-type="journal"><string-name><surname>Vitevitch</surname>, <given-names>M. S.</given-names></string-name>, &amp; <string-name><surname>Luce</surname>, <given-names>P. A.</given-names></string-name> (<year>1998</year>). <article-title>When words compete: Levels of processing in spoken word recognition</article-title>. <source>Psychological Science</source>, <fpage>9</fpage>&#8211;<lpage>325</lpage>. DOI: <pub-id pub-id-type="doi">10.1111/1467-9280.00064</pub-id></mixed-citation></ref>
<ref id="B64"><label>64</label><mixed-citation publication-type="journal"><string-name><surname>Vitevitch</surname>, <given-names>M. S.</given-names></string-name>, &amp; <string-name><surname>Luce</surname>, <given-names>P. A.</given-names></string-name> (<year>1999</year>). <article-title>Probabilistic phonotactics and neighborhood activation in spoken word recognition</article-title>. <source>Journal of memory and language</source>, <volume>40</volume>(<issue>3</issue>), <fpage>374</fpage>&#8211;<lpage>408</lpage>. DOI: <pub-id pub-id-type="doi">10.1006/jmla.1998.2618</pub-id></mixed-citation></ref>
<ref id="B65"><label>65</label><mixed-citation publication-type="journal"><string-name><surname>Vitevitch</surname>, <given-names>M. S.</given-names></string-name>, &amp; <string-name><surname>Luce</surname>, <given-names>P. A.</given-names></string-name> (<year>2004</year>). <article-title>A web-based interface to calculate phonotactic probability for words and nonwords in english</article-title>. <source>Behavior Research Methods, Instruments, &amp; Computers</source>, <volume>36</volume>(<issue>3</issue>), <fpage>481</fpage>&#8211;<lpage>487</lpage>. DOI: <pub-id pub-id-type="doi">10.3758/BF03195594</pub-id></mixed-citation></ref>
<ref id="B66"><label>66</label><mixed-citation publication-type="journal"><string-name><surname>Wang</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Seidl</surname>, <given-names>A.</given-names></string-name>, &amp; <string-name><surname>Cristia</surname>, <given-names>A.</given-names></string-name> (<year>2021</year>). <article-title>Infant speech perception and cognitive skills as predictors of later vocabulary</article-title>. <source>Infant Behavior and Development</source>, <volume>62</volume>, <fpage>101524</fpage>. DOI: <pub-id pub-id-type="doi">10.1016/j.infbeh.2020.101524</pub-id></mixed-citation></ref>
<ref id="B67"><label>67</label><mixed-citation publication-type="journal"><string-name><surname>Weber</surname>, <given-names>A.</given-names></string-name>, &amp; <string-name><surname>Cutler</surname>, <given-names>A.</given-names></string-name> (<year>2006</year>). <article-title>First-language phonotactics in second-language listening</article-title>. <source>The Journal of the Acoustical Society of America</source>, <volume>119</volume>(<issue>1</issue>), <fpage>597</fpage>&#8211;<lpage>607</lpage>. DOI: <pub-id pub-id-type="doi">10.1121/1.2141003</pub-id></mixed-citation></ref>
<ref id="B68"><label>68</label><mixed-citation publication-type="journal"><string-name><surname>Yip</surname>, <given-names>M. C.</given-names></string-name> (<year>2020</year>). <article-title>Spoken word recognition of l2 using probabilistic phonotactics in l1: evidence from cantonese-english bilinguals</article-title>. <source>Language Sciences</source>, <volume>80</volume>, <fpage>101287</fpage>. DOI: <pub-id pub-id-type="doi">10.1016/j.langsci.2020.101287</pub-id></mixed-citation></ref>
<ref id="B69"><label>69</label><mixed-citation publication-type="journal"><string-name><surname>Zamuner</surname>, <given-names>T. S.</given-names></string-name> (<year>2009</year>). <article-title>Phonotactic probabilities at the onset of language development: Speech production and word position</article-title>. <source>J Speech Lang Hear Res</source>, <volume>52</volume>(<issue>1</issue>), <fpage>49</fpage>&#8211;<lpage>60</lpage>. DOI: <pub-id pub-id-type="doi">10.1044/1092-4388(2008/07-0138)</pub-id></mixed-citation></ref>
<ref id="B70"><label>70</label><mixed-citation publication-type="journal"><string-name><surname>Ziegler</surname>, <given-names>J. C.</given-names></string-name>, &amp; <string-name><surname>Goswami</surname>, <given-names>U.</given-names></string-name> (<year>2005</year>). <article-title>Reading acquisition, developmental dyslexia, and skilled reading across languages: a psycholinguistic grain size theory</article-title>. <source>Psychological bulletin</source>, <volume>131</volume>(<issue>1</issue>), <fpage>3</fpage>. DOI: <pub-id pub-id-type="doi">10.1037/0033-2909.131.1.3</pub-id></mixed-citation></ref>
</ref-list>
</back>
</article>