1. Introduction

Sequences of vocalic sounds are attested in many languages. Two kinds of vocalic sequences are typically distinguished: diphthongs and hiatuses.1 This distinction is based on syllable affiliation: in diphthongs, two vowel qualities belong to one syllable, while in hiatuses they belong to two different syllables (Sands, 2004). In some languages, there is a phonological contrast between diphthongs and hiatuses; for example, Spanish [pi.ˈe] ‘I chirped’ vs. [ˈpi̯e] ‘foot’ (Hualde & Prieto, 2002, p. 217). For other languages, vocalic sequences are described as diphthongs in some contexts and as hiatuses in others; for example, in Fataluku (Timor-Alor-Pantar, East Timor), /hiʔamoi/ ‘climb’ varies between [hiʔamo.i] and [hiʔamoj] depending on speech rate (Heston, 2015, p. 91). In languages where there is no phonological contrast between hiatuses and diphthongs, the production of vocalic sequences as hiatuses or diphthongs typically depends on a range of factors, such as speech rate (vocalic sequences are more likely to be produced as diphthongs when the speech rate is faster) or the specific vowel qualities involved: vocalic sequences where the first vowel is lower than the second (falling/closing sequences) are more often described as diphthongs than vocalic sequences where the first vowel is higher than the second (rising/opening sequences). Examples of languages where falling sequences are described as diphthongs and rising sequences as hiatuses include Selepet (McElhanon, 1970) and Aimele (Aiton, 2016).

Various acoustic properties are implicated in the distinction between hiatuses and diphthongs. Hiatuses typically have greater duration than diphthongs (Chitoran, 2003; Hualde & Prieto, 2002; Marin, 2014) as well as longer formant trajectories (Limanni, 2014; Smith et al., 2008). For hiatuses, the formant values of the onset and offset typically overlap with those of corresponding monophthongs (Borzone de Manrique, 1979; Limanni, 2014; Smith et al., 2008), while for diphthongs this is not necessarily the case. Differences in the slope of F2 have also been identified; hiatuses typically have a steeper F2 slope, indicating a faster transition between vowel qualities, while diphthongs typically have a gradual F2 slope, indicating a gradual transition between vowel qualities (Aguilar, 1999; Chitoran, 2002).

Diphthongs and hiatuses are also characterised as having different patterns of intensity and F0 movements. Diphthongs have been described as having a single intensity peak in the middle of the sequence, while hiatuses either have an intensity plateau across the two vowels or an intensity peak on the first vowel followed by a decline in intensity across the second vowel (Edwards, 2016). The shape and alignment of F0 contours have also been implicated in the diphthong/hiatus distinction, with Torreira (2007) and Gubian et al. (2015) finding that hiatuses have a steeper F0 increase later in the sequence than diphthongs.

While various acoustic differences between diphthongs and hiatuses have been identified, these differences are not clear-cut, even in languages where there is a phonological contrast between the two. For example, in Spanish, there is some overlap between the durations (Hualde & Prieto, 2002; MacLeod, 2007), formant contours, and F0 contours of diphthongs and hiatuses (Gubian et al., 2015). Variation in their production has been shown to be affected by factors such as word position (vocalic sequences are often longer—i.e., more hiatus-like—in word initial-position [Cronenberg et al., 2024]), stress (stressed vocalic sequences are less likely to undergo reduction to a diphthong than unstressed ones [Alba, 2006; Jenkins, 1999]) and speech rate (i.e., vocalic sequences are shorter at faster speech rates [Alba, 2006; Borzone de Manrique, 1979; Chitoran & Hualde, 2007; Jenkins, 1999]).

Although it is established that the acoustic differences between diphthongs and hiatuses are not clear-cut, previous research has largely involved categorising tokens as instances of either diphthongs or hiatuses (impressionistically or based on orthography and/or stress distribution) before assessing the acoustic characteristics of each category. This research has also primarily been devoted to languages where there is a phonological contrast between diphthongs and hiatuses, in particular Spanish (Alba, 2006; Borzone de Manrique, 1979; Garrido, 2008; Gubian et al., 2015; Hualde & Prieto, 2002; Limanni, 2014; MacLeod, 2007; Torreira, 2007) and Romanian (Chitoran, 2003; Marin, 2014). Exploratory research that does not presuppose a diphthong/hiatus binary is lacking at present.

The present study examines the production of opening sequences /ia/, /ua/, /oa/, and /ea/ in te reo Māori (henceforth Māori), the indigenous Polynesian language of Aotearoa New Zealand. These sequences have been described as hiatuses based on auditory impressions and the observation that they do not operate as single phonological units for the purposes of stress assignment. Using data from the MAONZE corpus,2 we re-examine this classification through acoustic phonetic analysis. From the outset, we do not assume that the sequences are necessarily hiatuses, but develop a bottom-up approach which allows us to explore what kinds of variation are present in the data. To do this, we quantify formant, intensity, and pitch trajectories using functional Principal Components Analysis (fPCA, Gubian et al., 2015, 2019). We then perform what we term a uniting Principal Component Analysis (uPCA), which is a a principal components analysis that brings together the results of other analyses, in this case, the fPCA scores and duration. This is done to identify bundles of significant variation across the acoustic dimensions. Then, to determine what social and linguistic factors may underpin the variation observed, we conduct linear mixed-effects regression analysis on the uPCA scores. Lastly, to investigate what our findings mean for the phonological system of Māori, we examine the relationship between the sequences and their corresponding monophthongs. We also examine the speaker intercepts from the regression models to determine if there is a relationship between how different speakers produce the different sequences, which would provide evidence that variation in the production of the sequences is systematic.

The rest of this paper is structured as follows. Section 2 provides relevant language background. Section 3 introduces the present study, its research questions, and provides an overview of the data and measurements, as well as fPCA and uPCA. Section 4 reports the results. Section 5 discusses their interpretation and implications.

2. Language Background

Te reo Māori is the indigenous language of Aotearoa New Zealand. An Austronesian language, it belongs to the Polynesian branch of the Oceanic subgroup (Harlow, 2007). It has ten consonants, /p, t, k, m, n, ŋ, w, f, ɾ, h/, and five vowels, /i, e, a, o, u/ (Bauer, 1993; Harlow, 2007). Māori has a vowel length contrast that has been analysed phonologically as sequences of two identical vowels (see Harlow, 2007, p. 67, for details). All syllables are open, and consonant clusters are not attested. Onsets are optional. Sequences of three vowels or more are attested word-internally, and even longer sequences are possible across word boundaries, but this paper focuses exclusively on sequences of exactly two unlike vowels that occur within the same word.

All possible sequences of two vowels are permitted in Māori, but within a monomorphemic word, some vowel sequences, such as /uo/, are rare. Vowel sequences are attested in both monomorphemic words (e.g., /kea/ ‘kea, Nestor notabilis’; /ɾua/ ‘two’) and morphologically complex words. In morphologically complex words, vowel sequences arise as the result of affixation (e.g., /kite-a/ ‘see-pass’), compounding (e.g., /ŋutu-awa/ ‘river mouth, estuary, lit. lip-river’) and reduplication (e.g., /ako~ako/ ‘learn, teach:redup’ [Moorfield, 2011]).

Descriptions of Māori typically distinguish two types of vocalic sequences: closing sequences /ai, ae, ao, au, ei, oi, ou/ and opening sequences /ia, ua, oa, ea/. Closing sequences are described as diphthongs, while opening sequences are described as hiatuses (De Lacy, 2004; Harlow, 2007). This description of closing sequences as diphthongs refers to their syllabification as one syllable rather than their segmental status; the general consensus has been that all vocalic sequences in Māori are best analysed as sequences of two vowels (Bauer, 1993; Harlow, 2007) rather than as sequences of vowels and glides or as single complex segments (i.e., as diphthongs which are part of the vowel inventory). The classification of the opening sequences as hiatuses is based on auditory impressions and the observation that they do not operate as single phonological units for the purposes of stress assignment. The closing sequences are described as forming heavy syllables which attract stress, while the opening sequences are treated as two light syllables for the purposes of stress (Harlow, 2007; Schütz, 1985).

The presence of morpheme boundaries has been described as having an effect on the production of the closing vowel sequences in Māori (Bauer, 1993; Harlow, 2007). Bauer (1993, p. 536) describes all closing sequences that straddle a morpheme boundary as hiatuses. Harlow (2007), on the other hand, describes such sequences as more or less likely to be hiatuses, depending on the kind of morphological boundary involved. Although these accounts differ, the primary effect of morpheme boundaries described in the literature is that they cause sequences that are usually diphthongs to be realised as hiatuses. Given that the opening sequences are ostensibly already hiatuses, a morpheme boundary should not affect the production of opening sequences. In this study, however, we do not make a priori assumptions about the syllabification of opening sequences. As a result, we also consider the possibility that morpheme boundaries may affect their production (Section 4.2). More broadly, we did not enter this study with any specific predictions about how the opening sequences behave.

2.1. Sound changes relevant to vowels in Māori

Māori has been in significant contact with New Zealand English for ~200 years. Following colonisation, use of the language declined and, following a dramatic shift to English between the 1950s and 1980s, the language became seriously threatened (Benton, 1991). Since the 1980s, revitalisation efforts have led to a growing number of L1 and L2 speakers who are also fluent speakers of English (Benton & Benton, 2001; Harlow & Barbour, 2013). However, due to a break in the intergenerational transmission of Māori, many contemporary L1 speakers learned from parents and teachers who are themselves L2 speakers of the language (Watson et al., 2016).

Contact between English and Māori has resulted in various sound changes that previous studies have documented using the MAONZE corpus. Harlow et al. (2009) and Watson et al. (2016) demonstrate that Māori is undergoing vowel shifts similar to those attested in New Zealand English, including the raising of /e/ and fronting of /u/. Some of these shifts are also affecting vowel sequences; for example, the second target of /au/ and /ou/ is fronting over time, following the fronting of /u/. At the same time, some vowel sequences are undergoing additional changes that make them more similar to other sequences; for example, /ai/ and /ae/ are becoming closer together in the vowel space (Watson et al., 2016). In addition, the trajectories of the closing sequences /au/, /ou/, /ae/, /ai/ and /ao/ show evidence of having changed over time. For the oldest generation of speakers (Historical elders), the onsets and offsets of these sequences mostly overlap with their associated short vowels, while for younger speakers the formant trajectories have become shorter, and the onsets and offsets do not necessarily overlap with the corresponding monophthongs (see Watson et al., 2016, figure 8; Keegan et al., 2014, figure 2). Importantly, these changes have only been investigated for a subset of closing vowel sequences; no previous studies have investigated the production of the opening sequences (/ia, ea, oa, ua/), which are the focus of this paper.

3. The present study

The present study examines the production of the opening vowel sequences (/ia, ea, oa, ua/) in Māori. It addresses the following research questions:

  1. What variation is there in the production of the opening sequences?

  2. What social and linguistic factors may underpin the variation observed?

  3. What do the findings of (1) and (2) mean for the phonological system?

To do this, we use a novel, bottom-up approach to the acoustic analysis of vocalic sequences, consisting of several steps. The full analysis pipeline is shown in Figure 1.

Figure 1: Analysis pipeline.

The following subsections provide details about the corpus (Section 3.1), data selection and annotation (Section 3.2), and the acoustic measurements, normalisation and filtering (Section 3.3). The final subsection provides an overview of fPCA and uPCA (Section 3.4).

3.1. The MAONZE corpus and participants

This study uses data from the MAONZE (Māori and New Zealand English) corpus, which was compiled to study sound change in Māori and New Zealand English (King, Maclagan et al., 2010; King et al., 2011).

The corpus includes data from three generations of Māori speakers. The oldest generation, Historical elders, were born between ~1871 and 1916 and were mostly recorded in the 1940s in interviews of varying lengths (10–90 minutes), intended for radio broadcast. The next generation is Present day elders, who were born between 1920 and 1944. All are L1 speakers of Māori. For the Present day elders, the data consists of sociolinguistic interviews as well as read word lists and reading passages, most of which were recorded in the 2000s. Data in the MAONZE corpus was manually transcribed and force aligned using the HMM Tool Kit (HTK). See King et al. (2011) for the transcription protocol and Fromont (2019) for details about forced alignment.

The youngest generation of speakers in MAONZE were born between 1969 and 1992. For the present day younger speakers, some are L1 speakers of Māori (referred to here as Young L1), while others acquired the language later and are categorised as L2 speakers (referred to here as Young L2). All speakers are bilingual in Māori and New Zealand English. Like the Present day elders, the Young L1 and Young L2 speaker data consists of informal speech as well as read word lists and reading passages. The Young L1 and Young L2 speakers were all recorded in the 2000s. In total, the data used in this study comes from 61 speakers. An overview of the participants included in our study and the recording years are given in Table 1. For more details about the MAONZE corpus, see King, Maclagan et al. (2010) and King et al. (2011).

Table 1: Overview of speakers analysed.

Speaker category N speakers Birth year Recording year Age at recording
Men Historical elders 11 ~1871–1885 1946–1948 ~62–77
Present elders 10 1925–1939 2001–2006 64–79
Young L1 5 1970–1984 2001–2006 21–35
Young L2 5 1969–1983 2001–2006 21–35
Women Historical elders 7 1881–1916 1938–1950 55–67
Present elders 12 1918–1944 1981–2010 59–87
Young L1 6 1985–1992 2007–2009 17–32
Young L2 5 1975–1983 2007–2009 17–32

3.2. Data selection and annotation

We extracted all instances of /ia/, /ea/, /ua/ and /oa/ from the MAONZE corpus using the nzilbb.labbcat R library (Fromont, 2023). A subset of instances of /ia/, /ea/, /ua/ and /oa/ were then selected for analysis. We excluded all sequences that were part of a longer vowel sequence, either within a word or across a word boundary. For instance, tokens such as /kia oɾa/ ‘be well’ were excluded because the vowel sequence /ia/ is final in the word and the following word is vowel initial. Filtering out sequences of more than two vowels means that we excluded sequences containing long vowels (i.e., sequences of two identical vowels), such as the vowel sequence in pūangi /puuaŋi/ ‘balloon’. Words that contain a given sequence twice, such as /tuaɾua/ ‘repeat, second’ or /hiahia/ ‘desire’, were also excluded.

In order to include information about the position of a given sequence in a word, each sequence was annotated for its mora position, with each short vowel within a word counting as one mora, each long vowel (two identical vowels) within a word counting as two morae, and vowel sequences counting as the sum of the morae of their constituent vowels. Sequences were also annotated for whether they occur finally in the word and whether the vowel sequence straddles a morpheme boundary. Morpheme boundaries were manually coded, consulting the web-based version of Te Aka (Moorfield, 2011), Morfessor (Creutz & Lagus, 2007; Todd et al., 2022; Virpioja et al., 2011) and decomposition ratings from two fluent speakers of Māori (Panther et al., 2024).

The data set also includes monophthongs for each speaker, which were used to normalise the formant trajectories for the sequences data (see Section 3.3) and to compare the onsets and offsets of the sequences with their corresponding monophthongs (see Figure 5 and Figure 15). These were also taken from the MAONZE corpus. Monophthongs were only included if they did not occur adjacent to another vowel.

3.3. Acoustic measurements, normalisation and filtering

This study examines formant, intensity, and pitch trajectories as well as duration. This section outlines how each of these was measured, how the data was filtered, and in the case of the formant data, how the data was normalised. All acoustic measurements were made using Praat (Boersma & Weenink, 2021) as part of the nzilbb.labbcat package (Fromont, 2023) in R (R Core Team, 2021).

3.3.1. Formants

For each sequence, F1 and F2 trajectories were extracted based on 9 equally spaced time points across the trajectory, from .1 to .9 of the way through the sequence. For formants, the formant ceiling was 5500 Hz for women and 5000 Hz for men. The F1 and F2 trajectories were normalised using the speaker’s monophthongs as reference values.

For the monophthongs, a single midpoint measurement was taken for F1 and F2. Observations were excluded if either F1 or F2 was 2.5 standard deviations or greater above or below the grand mean for the relevant vowel and gender. The final monophthong dataset contains 96,139 tokens. The mean number of tokens per monophthong per speaker is 1,576.

Formant measurements were normalised using Lobanov 2.0 (Brand et al., 2021), a method based on the Lobanov method (Lobanov, 1971). For each speaker, this method calculates the mean values of each formant for each monophthong (i.e., mean F1 and mean F2 for /a/, mean F1 and mean F2 for /e/, etc.). It calculates the mean and standard deviation of these vowel-wise means for each formant to derive estimates of the centre and size of the vowel space that are unaffected by imbalances in observation numbers across vowels. Given these estimates, it normalises formant values for each token based on the token’s distance to the centre of the vowel space, relative to the size of the vowel space. This process was used to normalise both the formant trajectory measurements as well as the monophthong midpoint measurements. It is shown in equation (1), where i is an index which denotes a given speaker’s formant value (i.e., F1/F2), µ is the mean, and σ is the standard deviation.

Flobanov2.0,i=(Fraw,iμ(μvowel1,,μvoweln))σ(μvowel1,,μvoweln) (1)

(Brand et al., 2021, p. 8)

The advantage of Lobanov 2.0 over the Lobanov method (i.e., Lobanov 1.0, Lobanov, 1971) is that it is better suited towards imbalanced data sets (such as ours), where there are not equal token counts per vowel or speaker. Lobanov 1.0 is prone to bias towards vowels with larger token counts because it uses the mean and standard deviation of all formant values, while Lobanov 2.0 uses the “mean of the means,” with all vowels weighted equally.

After Lobanov 2.0 normalization, we removed observations that likely represent erroneous measurements. We identified two pervasive error types: large jumps in formant measurements (which are simply articulatorily implausible); and formant trajectories that consistently deviated from similar tokens throughout their length. Both of these error types are clearly visible in aggregate plots of formant trajectories and can be approached using automated filtering methods. We first removed individual formant measurements that made a suspiciously large jump compared to the surrounding data points. Such jumps can be defined as sudden changes in the velocity of the formant movement, which create deviations from an otherwise smooth contour. To create a baseline curve that allowed us to identify such points, we fitted a smooth to each formant curve using a Generalized Additive Model (GAM; Wood 2017). Assuming that the baseline curve is appropriately smooth, large jumps are identified by large deviations between the predicted formant value for a given point and its actual value. How do we ensure that the baseline curve is appropriately smooth? The smoothness of a GAM fitted to a single curve is typically estimated from that curve directly. However, this may result in overly wiggly curves for data with large jumps. Therefore, we forced the fitted curves to have the same degree of smoothness as a summary GAM (technically a quantile GAM for stability; Fasiolo et al., 2021; Tomaschek et al., 2018) fitted to all curves representing the same combination of formant, vowel, gender and speaker group as our target curve. Since this model is fitted to multiple curves, abrupt velocity changes in individual curves cancel each other out, yielding a smoothness estimate that is more representative of true formant movements. Individual GAMs with this smoothness value fitted to empirical formant trajectories therefore allowed us to identify observations that represent extreme changes in velocity. We removed all observations that were more than one unit away from the baseline curve measured on the Lobanov scale (this distance corresponds roughly to the difference between /i/ and /e/ on both F1 and F2). After this filtering step, we removed all individual trajectories that had fewer than four remaining data points.

The second filtering step consisted of removing trajectories that were consistently higher or lower than typical trajectories for a given combination of formant, vowel, gender and speaker type. We first calculated the median values at each time point for all trajectories representing the same formant, vowel, gender and speaker type as the target trajectory. We then calculated the difference between the target trajectory and these median values at each time point and averaged these differences. Finally, we calculated the absolute value of this average. This deviation value can only be high when the target trajectory is consistently higher or lower than a typical trajectory. We removed all trajectories with a deviation higher than 1.5 on the Lobanov scale. These processes resulted in the filtering out of 104 vowel sequence tokens.3

3.3.2. Pitch and intensity

We also extracted pitch and intensity trajectories for each sequence. Initially, 9 time points were extracted for pitch and intensity, but inspection of the data revealed that timepoints .1 and .9 often contained measurement errors, so the pitch and intensity analysis focuses just on timepoints .2 to .8. For pitch measurements, the pitch floor value was set to 30 Hz for men and 60 Hz for women. For intensity, the default Praat settings were used. Tokens with fewer than 3 measurement points for intensity or pitch were removed from the data set. Pitch and intensity trajectories were centered per speaker.

We used the same filtering methods as for formant contours (Section 3.3.1), removing overly large jumps in intensity (with a threshold value of 15 dB) and contours that were consistently higher or lower than typical for similar trajectories (also with a threshold of 15 dB). As a result of this filtering process, 1,064 tokens were removed.4

For pitch, trajectories were first manually filtered to remove measurement errors. After manually inspecting the data, pitch measurements below 70 Hz or above 460 Hz for women and below 60 Hz or above 230 Hz for men were removed, as the majority of pitch measurement errors lay outside of this range. Following manual filtering of the data, we used the method described above to remove observations representing large jumps (with a threshold of 100 Hz). This filtering process resulted in the removal of 1,283 tokens.5

3.3.3. Duration

We also extracted the duration of each sequence. Duration measurements were made by subtracting the end time of the sequence from the start time. Sequences greater than 500ms were filtered out, as a large number of forced alignment errors occurred above this threshold (for tokens greater than 500ms, we found that 15% of the data or more had forced alignment errors). The duration data was also z-scored across all sequences, i.e., the whole data set, not by speaker. The duration data was z-scored after performing the fPCA so that the z-scored values were derived only for tokens included in the final data set inputted into the uPCA. The final data set inputted to the uPCA contained 10,500 vowel sequence tokens.

3.4. fPCA and uPCA overview

For each vowel, we are interested in finding patterns of variation between F1, F2, pitch, and intensity, which are time-varying contours, and duration, which is a single scale measure. To do this, we used two dimensionality reduction techniques: fPCA (see 1e) and uPCA (see 1f).

In fPCA, the input observations consist of sequences of measurements, or functions—in this case, representing trajectories of F1, F2, intensity and pitch. fPCA identifies the main dimensions of variation in these curves by calculating the mean curve and a set of functional principal components (fPCs) that capture how curves vary around this mean curve. Each curve can be represented as a sum of the mean curve, and the fPCs are multiplied by corresponding weights in a way similar to how a weighted sum of sine waves of different frequencies can represent a complex sound wave. In this sum, the weight or score of each fPC indicates how strongly the function displays the characteristics of that fPC. A score of zero indicates that the characteristics of the fPC are not present in the curve, a positive score indicates that the characteristics are present and oriented in the same direction (i.e., the curve increases wherever the fPC increases), and a negative score indicates that the characteristics are present but oriented in the opposite direction (i.e., the curve decreases wherever the fPC increases). The fPCA scores reduce each input trajectory to a set of numbers which locate it along different dimensions of trajectory shape variation.

An fPCA analysis also tells us what proportion of the total variation around the mean curve is explained by each fPC. The first fPC explains the most variance, and subsequent fPCs explain decreasing amounts of variance. We set a cumulative variance threshold of 90%, meaning that we do not set in advance the number of fPCs to be extracted, but rather extract them one at a time until they collectively explain at least 90% of the variance around the mean curve in the original data. Explaining 90% of the variance across all sequences required three fPCs for formant trajectories and one to two fPCs for pitch and intensity (see the Supplementary Materials for details); for consistency, we chose to include the top three fPCs across all trajectory types.

Analyses were performed using the fdapace package in R (Y. Zhou et al., 2022). They were done separately for each of the different trajectory types (F1, F2, pitch, intensity), for each of the vowel sequences (/ia, ea, oa, ua/). This yielded a total of 16 analyses. This means that the functional principal components and their corresponding scores were acquired independently. As a result, the fPCA provides insights into the main dimensions of variation for each trajectory, but no insights into any relationships that might exist across the different trajectories. Then, to examine the relationships between formant, pitch, and intensity trajectories, and their correlations with duration, we conducted a uPCA. As mentioned earlier, this is a Principal Components Analysis which brings together results from other analyses; in this case, the scores for the top three fPCs for each trajectory—which explain more than 90% of variance observed—as well as the scaled duration of the vowel sequence.

Principal Components Analysis is a method for reducing the dimensionality of quantitative multivariate datasets, which is similar to fPCA, except that its input data consists of single measurements rather than sequences of measurements. Our uPCA identifies a set of uniting Principal Components (uPCs), which are bundles of covarying acoustic cues that best explain variation across speakers and vowel sequence tokens. The uPCs can be interpreted by examining their loadings, which reflect the extent to which each acoustic cue contributes to the bundle and the direction in which it covaries with the other cues in the bundle. As in fPCA, each observed token can be decomposed into a weighted sum of uPCs, where the weight or score of the uPC indicates how strongly the token shows the characteristics of the bundle of acoustic cues represented by the uPC. A positive score indicates that the cues are stronger in the token than on average (e.g., longer duration than average and a more steeply declining pitch trajectory than average), while a negative score indicates that the cues are weaker in the token than on average (e.g., shorter duration than average and a less steeply declining pitch trajectory than average, which may mean that pitch actually increases). This analysis was performed using the pca_test() function from the nzilbb.vowels R package (Wilson Black & Brand, 2022).

As an analogy for understanding the roles of fPCA and uPCA in our analysis, consider a hypothetical scenario of analysing the acoustic correlates of the continuum between hypo- and hyper-articulation of vowels. Such an analysis would first characterise each vowel via a small number of formant frequencies and would then construct a measure such as vowel space centrality based on these formant frequencies. When thinking about hypo- and hyper-articulation, it is this construct of centrality that would be of primary concern, not the specific frequencies of each formant of a token. In our analysis, the fPCs are analogous to the formant frequencies from this scenario, and the uPCs are analogous to constructs like centrality: the fPCs are primarily a means to obtain interpretable uPCs.

As this analogy demonstrates, our analysis focuses on the uPCs. We are not primarily interested in the fPCs in and of themselves, but rather the ways in which they work together in the uPCs to define characteristic bundles of covarying acoustic cues. Correspondingly, we use fPCA simply as a tool to represent trajectories with a few characteristic numbers (in the form of fPC scores) for input into the uPCA so that we can think of uPCA results as bundling acoustic characteristics of entire trajectories (e.g., shape). The specific details of what each fPC represents are secondary to the details of what happens when they cluster together in meaningful ways in the uPCA. For details of what is represented by each fPC implicated in the uPCA results, see Appendix A.

4. Results

4.1. RQ1: Variation in the production of the opening sequences

To investigate co-variation in the production of the opening sequences, we employ fPCA and uPCA (see Section 3.4 for details).

The main variation in our data is identified by uPC1; none of the other uPCs were found to identify meaningful variation in our data.6 The estimated null and sampling distribution for index loadings of Principal Component 1 (uPC1) for each sequence are shown in Figure 2. These visualisations are created based on the plot_loadings() function from Wilson Black and Brand (2022). Variables for which the distribution (shown in red) sits safely outside the null distribution (shown in blue) can be interpreted as contributing meaningfully to the uPC. As stated above, the loadings capture the direction in which each variable covaries with the other cues; a positive loading means that as the principal component score increases, that variable tends to increase. On the other hand, a negative loading entails that as the principal component score increases, that variable decreases. For example, if duration has a positive loading, it increases as the principal component score increases. Variables sharing a sign correlate with each other; i.e., they move in the same direction along the principal component.

Figure 2: Index loadings of uPC1: Bootstrapped Sampling and Permutation-Based Null Distributions after 500 iterations. Loading direction is indicated by ‘+’ or ‘or ‘–’.

For all sequences, duration, intensity (fPC2) and F1 (fPC1) are loaded onto uPC1 and correlate together. Even though the fPCAs and the uPCA were conducted separately for each sequence, the emerging patterns of variation are strikingly similar. As a result, we interpret the loadings for all sequences together. This is done by describing what high and low uPC scores mean for each dimension. In the interests of readability, specifics about the fPCA results as they relate to uPC loadings are found in Appendix A. Our description of what each fPC represents can be understood through the effects of varying uPC1 on the corresponding trajectory, which are outlined in the next section.

4.1.1. Interpreting the loadings

Sequences that have low uPC1 scores have greater duration. This can be seen in Figure 3a, which plots, per sequence, the scaled duration of tokens with high (>1) and low (<1) scores for uPC1 (with the mean shown in white and the median shown in black). As we can see in Figure 3b, sequences with low uPC1 scores also tend to have high initial intensity followed by a steep drop. Sequences with low uPC1 scores also have overall higher F1 and more F1 movement, which can be seen in Figure 3c. Note that Figure 3 and the other visualisations in this section are generated using the filtered raw data (i.e., they are not reconstructions of the fPCs). Note that the filtered intensity data are centered by speaker. The formant and intensity trajectories are generated using geom_textsmooth() with Locally Estimated Scatterplot Smoothing (LOESS), with the span value set to 0.5.

Figure 3: Duration, intensity trajectories, and F1 trajectories of tokens with high (>1) and low (<1) uPC1 scores.

On the other hand, sequences that have high uPC1 scores have shorter duration (Figure 3a), relatively flat intensity trajectories (Figure 3b) and relatively low and flat F1 trajectories (Figure 3c).

F2 is only influential in the uPCA for /ia/, as shown in Figure 4. Tokens with low uPC1 scores (shown in blue in Figure 4) tend to have a steep falling F2 trajectory until the midpoint, followed by a steadier F2. On the other hand, tokens with high uPC1 scores tend to have a more gradual falling slope (shown in red in Figure 4). Figure 4 captures the combined effects of fPC1 and fPC3 on F2. Details about the effects of each individual fPC can be found in the Appendix.

Figure 4: F2 trajectories: high and low uPC1 tokens of /ia/.

An overview of what high and low uPC1 score means for each dimension is given in Table 2.

Table 2: uPC1 loadings overview.

variable captures low uPC1 score high score
duration: z-scored longer duration shorter duration
intensity: fPC2 differences in shape of trajectory steep drop stable intensity
F1: fPC1 overall height and shape of trajectory lower overall F1 higher overall F1
F2: fPC1 (/ia/) overall height and steep drop until more gradually
F2: fPC3 (/ia/) shape of transition the midpoint falling

4.1.2. Interpreting the loadings: Movement in the vowel space

So far, we have examined variation in the shape of intensity, F1 and F2 trajectories which are implicated in the uPCA. This gives us an idea of the kinds of intensity and F1 and F2 trajectories associated with low and high uPC1 scores, but not the overall shape of the movement in the vowel space or the relationship between the onset and offsets of the sequences and their corresponding monophthongs. In order to do this, we plot the formant trajectories of tokens with low (blue) and high (red) uPC1 scores, along with the monophthongs in Figure 5. In Figure 5, the width of the trajectory line shows the intensity. The length of the formant trajectory is only indicative of the distance between the onset and offset, and should not be interpreted as indicative of the duration of the sequence.7

Figure 5: Formant trajectories of tokens with low and high uPC1 scores, with monophthongs, by speaker gender. The width of the trajectory line shows intensity.

We can see in Figure 5 that high uPC1 scores are associated with shorter formant trajectories, which do not reach /a/. The trajectories of high uPC1 tokens /ua/ and /oa/ are particularly short. Lower PC1 scores, on the other hand, are associated with longer formant trajectories. For /ia/ and /ea/, the formant trajectories start lower than they do for the high PC1 tokens, and for /ia/, the starting position is not quite inside its corresponding monophthong /i/.

To summarise, sequences with low uPC1 scores are longer in duration, have longer formant trajectories and a higher initial intensity followed by a steep drop. Furthermore, low uPC1 tokens are more likely to end in the /a/ monophthong space, as seen in Figure 5. These are acoustic properties associated with hiatuses (Section 1).

In contrast, sequences with a higher uPC1 score have shorter formant trajectories and a more stable intensity, aside from a slight peak toward the middle of the vowel. These are properties associated with diphthongs (Section 1). Taken together, uPC1 identifies continuum of variation between more hiatus-like and more diphthong-like realisations. It is worth stressing that the results of the uPC cannot be interpreted as identifying tokens of hiatus and diphthongs, respectively, which are phonological categories. Rather, they capture the acoustic variation present in our data set—patterns of variation that can best be described as more hiatus-like and more diphthong-like. It is also worth stressing that the techniques we use do not allow us to determine exactly which part of the hiatus-like to diphthong-like continuum the variation in our data spans. What is clear, however, is that the properties of high- and low-scoring tokens correspond to the acoustic properties of hiatuses and diphthongs observed cross-linguistically, and manual inspection of the data confirms that tokens of one end of the continuum appear to be clear instances of hiatus, while tokens at the other end are not. Figure 6 shows spectrograms of tokens with low and high uPC1 scores. Tokens of /mea/ ‘say, thing’, which received low and high uPC1 scores, are shown in Figure 6a and 6b, and examples of /noa/ ‘merely’, which received low and high uPC1 scores, are shown in 6c and 6d. These spectrograms show the intensity contours in pink.

Figure 6: Spectograms of low and high uPC1 scoring tokens of /mea/ ‘say, thing’ (top) and /noa/ ‘merely’ (bottom). The intensity trajectories of these tokens are shown in pink.

We can see in Figure 6 that low uPC1 tokens have a high falling intensity, while high uPC1 tokens have a very flat intensity. We can also see that the low uPC tokens are much longer. In the low uPC1 token of /mea/, /ea/ is 187ms, and in the low uPC1 token of /noa/, /oa/ is 282ms. On the other hand, high uPC1 tokens are much shorter; the /ea/ is 61ms and the the /oa/ is 78ms.

4.1.3. What about pitch?

Pitch is not implicated in the uPCA for any of the sequences. We know from previous research, however, that pitch differences between hiatuses and diphthongs are attested (see Section 1).

It is possible that pitch is involved in the variation between more hiatus- and diphthong-like realisations of the opening sequences, but the uPCA does not pick it up. This is due to the way the data was processed; namely, that the pitch trajectories are centred within speakers (Section 3.3.2). This means that absolute differences in pitch would not be picked up if they are only attested across speakers, not within speaker. This situation may be plausible if speakers only attested diphthong-like or hiatus-like sequences. This, however, is not the case: all speakers demonstrate variation (see Section 4.3.1). Given this, and that the fPCA (namely fPC1, see Supplementary Materials) does pick up within speaker absolute pitch differences, which are not implicated in the uPCA, we are quite confident that pitch is not making a major contribution, especially compared to the other variables.

4.2. RQ2: Social and linguistic factors underpinning the variation

In the previous section, we see that there are robust patterns of variation between formant trajectories, duration and intensity, which form a clear continuum between more hiatus-like and more diphthong-like variants. Here, we examine what social and linguistic factors may underpin this variation. To do so, we perform linear mixed-effects regression analysis on the uPCA scores, using the lme4 R package (Bates et al., 2015). We perform regression analyses for each vowel sequence separately, with uPC1 as the response and a range of variables as predictors. All predictors discussed here are denoted using small caps.

We examine the effects of speaker category and the interaction of speaker category and speaker gender. Recall from Section 3.1 that there are four categories of speakers in MAONZE, representing three generations of speakers.

Examining effects of speaker category enables us to identify whether change over time has occurred. speaker category is coded using custom contrasts, with normalised weights to ensure interpretable scaling. The following contrasts are defined: a) Present elders vs. Historical elders, b) Young L1 & Young L2 vs. Historical elders & Present elders, and c) Young L2 vs. Young L1 speakers.

Examining the interaction of speaker category and speaker gender allows us to see whether men and women behave differently from each other. speaker gender is a sum-coded binary factor, where F (denoting women) is positive.

We also examine the possible roles of two factors relating to the position of the vocalic sequence in the word: mora position and morpheme boundary. Mora position reflects linear position, which affects the duration and quality of vowels in many languages (e.g., Johnson & Martin, 2001; van Santen, 1992) and has been shown to play a role in dipthongisation, with vowel sequences in Romance languages being less likely to diphthongise in word-initial position than in other positions (Chitoran & Hualde, 2007). Mora position is a simple-coded ternary factor, which indicates the position of the first vowel in the sequence. Mora position 1 (word-initial) serves as reference and mora positions 2 and 3+ as alternative levels; positions 3 and later are collapsed because the vast majority of sequences in the data occur in mora positions 1–3. Morpheme boundary reflects whether the vowel sequence straddles a morpheme boundary (e.g., /patu-a/ ‘hit-pass’), which descriptions of Māori have argued is associated with more hiatus-like productions (Bauer, 1993; Harlow, 2007). Morpheme boundary is a sum-coded binary factor, where the positive direction represents the presence of a morpheme boundary in the middle of the vowel sequence, and the negative direction represents the absence of a morpheme boundary.

Including both morpheme boundary and mora position as predictors is problematic for /ea/ and /ua/, as the vast majority of instances of /ea/ and /ua/ that straddle a morpheme boundary happen to occur in mora position 2. This can be seen in Tables 3 and 4. This confound of mora position 2 and morpheme boundaries in Māori has been noted previously (Krupa, 1966). For /ea/, this confound of morpheme boundary and mora position causes multicollinearity issues, as there are also virtually no tokens in mora position 2 that do not straddle a morpheme boundary.

Table 3: Occurrences of /ea/ across morpheme boundaries, by mora position.

mora position
morpheme boundary 1 2 3 4 5
absent 2234 1 301 3 5
present 3 197 3 9 3

Table 4: Occurrences of /ua/ across morpheme boundaries, by mora position.

mora position
morpheme boundary 1 2 3 4 5
absent 2389 273 1036 40 87
present 0 120 0 8 0

Given this confound, we exclude mora position from our models for /ea/ and /ua/ to examine the possible effects of morpheme boundary, which has been implicated in previous work on Māori. For /oa/ and /ia/, it is feasible to include both mora position and morpheme boundary. These issues are discussed further in Section 4.2.3.

To determine whether the linguistic environment effects have changed over time, we examine the interaction of speaker category and mora position (for the vowel sequences in which mora position was included), as well as the interaction of speaker category and morpheme boundary. Although Māori is described as having stress, we do not include stress as a predictor in these models. This is because stress is generally described as falling on the first vowel of a sequence if it is in word initial position, but it will not attract stress in other word positions; see Harlow (2007, p. 82) and Schütz (1985) for more details.8 As we include annotations for word position, it is redundant to also annotate whether sequences are stressed.

We include several control predictors to ensure that resultant target linguistic and social effects are not simple reflections of phonetic implementation. These include controls for local speech rate, coarticulation, and phrase-final lengthening. To control for local speech rate, we include a numeric predictor, local speech rate, which denotes the inverse of the mean duration of monophthongs within a window of ten seconds on either side of a given sequence. The use of an inverse transformation and averaging across a local window reduces the influence of isolated monophthongs with long durations. The local speech rate predictor is also centered across all data in the models to help with model convergence and interpretability. We also include the interaction of speaker category and local speech rate as a control predictor to control for speech rate differences between different speaker groups, as the recordings in MAONZE for older speakers are more formal than for the younger speakers. To control for coarticulation, we include a predictor, prev c tongue, which is a simple-coded ternary factor indicating the place of articulation of the preceding consonant (coronal, dorsal, or neutral [reference level]).9 Finally, to control for effects of phrase-final lengthening, we include a predictor, final pause, which captures instances where a sequence occurs finally in the word and is followed by a pause of 250ms or more. This cut-off point was set after examining the distribution of pauses in our data set, which were predominantly shorter than 250ms. Final pause is a sum-coded binary factor, where the positive direction represents word-final tokens followed by a 250ms+ pause and the negative direction represents tokens that are not word-final and/or not followed by a 250ms+ pause.

All models include random intercepts for speaker and word. We do not use variable selection to prune non-significant terms from models; instead, we leave all terms in each model because they are all motivated by theoretical considerations. After fitting each model, we check it for multicollinearity, to ensure that inferences are not invalidated by correlations between predictors. To do so, we calculate the general variance inflation factors (GVIFs) for the fixed effects in each model and check that they are all below five. To aid the interpretation of significant interactions, we conduct omnibus testing via model comparisons. This process involves removing a given predictor and all its interactions from the model, and comparing this model with the full model using the anova() function. In cases where a predictor makes a significant contribution to the model’s explanatory value as indicated by a significant omnibus test, we interpret the significant effects of each term relevant to that predictor, based on the way the contrasts have been constructed. For additional clarity, we conduct EMM tests (Lenth, 2024) to examine comparisons that are not directly instantiated by the contrasts; that is, to check if differences between any given pair of levels of a given predictor are significant when averaging across the interacting predictors. Note that in the following discussion, we refer to a predictor as non-significant in an interaction when we find no significant overall effect when we compare the model to a model without that predictor, and there are no significant pairwise comparisons for that predictor at any level of the interacting predictor.

The following subsections report the main findings of the regression analysis. The effects we report hold above and beyond the influence of the control variables (speech rate, final lengthening); see the Appendix for full models. The plots in this section are made using the sjPlot package (Lüdecke, 2024) and show the marginal effects of the specified predictors on uPC1, while holding all other factors constant at their reference level.

4.2.1. Social factors: speaker category and gender effects

For /oa/ there is a significant interaction of speaker gender and speaker category (speaker category|Young L1 & Young L2 vs. Historical & Present elders: speaker gender: β = –0.447, SE = 0.118, t = –3.779 , p < 0.001). There is also a significant interaction of speaker gender and speaker category for /ua/ (speaker category|Young L1 & Young L2 vs. Historical & Present elders:speaker gender: β = –0.203, SE = 0.088, t = –2.315, p = 0.025). The predicted values of uPC1 for /oa/ and /ua/ showing the interaction of speaker category and speaker gender can be seen in Figure 7.

Figure 7: Predicted uPC1 values of speaker category and speaker gender. Points show model estimates, bars show 95% confidence intervals.

We also find that there is a significant main effect of speaker category for /ia/ (speaker category|YoungL1 & Young L2 vs. Historical & Present elders: β = –0.817 , SE = 0.240, t = –3.408 , p < 0.001). This marginal effect is shown in Figure 8.

Figure 8: Predicted uPC1 values by speaker category for /ia/. Points show model estimates, bars show 95% confidence intervals.

Model comparisons indicate that there are significant effects of speaker category for /ia/ (χ2 = 66.28, df = 18, p < 0.001), /oa/ (χ2 = 42.71, df = 18, p < 0.001) and /ua/ (χ2 = 13.23, df = 6, p = 0.04). Overall, the significant effects of category tend to be negative, which would suggest that younger speakers tend to have lower scores than older speakers. To examine this further and to investigate speaker category differences which are not captured by the model terms (Historical elders vs. Young L1, Historical elders vs. Young L2, Present elders vs. Young L1, Present elders vs. Young L2), we conduct EMM tests. For /ia/, an EMM test reveals a significant difference between Present elders and Young L1 speakers (Present elders -Young L1: β = 1.034, SE = 0.233, t(152) = 4.439, p = < 0.001, Present elders -Young L2: β = 0.919, SE = 0.241, t(145) = 3.810, p = 0.001). This indicates that Young L1 speakers and Young L2 speakers have significantly lower uPC1 scores than Present elders for /ia/. For /oa/, conducting an EMM test indicates no significant differences between categories. For /ua/, there is a significant difference between Historical elders and Present elders, which is already captured by the model coefficients. While some category differences are found—especially for /ia/—suggesting that younger speakers tend to have lower scores than older speakers, we do not find clear differences across multiple speaker categories, which would suggest a pattern of change over time.

Given that we do not find clear differences between speaker categories, we suspect that the negative effects of speaker category for /oa/ and /ua/ come from the interaction with gender, the effects of which differ between categories. To investigate gender effects within each speaker category, we conduct an EMM test on the interaction of speaker gender and speaker category for /oa/ and /ua/. For /oa/, we find that men have significantly higher uPC1 scores in the Young L1 (M – F: β = 0.581, SE = 0.265, t(61.7) = 2.194, p = 0.032) and Young L2 categories (M – F: β = 0.701, SE = 0.286, t(67.1) = 2.451, p = 0.017). For /ua/, we find that men have significantly higher uPC1 scores in the Present elders (M – F: β = 0.344, SE = 0.133, z = 2.582, p = 0.01) Young L1 (M – F: β = 0.375, SE = 0.186, z = 2.023, p = 0.043) and Young L2 categories (M – F: β = 0.763, SE = 0.199, z = 3.833, p < 0.001) but not in the other speaker categories. These differences between men and women speakers can be seen in Figure 7a.

For /ia/, we also find significant effects of speaker gender, which is confirmed by conducting a model comparison (χ2 = 11.26, df = 4 , p = 0.024). The marginal effects are shown in Figure 9. To investigate differences between men and women—which cannot be inferred from the coefficients due to the way the contrasts were coded—we conduct an EMM test. This indicates that there is a significant difference between men and women when averaging across all the different interactions, with men having higher uPC1 scores (i.e., more diphthong-like sequences, M-F: β = 0.386, SE = 0.148, t(71.4) = 2.610, p = 0.011.)

Figure 9: Predicted uPC1 values for /ia/ by speaker gender. Points show model estimates, bars show 95% confidence intervals.

To summarise, there is evidence that social factors are playing a role. In particular, we find that for the youngest generation of speakers, women have more hiatus-like sequences than men. Although no robust category effects are found, what is clear is that there is variation amongst all groups of speakers in terms of their uPC1 scores, which can be seen in the marginal effects plots.

We have not discussed social factor effects for /ea/ because no significant effects of speaker gender are found. There is an overall effect of category (χ2 = 40.52, df = 12, p < 0.001), but there is no evidence of differences between the speaker categories when we average across the interacting predictors for pairwise comparisons. This means that the category effects for /ea/ reside in interactions with other predictors (other than speaker gender). This is discussed in Section 4.2.3.

4.2.2. Linguistic factors: word position effects

There are significant interactions of speaker category and mora position for /oa/ (speaker category|Present vs. Historical elders:mora position|2: β = 0.417, SE = 0.197 t = 2.121, p = 0.034 and speaker category|Young L1 & Young L2 vs. Historical & Present elders:mora position|3+: β = 0.685, SE = 0.192, t = 3.561, p = < 0.001). There is also a significant interaction of speaker category and mora position for /ia/ (speaker category|Young L1 & Young L2 vs. Historical & Present elders:mora position|3+: β = 0.309, SE = 0.137, t = 2.258, p = 0.024). The marginal effects are shown in Figure 10.

Figure 10: Predicted uPC1 values of /ia/ and /oa/ by speaker category and mora position. Points show model estimates, bars shown 95% confidence intervals.

For /ia/, a model comparison finds no significant overall effect of mora position2 = 14.37, df = 8, p = 0.072). This suggests the mora position effect for /ia/ lies in the interaction with speaker category.

For /oa/, conducting a model comparison suggests that there is an overall effect of mora position2 = 21.04, df = 8, p = 0.007).

To investigate how strong the effect of mora position may be within each speaker category, we conduct an EMM test on the interaction of mora position and speaker category for /ia/ and /oa/. For /ia/, this reveals one significant difference between mora positions, and only for Young L2 speakers (1 – (3+): β = –0.650, SE = 0.227, t = –2.859, p = 0.013). For /oa/, there are no significant differences between mora positions for any of the speaker categories.

From the model coefficients, we can surmise that for both /ia/ and /oa/, the difference between mora position 1 and mora position 3+ is positive (i.e., uPC1 tends to be greater in mora position 3+ than in mora position 1). For /oa/ we can also surmise that for the Present elders, mora position 2 is associated more with higher uPC1 scores than for Historical elders. The effects between groups are significantly different from each other, but mora position 3+ is not significantly different from mora position 1 or 2 within any generation (with the exception of Young L2 speakers for /ia/).10

4.2.3. Linguistic factors: Morpheme boundary effects

For /ea/, we find a significant interaction between speaker category and morpheme boundary (speaker category|Present vs. Historical elders:morpheme boundary: β = –0.362, SE = 0.155, t = –2.335, p = 0.020 and speaker category|Young L1 &Young L2 vs. Historical & Present elders:morpheme boundary: β = –0.233, SE = 0.102 t = –2.295, p = 0.022). Model comparison confirms that there is an effect of morpheme boundary2 = 15.246, df = 4, p = 0.004). Pairwise comparisons indicate that there is an overall difference between sequences that do straddle a morpheme boundary and those that do not (false-true: β = 0.479, SE = 0.174, t(53.6) = 2.758, p = 0.008), whereby morpheme boundaries are associated with lower uPC scores (hiatus-like realisations). This finding is striking because it is the boundary effect described in the literature for the closing sequences in Māori (Section 2). A further EMM test on the interaction of speaker category and morpheme boundary indicates that morpheme boundaries are associated with significantly lower uPC scores (hiatus-like realisations) for all groups except Historical elders, which can be seen in Figure 11.

Figure 11: Predicted uPC1 values for /ea/, by speaker category and morpheme boundary. Points show model estimates, bars show 95% confidence intervals.

As noted in the previous section, there is a confound of morpheme boundary and mora position, as the vast majority of instances of /ea/ which straddle a morpheme boundary occur in mora position 2 (see Table 3). As a result, this morpheme boundary effect is somewhat difficult to disentangle from possible word position effects. We note, however, that the effect of morpheme boundaries for /ea/ is different from the mora position 2 effects for other sequences: second position is generally associated with more diphthong-like sequences, while for /ea/ morpheme boundaries are associated with more hiatus-like sequences. This lends weight to the interpretation that this is an effect of the morpheme boundary rather than one of mora position.

We also find a significant effect of morpheme boundary for /ua/ (β = –0.377, SE = 0.17, t = –2.218, p = 0.028). Model comparison confirms this (χ2 = 11.133 , df = 4, p = 0.025). An EMM test shows that the presence of a morpheme boundary in the middle of a vowel sequence is associated with significantly lower uPC scores (i.e., more-hiatus like realisations) when averaging across interacting predictors for pairwise comparisons (false-true: β = 0.377, SE = 0.175, t(209) = 2.154, p = 0.032). Figure 12 shows the marginal effects.

Figure 12: Predicted uPC1 values for /ua/ by morpheme boundary. Points show model estimates, bars show 95% confidence intervals.

Again, this morpheme boundary effect is somewhat difficult to disentangle from possible word position effects because the vast majority of tokens of /ua/ across a morpheme boundary are also in mora position 2 (Table 4). As above, the effect of morpheme boundaries for /ua/ is different to the mora position 2 effects for other sequences, which lends support for this being a morpheme boundary effect.

4.2.4. Summary of regression analysis results

The results of the regression models indicate that variation between hiatus- and diphthong-like productions of the opening sequences is associated with several linguistic and socio-linguistic factors. Investigating effects of speaker category and speaker gender reveals that men have more diphthong-like sequences than women, but in most cases, this difference is only found in the Young L1 and Young L2 speaker categories.

The results of the regression analysis also suggest that there are morpheme boundary effects on the production of /ea/ and /ua/, whereby morpheme boundaries are associated with more hiatus-like realisations. For /ea/, this morpheme boundary effect is only found for the Present elders, and Young L1 and Young L2 speakers. This is striking because this is the morpheme boundary effect described in the literature for the closing sequences, not the opening sequences (see Section 2 for details).

The significant effects per sequence are summarised in Table 5. This includes main effects that are involved in interactions.

Table 5: Regression analysis summary.

sequence
predictor /ia/ /ea/ /oa/ /ua/
speaker category
speaker category:speaker gender
speaker gender
speaker category:mora position
mora position
speaker category:morpheme boundary
morpheme boundary

4.3. RQ3: Implications for the phonological system

So far, our analysis examines variation for each vowel sequence separately. More broadly, we are also interested in what our findings mean for the phonological system.

In Section 4.2.1, we see that there has likely always been variation in the production of the sequences, with all generations of speakers showing variation between more hiatus-like and more diphthong-like realisations. Here, we analyse speaker intercepts from the regression models to examine whether there is evidence that variation in the production of the sequences is systematic. We also visualise monophthongs and vowel sequence trajectories to examine the relationship between the sequences and their corresponding monophthongs.

4.3.1. Co-variation between the vowel sequences

No speaker has exclusively high or low uPC1 scoring tokens for any sequence type; all speakers demonstrate variation. However, some speakers tend to have higher uPC scores across all four sequences, whereas others have lower uPC scores across the sequences. This can be seen in Figure 13 which shows the distribution of high (> 1) and low (<1) tokens per speaker. Note that only 55 of the 61 speakers in MAONZE are included in Figure 13 and in the speaker intercept correlations shown in Figure 14. Speakers who did not produce all four vowel sequences are excluded from this part of the analysis.

Figure 13: Bar plot showing the distribution of distribution of high (> 1) low (<1) scoring tokens per speaker and per vowel sequence.

Figure 14: Barplots (diagonal) show speaker intercepts from the regression models by vowel sequence. Scatterplots (right of diagonal) show correlations between speaker intercepts for pairs of sequences. Pearson (r) and Spearman (rs) correlation coefficients, as well as p-values, are shown left of the diagonal.

To examine statistically if there are relationships between how diphthong- or hiatus-like a given speaker’s sequences are, we took the random intercepts from the regression models, which can be used as a measure of how speakers differ from the rest of the population for a given feature, controlling for their gender and speaker category, which are also controlled in the regression models. For these data, the random intercepts capture between-speaker variation in uPC1 that is not explained by the fixed effects (i.e., effects other than gender, mora position, speech rate, etc.). A positive speaker intercept indicates that a speaker’s uPC1 scores for that sequence are consistently higher than the population average, after accounting for the fixed effects, while a negative speaker intercept indicates that a speaker’s uPC1 scores for that sequence are consistently lower than the population average.

Figure 14 shows the correlations between speaker intercepts for the opening sequences. This plot was generated using a slightly adapted version of the pairscor.fnc() in the languageR library (Baayen, 2022). The top-right section of Figure 14 shows scatterplots of the speaker intercepts for each of the different pairs of sequences, with each point representing a single speaker. The bottom left section gives the Pearson (r) and Spearman (rs) correlation coefficients. The cells along the diagonal contain histograms that show the distribution of speaker intercepts for each sequence.

The results in Figure 14 indicate a moderate but statistically significant correlation between the speaker intercepts for all four sequences. This is a positive correlation, indicating that speakers with higher uPC1 scores for one sequence (i.e., more diphthong-like) also tend to have higher uPC1 scores for the other sequences. These results suggest that there is indeed a relationship between speakers’ productions of the different opening sequences.

4.3.2. Changes in the relationship between the sequences and their corresponding monophthongs

As mentioned above, we are also interested in the relationship between the sequences and their corresponding monophthongs. To examine this, we visualise the monophthongs and vowel sequence trajectories for each speaker category by gender. These visualisations are created using Generalized Additive Mixed Models (GAMMs). These models are fitted separately to F1 and F2 trajectories, with the measurement number as the time variable. The predicted curves are allowed to vary as a function of speaker category, speaker gender and their interaction. We also include two separate sets of random smooths by preceding and following phoneme, which were allowed to vary across speaker category and speaker gender. Specific preceding and following phonemes often co-occur because they are in the same lexical item, which would normally result in a high degree of collinearity (or concurvity, in the case of smooths). For this reason, we transform the time dimension for the random smooths in such a way that the influence of the preceding phoneme was strongest at the start and fell to zero by about two-thirds of the way into the vowel sequence and vice versa for the following phoneme (this was achieved by applying a decay function to the time variable). For the ellipses, the level parameter is set to 0.67, which corresponds to about one standard deviation.

The monophthongs of Historical elders and the trajectories of the relevant vowel sequences are shown in Figure 15a. We can see that for the Historical elders, the first and second targets of the sequences overlap with their corresponding monophthongs. For Present elders, the sequences overall look similar to those of the Historical elders. We can see that for the women, /u/ has moved forward slightly in the vowel space, and the first target of /ua/ only just overlaps with /u/. This can be seen in Figure 15b.

Figure 15: F1 and F2 trajectories for the opening sequences with monophthongs, per speaker category and gender.

The monophthongs and sequences of Young L1 speakers can be seen in Figure 15c. For Young L1 speakers, we can see that there is quite a lot of overlap of the two front vowels /i/ and /e/, especially for women. We can also observe that, in comparison with the Historical and Present elders, /u/ has undergone fronting, especially for women. These vowel shifts in Māori have been previously documented and attributed to contact with New Zealand English (Harlow et al., 2009; Maclagan, Watson, Harlow et al., 2009; Watson et al., 2016). The first targets of /ia/ and /ea/ are very close for women, and there is more overlap of the formant trajectories of these sequences than in the Historical or Present elders. This shift parallels the ongoing merger of the near /iə/ and square /eə/ diphthongs in New Zealand English (Gubian et al., 2019; Maclagan & Gordon, 2000). It is worth noting that these figures only show the formant trajectories and do not show the duration or intensity of /ea/ and /ia/. Thus, while the two sequences have become close in terms of their formant trajectories, we do not know to what extent they are similar in terms of other aspects of their production. We can also see in Figure 15c that for women, the first targets of /ua/ do not overlap with /u/.

The monophthong shifts demonstrated by Young L2 speakers are similar to those of Young L1 speakers, which which can be seen in Figure 15d. /e/ has raised and /u/ has fronted, especially for women. We can also see considerable overlap of /ia/ and /ea/ for both women and men. For the women, the first target of /ua/ does not overlap with /u/.

These visualisations show that the monophthongs of Māori speakers have shifted over time, but the sequence targets have not necessarily moved with them. We see an increasing decoupling between the realisation of the monophthongs and the beginning and end of the vowel sequence trajectories. We have also seen that the formant trajectories of /ia/ and /ea/ overlap considerably for younger speakers, which is evidently a contact effect, paralleling the merger of near and square in New Zealand English. Given that shifts to New Zealand English diphthongs have affected the production of the sequences, but shifts to monophthongs have not affected the sequence targets, this supports the interpretation that the sequences are becoming phonologically independent of their putative components.

Another possible explanation for the moving together of /ia/ and /ea/ is that the onset of /ea/ has risen in parallel with /e/. To investigate this possibility, we examine whether there is a correlation between the mean F1 of /e/ for each speaker (representing the average height of their /e/) and their mean uPC1 score for /ea/ (representing where they sit overall with respect to the dimension of variation captured by uPC1). Since uPC1 includes variation in F1 trajectory height, we would expect to see a correlation if there was a systematic relationship between the F1 of /e/ and F1 of the sequence onset. No evidence of a strong correlation between the two is found (see Section 14 of the Supplementary Materials for details). We also note that the per speaker intercepts for /ia/ and /ea/ show the strongest correlation of any of the sequences (see Figure 14); that is, there is a stronger correlation between how speakers produce /ia/ and /ea/ than between any of the other sequences. Given this, we believe that the reduction of the distinction between /ea/ and /ia/ seems to be a distinct change from raising of /e/.

5. Discussion

This study employed a novel bottom-up approach to study variation in the production of Māori opening vowel sequences. To do this, we used fPCA and uPCA to examine patterns of systematic variation between formant, intensity, and pitch trajectories, along with duration. Our approach revealed a continuum of variation between hiatus-like and diphthong-like realisations, which contrasts with how these sequences have been described in the literature—namely, as hiatuses.

Our study revealed that a range of factors, including the presence of morpheme boundaries, speaker category, and gender, affect where tokens sit on this continuum. We also found evidence that the variation between more diphthong-like and hiatus-like realisation is structured; speakers who have more diphthong-like variants for one sequence also tend to have more diphthong-like variants for the others. Finally, there was evidence of sound change in the sequences under investigation.

Overall, these findings show that the opening sequences may be realised phonetically like diphthongs, raising questions about their phonological classification. The systematicity of diphthong-like realisations across opening sequences within speakers, together with the presence of a morpheme boundary effect that has previously been associated only with closing sequences and the parallels with sound change affecting diphthongs in New Zealand English, make these questions even more salient. The finding that, among younger speakers, women tend to have more hiatus-like productions than men raises further questions about the nature and motivation of sound changes affecting opening sequences and the role that gender and other social factors play in them.

We now discuss these questions and other potential implications of our findings.

5.1. The phonological representation of vocalic sequences in Māori

We show that the opening sequences in te reo Māori span a continuum from more hiatus-like to more diphthong-like realisations. Although the existence of a continuum itself does not prove that the most diphthong-like of our sequences are actually diphthongs, they certainly contain a set of characteristics described as associated with diphthongs. Inspection of the most diphthong-like trajectories suggests they are more diphthongal than hiatus-like, and this is confirmed by auditory checking. From a phonetic perspective, we would argue that there is nowhere on this continuum where one can draw a clear line between diphthongs and hiatuses. This is a true continuum. It raises the question: what does the existence of a hiatus-diphthong continuum in the opening sequences imply about the underlying phonology of Māori? This question is particularly relevant with regards to the distribution of stress in Māori.

It is worth noting that the literature on the cross-linguistic patterning of diphthongs distinguishes between underlying phonological diphthongs and surface phonetic diphthongs. The former are sometimes referred to as ‘segmental diphthongs’ (Andersen, 1972), or ‘true diphthongs’ (Rehg, 2007). On the other hand, ‘sequential diphthongs’ (Andersen, 1972), or ‘apparent diphthongs’ (Rehg, 2007), are surface diphthongs, which come from the phonetic coalescence of two underlying phonemes into a single syllable. Māori is not generally described as having true, segmental diphthongs (Bauer, 1993; Harlow, 2007), but as containing five short monophthongs and five long counterparts, which can form sequences. Evidence that both opening and closing sequences are not underlying diphthongs—but a sequence of segments—comes from reduplication. When it suits the template, reduplication will copy just the first mora from a sequence (e.g., /hoata/ – /ho~hoata/ ‘the moon on the third day’, /whai/ – /wha~whai/ ‘fight, quarrel’) (Keegan, 1996), although this process is not necessarily synchronically productive.

In terms of their surface phonetic realisation, Māori vowel sequences are described as either spanning two syllables (hiatus) or coalescing into a single syllable in a phonetic surface diphthong. Many accounts claim only the closing sequences can do the latter, not the opening ones (De Lacy, 2004; Harlow, 2007). There are some distinctions in the behaviour of stress that are attributed to this difference. Opening sequences are described as being stressed on the first or second vowel, depending on the presence or absence of sentence-final intonation, for example, whereas the stress in closing sequences consistently falls in one place (Schütz, 1985, p. 19). Moreover, the lexical stress rules invoke attraction to heavy syllables involving closing sequences but not opening ones (Harlow, 2007).

If the basis of the stress rules is the realisation of the closing sequences as diphthongs, and opening sequences are increasingly phonetically diphthongal, then this raises questions about the phonology itself. If opening sequences can be surface diphthongs, an account of stress which relies on dipthongs attracting stress does not adequately distinguish between the closing and opening sequences in terms of their stress behaviours. If the stress distinction remains as described, the stress rules would need to be reformulated in a way that does not explicitly rest on diphthongs—in particular, the assumption that closing sequences are diphthongs and opening sequences are not. However, it is also possible that the distinction between closing and opening sequences in terms of the distribution of stress may not remain robust. Empirical work on stress as produced by contemporary speakers of Māori is scarce and, as argued by Culhane (2024, p. 1070), “te reo Māori has undergone numerous sound changes due to contact with English and it is likely that the production and distribution of stress has also been impacted by language contact.” Members of our team are actively investigating this topic.

Our findings raise another question: Is it possible that Māori is developing true, segmental diphthongs? From Figure 5, we see that the most diphthong-like of our sequences do not start in one vowel and end in the other. In Figure 15, we also see cases where the monophthongs have undergone sound change and the diphthongs have not. Indeed, work on the closing sequences also shows that they are not all well-aligned with their component vowels (Watson et al., 2016). There is no conclusive evidence that such sequences are separate phonemes, but it does seem fairly likely that they are separate cognitive representations, which would certainly set the scene for potential phonological reinterpretation. Looking at the surface phonetic realisation alone, as we have done, cannot provide solid evidence on this question. This would require future study, such as experimental work investigating what fluent speaker intuitions and behaviours may reveal about the underlying phonology and empirical work into speaker implementation of stress. It would also be fruitful to examine whether metalinguistic judgements about syllabification affect production (like those investigated by Tilsen and Cohn, 2016). Specifically, this would involve examining whether speakers’ judgements about the number of syllables in words are reflected in their production of their opening sequences. Studies on Spanish have shown that speakers’ intuitions about syllabification largely correlate with the duration of their vocalic sequences, with sequences judged to be heterosyllabic a having greater duration (Face & Alvord, 2004; Hualde & Prieto, 2002).

5.2. Within-speaker covariation across the opening sequences

In Figure 14, we show that there is a relationship between a speaker’s production of one opening vowel sequence and their production of all the others. Even when we control for the speaker category, gender, speech rate, and other factors that we know affect the likely production, some speakers are, on the whole, more diphthongal and others more hiatus-like in their production across all opening sequences.

We can interpret these systematic patterns across speakers in three ways. One interpretation is that some factor we have not fully controlled for in our models that affects all of the sequences. We have a speech rate control, for example, but if it is only an approximation of speech rate, perhaps it does not adequately capture the effect of speech rate on production and an additional uncontrolled speech rate effect (or similar) is surfacing across all of the sequences.

A second possible explanation is that within-speaker covariation across the opening sequences is social. Recent work on New Zealand English (Brand et al., 2021; Hurring et al., 2025) shows that sets of vowels can pattern similarly across speakers. For example, some speakers consistently lead change in certain sets of vowels (even controlling for their age and gender), whereas others lag in these changes. Moreover, listeners evaluate realisation of these vowels as socially meaningful. For te reo Māori opening vowel sequences, we have observed different realisations by men and women, and across different speaker groups. Perhaps the degree of diphthongisation of these vowels carries some social meaning, and it is distributed across the set of vowels.

The final possibility is that these associations reflect emergent phonological change, and some speakers have a greater tendency to reanalyse surface phonetics as reflective of an underlying diphthong. These speakers may be most likely to show decoupling, for example, between the realisation of the vowel sequences and the associated monophthongs, and be most likely to produce consistently diphthongal realisations.

Our current analysis does not allow us to disentangle these possibilities. But it is clear that the variation in all of the opening sequences is linked in an interesting way, eliminating the possibility that these are simply a set of completely unrelated sound changes.

5.3. A potential morpheme boundary effect

Our analysis identifies a morpheme boundary effect on the production of /ea/ and /ua/, where these sequences are more hiatus-like when they straddle a morpheme boundary. This finding is striking because it is the morpheme boundary effect described in the literature for the closing sequences, which are otherwise described as being surface diphthongs and having different phonological behaviour from the opening sequences. The presence of a morpheme boundary is thought to prevent the formation of a one syllable sequence, forcing the sequence into hiatus.

The fact that we observe a morpheme effect in opening sequences forces us to accept one of two possibilities about the nature of the morpheme boundary effect. The first is that opening sequences do indeed sometimes form a diphthong and the morpheme boundary can prevent this, just as with the closing sequences. As described above, the phonetic patterns we observe are consistent with the interpretation that the opening sequences can be diphthongs. The second possibility is that the morpheme boundary effect does not simply categorically force a diphthong to a hiatus, but rather it is a more continuous effect, which affects the degree of vowel coalescence, moving the sequence further down the continuum in a hiatus-ward direction. These interpretations are not mutually exclusive.

However we interpret the nature of the effect, these results contribute to a growing body of evidence that morphological structure can impact the phonetic realisation of complex words (e.g., Bell et al., 2021; Plag et al., 2017; Seyfarth et al., 2018), providing tentative further evidence in support of a direct effect of the morphology on phonetic realisation. It is tentative because, as we have noted in Section 4.2.3, this result cannot be regarded as fully conclusive due to the confound of morphological boundaries and mora position 2 in the corpus data. This is an area which certainly warrants further work and will likely require elicited data designed to disentangle these confounds.

5.4. Links with New Zealand English Changes

In this study, we have seen changes to formant trajectories of /ia/ and /ea/ and fronting of /u/, which are likely contact effects from New Zealand English (NZE). Here we discuss whether these contact effects are phonetic transfer or show evidence of a more extensive phonological restructuring.

As previous studies have demonstrated, there is evidence of a significant phonetic transfer in the monophthongs of te reo Māori for younger generations of speakers, which largely align with the monophthongs of NZE (Maclagan, Watson, King et al., 2009; Watson et al., 2016). In particular, “the movements of Māori /e eː/ are highly correlated with the movements of the NZE dress and trap, as are the movements of the Māori /u uː/ with those of NZE goose.” (Watson et al., 2016, p. 203). Changes to the realisations of the closing sequences in Māori due to influence from NZE have also been documented, in particular changes in the production of /ou/, linked to the NZE goat vowel (Watson et al., 2016) and /ai/ to price (Harlow et al., 2009).

This evidence that the realisation of diphthongs in English has affected the production of a sequence in Māori is notable because Māori vowel sequences are not thought to be underlyingly single segments (Section 2). However, the type of influence we see here—that speakers are collapsing /ia/ and /ea/ in a similar way to NZE near/square—entails that there is some kind of category similarity in the mind of the speaker. Similarly, in the case of /ou/ and /ai/, the effects reported by Watson et al. (2016) and Harlow et al. (2009) would suggest a stored representation of /ou/ and /ai/ that is influenced by NZE representations of goat and price. This is already somewhat challenging in the context of an analysis in which Māori does not contain segmental diphthongs, as it entails a possible reanalysis of sequential diphthongs as segmental diphthongs. But our work contains related evidence from the opening sequences, which are not reported to even be surface diphthongs.

The phenomena we cite as evidence that phonological restructuring is also occurring in Māori can be understood as arising from shared representations in bilingual phonological systems, as all speakers of Māori are bilingual and most are late bilinguals. Some degree of sharing or influence of phonological representations across languages by bilinguals has been widely documented (e.g., Jared & Kroll, 2001; Nakayama et al., 2012; Roelofs, 2003; H. Zhou et al., 2010). Sharing of phonological representations is especially expected for L2 speakers, where we can expect some mapping of L2 sounds onto L1 categories, with the phonetic realisations associated with these L1 phonological representations influencing how they are realised in the the L2 (as proposed by models such as the Perceptual Assimilation Model [Best, 1995; Tyler & Best, 2024] and the Speech Learning Model [Flege, 1995; Flege & Bohn, 2021]). Previous studies on contact between two languages—one containing two similar vowel sounds (e.g., Catalan /o/ and /ɔ/) and the other only one (e.g., Spanish /o/)—show that speakers dominant in the language with only one of the vowels tend to collapse the contrast, presumably due to sharing of representations across languages. Spanish-dominant bilinguals merge /o/ and /ɔ/ in Catalan to a Spanish-like /o/ (Simonet, 2011). Similarly, Spanish-dominant Galician-Spanish bilinguals tend to merge Galician /e/ and /ɛ/; Spanish attests only /e/ (Amengual & Chamorro, 2015). In the case of Māori, it is plausible that Māori-English bilinguals dominant in English collapse the contrast between /ia/ and/ea/ in Māori due to sharing of phonological representations with merged NZE near/square. But this would imply that /ia/ and /ea/ are segmental diphthongs, not merely sequential.

If speakers share phonological representations for /ia/ and /ea/ with NZE diphthongs near/square (and also some closing sequences, cf. Watson et al., 2016), this could result in the emergence of a segmental diphthong category more broadly, which also includes /oa/ and /ua/. This is a highly speculative proposal, but cases of sequence to single segment reanalysis have been documented in language contact situations, in particular for languages spoken in the Balkans (Friedman & Joseph, 2025). While we cannot know for sure if such an extreme change has occurred, the evidence nonetheless suggests that for vocalic sequences, contact effects from NZE likely extend beyond phonetic transfer to phonological restructuring.

5.5. The social meaning of variation and motivation for gender differences

The regression analysis on the uPCA scores reveals that, in general, amongst the youngest generation of speakers (born 1969–1992), women tend to have more hiatus-like productions. In particular, women have more hiatus-like productions of /ia/, and for /ua/ and /oa/, this is true for the youngest speakers. It is likely that these gendered patterns reflect some social evaluation of the variation. We note that hiatus-like pronunciation of the opening sequences is generally considered to be “correct”; elders and teachers generally describe the opening sequences as articulated with each vowel pronounced separately (Stoakes et al., 2019). It has been observed that women are often more likely to use forms overtly prescribed as “correct” (Labov, 2001, p. 293), which could be the case here, especially in the current context of high rates of L2 acquisition and an apparent shift toward diphthongization by younger men.

In the case of /oa/, women were actually more diphthongal than men amongst the Historic elders, but then this shifted for the younger speakers. This parallels other shifts observed in te reo Māori. King, Watson et al. (2010) describe the fronting of /u/ as being initially led by women, but when it became salient, women led a change toward a more conservative pronunciation before the change garnered forward momentum again. They claim that the trajectory of Māori sound changes in their study confirms “other diachronic sociolinguistic studies which reveal that women are typically in the vanguard of a social change as it begins, but if the sound change becomes salient and therefore stigmatised, women’s speech therefore becomes more conservative with respect to that feature” (King, Watson et al., 2010, p. 208).

We cannot, of course, say for certain why we observe these gender effects, but the gender effects do point to the variation being partially social. More broadly, our findings raise questions about how variation in the production of the opening sequences is socially interpreted. We note that we do not have any evidence that speakers are aware of the kinds of variation we observe. This would be an interesting area for future research.

6. Conclusion

The present study has examined the production of Māori opening sequences /ia/, /ea/, /oa/ and /ua/. It has shown that they are not uniformly hiatus-like in their production, as previously described in the literature. Instead, there is a continuum of variation between hiatus-like and diphthong-like realisations. Our analysis reveals that all generations of Māori speakers demonstrate variation in the production of the opening sequences. For the youngest generation of speakers, we observe a significant gender difference, with men producing more diphthong-like variants, while women have more hiatus-like variants.

The realisation of opening sequences as phonetic diphthongs raises many questions about their underlying representation and the phonological structure of Māori in general. For example, the Māori stress rules, as described, rest on a clear surface distinction between closing sequences (realised as diphthongs) and opening sequences (as hiatuses). We have shown that the empirical reality is less clear-cut.

In addition to providing a nuanced picture of the variation in the production of the opening sequences in Māori, this paper also makes a significant methodological contribution. Unlike most previous acoustic research on vocalic sequences, the methods used here do not require presupposing categorical distinctions between different kinds of tokens (e.g., hiatuses and diphthongs). Instead, this approach allowed us to explore variation between different cues without pre-supposing what kind of variation is present in the data. Our approach has promising applications for future studies of vocalic sequences, as well as other gradient acoustic phenomena involving both trajectories (formants, intensity, etc.) and static measures (e.g., duration).

Appendix A: Relevant fPCA results

For all sequences, intensity (fPC2) and F1 (fPC1) are loaded onto uPC1 and correlate together (see Section 4). For /ia/, F2 (fPC1 and fPC3) are also loaded onto uPC1. Here we discuss what these fPCs capture in more detail, as well as how they are loaded onto uPC1.

A.1 Intensity fPC2

fPC2 for intensity captures the overall shape of the trajectory. A high score corresponds to a steep drop, while a low fPC2 score indicates a relatively stable intensity across the sequence. This can be seen in Figure A1 which shows the effect of fPC2 when added or subtracted from the mean intensity curve. The mean curve is shown as a thick black line, high fPC scores are indicated with a plus <+> and low fPC scores are indicated with a minus <->. These curves are weighted by the standard deviation of the scores for each fPC, the same method implemented in Gubian et al. (2015) and Gubian et al. (2019).

Figure A1: Effects of fPC2 on intensity.

The negative loading of intensity fPC2 (see Figure 2) entails that as uPC1 decreases, fPC2 for intensity increases. This means that sequences with lower uPC1 scores tend have high fPC2 scores for intensity (a steep drop) and sequence with high uPC1 scores have lower fPC2 scores for intensity (more stable intensity).

A.2 F1 fPC1

fPC1 for F1 captures the height and general shape of the trajectory, with high scores corresponding to higher F1 and more F1 movement, and lower scores to lower F1. This can be seen in Figure A2, which shows effect of fPC1 when added or subtracted from the mean F1 curve.

Figure A2: Effects of fPC1 on F1.

Sequences with lower uPC1 scores tend have higher fPC1 for F1, due to tThe negative loading of F1 fPC1 onto uPC1 (as shown in Figure 2). This entails that as uPC1 decreases, fPC1 scores for F1 increase (higher F1) and as uPC1 increases, fPC1 scores for F2 decrease (lower F1).

A.3 F2 fPC1 and fPC3 for /ia/

fPC1 for F2 captures the overall height and shape of the trajectory. High fPC1 scores correspond to higher F2 and lower scores to lower F2. This can be seen in the left panel of Figure A3.

fPC2 for F2 captures differences in the shape of the transition portion of the sequence, in particular whether it has more of a concave or convex shape. Lower fPC2 corresponds to more of a concave shape and higher fPC2 corresponds to more of a convex shape. This can be seen in the right panel of Figure A3.

Figure A3: Effects of fPC1 and fPC1 on F1.

Both fPC1 and fPC3 capture variation in the F2 trajectory, but they have different loadings (see Figure 2). The positive loading of fPC1 for F2 means that as uPC1 increases, fPC1 also increases. This means that tokens with higher uPC1 scores have higher F2, On the other hand, the negative loading of fPC3 means that as uPC1 increases, fPC3 decreases. This means that for tokens with lower uPC1, the F2 trajectory has more of a convex shape, and for higher uPC1 tokens, the trajectory has more of a concave shape. The combined effects of fPC1 and fPC3 for /ia/ are plotted in Figure 4 which can be found in Section 4.1.

In addition to the fPCs discussed here, there are numerous other fPCs which captured variation in our data but which were not meaningfully implicated in the variation captured by uPC1. The full fPCA results can found in the Supplementary Materials.

Appendix B: Full regression results

Table B1: Regression results: uPC1 /ia/.

Estimate SE df t p
(Intercept) –0.911 0.198 51 –4.603
sp. cat.|Pres.Elder.v.Hist. 0.478 0.266 342 1.799
sp. cat.|L1L2.v.Hist.Pres.Elder –0.817 0.240 448 –3.408 <0.001
sp. cat.|L2.v.L1 0.069 0.395 503 0.175
sp. gender –0.193 0.068 55 –2.828 0.007
mora pos.|2 0.582 0.301 75 1.935
mora pos.|3+ 0.377 0.156 27 2.422 0.023
morph. boundary 0.103 0.117 109 0.881
local speech rate 0.019 0.010 2002 1.938
final pause –0.553 0.059 1979 –9.394 <0.001
prev c tongue|coronal.v.neutral –0.131 0.125 105 –1.046
prev c tongue|dorsal.v.neutral 0.051 0.133 91 0.382
sp. cat.|Pres.Elder.v.Hist.:sp. gender –0.137 0.168 58 –0.817
sp. cat.|L1L2.v.Hist.Pres.Elder:sp. gender –0.125 0.136 55 –0.915
sp. cat.|L2.v.L1:sp. gender –0.412 0.215 53 –1.918
sp. cat.|Pres.Elder.v.Hist.:mora pos.|2 0.742 0.472 1478 1.570
sp. cat.|L1L2.v.Hist.Pres.Elder:mora pos.|2 0.352 0.466 1240 0.755
sp. cat.|L2.v.L1:mora pos.|2 –0.099 0.768 1973 –0.129
sp. cat.|Pres.Elder.v.Hist.:mora pos.|3+ 0.102 0.156 1897 0.654
sp. cat.|L1L2.v.Hist.Pres.Elder:mora pos.|3+ 0.309 0.137 1645 2.258 0.024
sp. cat.|L2.v.L1:mora pos.|3+ 0.237 0.216 1917 1.097
sp. cat.|Pres.Elder.v.Hist.:morph. boundary –0.164 0.198 1342 –0.829
sp. cat.|L1L2.v.Hist.Pres.Elder:morph. boundary 0.217 0.192 1627 1.132
sp. cat.|L2.v.L1:morph. boundary 0.036 0.324 1974 0.110
sp. cat.|Pres.Elder.v.Hist.:local speech rate 0.003 0.023 1912 0.133
sp. cat.|L1L2.v.Hist.Pres.Elder:local speech rate –0.017 0.019 2001 –0.878
sp. cat.|L2.v.L1:local speech rate 0.019 0.032 2012 0.607

Table B2: Regression results: uPC1 /ea/.

Estimate SE df t p
(Intercept –0.991 0.106 112 –9.334
sp. cat.|Pres.Elder.v.Hist. 0.423 0.220 152 1.926
sp. cat.|L1L2.v.Hist.Pres.Elder –0.068 0.163 107 –0.415
sp. cat.|L2.v.L1 0.433 0.241 83 1.796
sp. gender –0.074 0.070 59 –1.051
morph. boundary –0.239 0.078 52 –3.084 0.003
local speech rate 0.034 0.009 2704 3.953 <0.001
final pause –0.665 0.046 2737 –14.533 <0.001
prev c tongue|coronal.v.neutral –0.155 0.129 58 –1.201
prev c tongue|dorsal.v.neutral –0.052 0.198 65 –0.262
sp. cat.|Pres.Elder.v.Hist.:sp. gender –0.258 0.181 72 –1.427
sp. cat.|L1L2.v.Hist.Pres.Elder:sp. gender –0.099 0.140 59 –0.707
sp. cat.|L2.v.L1:sp. gender –0.142 0.214 52 –0.663
sp. cat.|Pres.Elder.v.Hist.:morph. boundary –0.362 0.155 2721 –2.335 0.02
sp. cat.|L1L2.v.Hist.Pres.Elder:morph. boundary –0.233 0.102 2752 –2.295 0.022
sp. cat.|L2.v.L1:morph. boundary –0.208 0.131 2720 –1.593
sp. cat.|Pres.Elder.v.Hist.: local speech rate –0.012 0.025 2555 –0.474
sp. cat.|L1L2.v.Hist.Pres.Elder:local speech rate 0.037 0.017 2700 2.135 0.033
sp. cat.|L2.v.L1: local speech rate 0.001 0.024 2757 0.023

Table B3: Regression results: uPC1 /oa/.

Estimate SE df t p
(Intercept) –0.626 0.195 48 –3.210
sp. cat.|Pres.Elder.v.Hist. –0.035 0.269 535 –0.130
sp. cat.|L1L2.v.Hist.Pres.Elder –0.276 0.234 556 –1.179
sp. cat.|L2.v.L1 0.258 0.363 626 0.710
sp. gender –0.097 0.059 59 –1.639
mora pos.|2 0.303 0.241 17 1.257
mora pos.|3+ 0.175 0.200 34 0.875
morph. boundary –0.173 0.126 44 –1.375
local speech rate 0.036 0.011 1618 3.218 0.001
final pause –0.232 0.058 1671 –3.994 <0.001
prev c tongue|coronal.v.neutral 0.577 0.174 21 3.316 0.003
prev c tongue|dorsal.v.neutral 0.226 0.233 25 0.970
sp. cat.|Pres.Elder.v.Hist.:sp. gender –0.261 0.153 70 –1.699
sp. cat.|L1L2.v.Hist.Pres.Elder:sp. gender –0.447 0.118 60 –3.779 <0.001
sp. cat.|L2.v.L1:sp. gender –0.060 0.180 52 –0.336
sp. cat.|Pres.Elder.v.Hist.:mora pos.|2 0.417 0.197 1689 2.121 0.034
sp. cat.|L1L2.v.Hist.Pres.Elder:mora pos.|2 0.280 0.158 1684 1.772
sp. cat.|L2.v.L1:mora pos.|2 –0.205 0.242 1660 –0.846
sp. cat.|Pres.Elder.v.Hist.:mora pos.|3+ 0.356 0.273 1628 1.301
sp. cat.|L1L2.v.Hist.Pres.Elder:mora pos.|3+ 0.685 0.192 1576 3.561 <0.001
sp. cat.|L2.v.L1:mora pos.|3+ 0.271 0.259 1687 1.044
sp. cat.|Pres.Elder.v.Hist.:morph. boundary –0.185 0.206 1603 –0.896
sp. cat.|L1L2.v.Hist.Pres.Elder:morph. boundary –0.044 0.187 857 –0.235
sp. cat.|L2.v.L1:morph. boundary –0.177 0.292 1616 –0.606
sp. cat.|Pres.Elder.v.Hist.:local speech rate –0.013 0.027 1520 –0.479
sp. cat.|L1L2.v.Hist.Pres.Elder:local speech rate –0.007 0.022 1611 –0.298
sp. cat.|L2.v.L1:local speech rate –0.029 0.035 1643 –0.835

Table B4: Regression results: uPC1 /ua/.

Estimate SE df t p
(Intercept) –0.758 0.113 122 –6.708
sp. cat.|Pres.Elder.v.Hist. 0.483 0.166 298 2.904 0.004
sp. cat.|L1L2.v.Hist.Pres.Elder 0.133 0.162 489 0.818
sp. cat.|L2.v.L1 0.292 0.265 608 1.103
sp. gender –0.183 0.044 52 –4.168 <0.001
morph. boundary –0.189 0.085 145 –2.218 0.028
local speech rate 0.052 0.007 3414 7.554 <0.001
final pause –0.413 0.044 3924 –9.279 <0.001
prev c tongue|coronal.v.neutral 0.338 0.110 46 3.083 0.003
prev c tongue|dorsal.v.neutral 0.344 0.232 34 1.482
sp. cat.|Pres.Elder.v.Hist.:sp. gender –0.182 0.111 62 –1.640
sp. cat.|L1L2.v.Hist.Pres.Elder:sp. gender –0.203 0.088 53 –2.315 0.025
sp. cat.|L2.v.L1:sp. gender –0.194 0.136 47 –1.425
sp. cat.|Pres.Elder.v.Hist.:morph. boundary 0.078 0.139 3216 0.565
sp. cat.|L1L2.v.Hist.Pres.Elder:morph. boundary 0.228 0.143 1347 1.589
sp. cat.|L2.v.L1:morph. boundary 0.253 0.235 3096 1.078
sp. cat.|Pres.Elder.v.Hist.:local speech rate –0.025 0.017 3312 –1.440
sp. cat.|L1L2.v.Hist.Pres.Elder:local speech rate 0.002 0.014 3404 0.135
sp. cat.|L2.v.L1:local speech rate 0.033 0.021 3476 –1.567

Additional files

The full data set and code can be found in the Supplementary Materials, hosted by NZILBB on Github at nzilbb.github.io/OpeningSeqsSupplementaries.

Acknowledgements

This research was funded by a Royal Society of New Zealand Marsden Research Grant awarded to Jennifer Hay, Jeanette King, Forrest Panther, Simon Todd, and Márton Sóskuthy (Proposal number 22-UOC-065/ Grant Number M1254). Additional funding to Márton Sóskuthy was provided by the MTA Distinguished Guest Scientist Fellowship Programme 2025 (Hungarian Academy of Sciences).

The MAONZE corpus was created by Jeanette King, Margaret Maclagan, Ray Harlow, Peter Keegan and Catherine Watson. We would like to thank the MAONZE project team for granting us permission to use it.

Competing Interests

The authors have no competing interests to declare.

Authors’ contributions

This project is funded by a Marsden Grant awarded to Jennifer Hay (PI), and Jeanette King, Forrest Panther, Márton Sóskuthy, and Simon Todd (AI). Peter Keegan is an advisor on that grant and Kirsten Culhane is a postdoctoral fellow.

The project described in this paper originated as a joint class project in a graduate sociophonetics class in which Allie Osborne, Penny Harris, and Kate Maindonald were students. Jennifer Hay taught the class, and Kirsten Culhane participated as a research advisor. That classwork led to a preliminary analysis, including data annotation, creation of visualisations which are in the final paper, and development of the fPCA pipeline, with significant code contributions from Allie Osborne.

Subsequent to the class, Kirsten Culhane oversaw the full analysis pipeline, including new innovations from Márton Sóskuthy, who devised the trajectory filtering processes, and created the trajectory visualisations presented in the final section, and from Simon Todd, who devised the reported regression analysis. All authors except the students attended regular meetings to discuss aspects of the analysis and provide advice.

The final draft was written by Kirsten Culhane, with significant contributions from Simon Todd, Márton Sóskuthy and Jennifer Hay. All authors reviewed the submitted draft. Kirsten Culhane compiled the supplementary materials.

Notes

  1. Throughout this paper, we use the terms vocalic sequences and vowel sequences. Vocalic sequences is used as a general term to refer to sequences of vocalic sounds, irrespective of their phonological status. Vowel sequences refers to the phonological analysis of these phenomena as sequences of two monophthongs. [^]
  2. https://maonze.blogs.auckland.ac.nz/. [^]
  3. We note that after this process, the majority of tokens had all 9 data points (71.1% of F1 trajectories and 83.84% of F2 trajectories). For Historical elders, 99 of the remaining F1 trajectories and 55 of the remaining F2 had 4 or fewer time points. For Present elders, 126 of the remaining F1 trajectories and 43 of the remaining F2 had 4 or fewer time points. For Young L1, 91 of the remaining F1 trajectories and 13 of the remaining F2 had 4 or fewer time points. For Young L2, 42 of the remaining F1 trajectories and 14 of the remaining F2 had 4 or fewer time points. Trajectories with 4 or fewer time points make up 2.88% of F1 and 0.92% of F2 trajectories the data, and as a result we do not anticipate that they have affected the results of the fPCA analysis. [^]
  4. After this process, the majority of tokens had all 7 data points (79.43%). For Historical elders, 54 of the remaining intensity trajectories had 4 time points. For Present elders, 161 of the remaining intensity trajectories had 4 time points. For Young L1, 45 of the remaining intensity trajectories had 4 time points. For Young L2, 44 of the remaining intensity trajectories had 4 time points. Trajectories with fewer than 4 time points make up 2.51% of intensity trajectories in the data. Given this, we do not anticipate that trajectories of 4 time points impacted the results of the fPCA analysis. [^]
  5. Again, after this process, the majority of tokens had all 7 data points (77.7%). For Historical elders, 65 of the remaining pitch trajectories 4 or fewer time points. For Present elders, 168 of the remaining pitch trajectories 4 or fewer time points. For Young L1 speakers, 45 of the remaining pitch trajectories 4 or fewer time points. For Young L2 speakers, 52 of the remaining pitch trajectories 4 or fewer time points. Trajectories with fewer than 4 time points make up 2.78 % of pitch trajectories in the filtered data. [^]
  6. Information about the other uPCs can be found in the Supplementary Materials. [^]
  7. Figure 5 shows the monophthongs of all women and men. It could be possible that speakers with predominantly high uPC1 scores and predominantly low uPC1 scores have different monophthong spaces, and that the different trajectories we see between high and low uPC1 tokens are actually a reflection of these differences. To investigate this, we plot the vowel spaces of men and women with predominantly low uPC1 scores. These plots can be found in the Supplementary Materials. Plotting the data in this way does not change the interpretation of the results, however. Even if we separate the monophthongs of speakers who have predominantly high and low scores, the interpretation remains the same: the formant trajectories of low scoring tokens tend to be longer and overlap with /a/ while for high scoring tokens, the formant trajectories tend to be shorter. [^]
  8. Recall from Section 2 that closing sequences are described as attracting stress when in non-initial position, but opening sequences are not. This claim relates to the phonological status of opening sequences as hiatuses rather than diphthongs, and we will return to discuss implications of our results for its status in the next section. [^]
  9. Finer-grained distinctions were not possible due to data sparsity. [^]
  10. The mora position effect may also be affected by pooling of data across speaker groups (due to the way the factor was coded). [^]

References

Aguilar, L. (1999). Hiatus and diphthong: Acoustic cues and speech situation differences. Speech Communication, 28(1), 57–74.  http://doi.org/10.1016/s0167-6393(99)00003-5

Aiton, G. (2016). A grammar of Eibela: A language of the Western Province, Papua New Guinea [Doctoral dissertation, James Cook University].

Alba, M. (2006). Accounting for variability in the production of Spanish vowel sequences. In N. Sagarra & A. J. Toribio (Eds.), Selected proceedings of the 9th Hispanic Linguistics Symposium (pp. 273–285). Cascadilla Proceedings Project.

Amengual, M., & Chamorro, P. (2015). The effects of language dominance in the perception and production of the Galician mid vowel contrasts. Phonetica, 72(4), 207–236.  http://doi.org/10.1159/000439406

Andersen, H. (1972). Diphthongization. Language, 48(1), 11–50.  http://doi.org/10.2307/412489

Baayen, R. H. (2022). languageR: Data sets and functions with “Analyzing linguistic data: A practical introduction to statistics” [Computer software].  http://doi.org/10.32614/CRAN.package.languageR

Bates, D., Mächler, M., Bolker, B., & Walker, S. (2015). Fitting linear mixed-effects models using lme4. Journal of Statistical Software, 67(1), 1–48.  http://doi.org/10.18637/jss.v067.i01

Bauer, W. (1993). Maori. Routledge.  http://doi.org/10.4324/9780203403723

Bell, M. J., Ben Hedia, S., & Plag, I. (2021). How morphological structure affects phonetic realisation in English compound nouns. Morphology, 31, 87–120.  http://doi.org/10.1007/s11525-020-09346-6

Benton, R. A. (1991). The Māori language: Dying or reviving? (Alumni-in-Residence Working Paper Series. Reprinted by New Zealand Council for Educational Research in 1997). East West Center. https://www.nzcer.org.nz/research/publications/maori-language-dying-or-reviving

Benton, R. A., & Benton, N. (2001). RLS in Aotearoa/New Zealand 1989–1999. In J. Fishman (Ed.), Can threatened languages be saved? (pp. 423–450). Multilingual Matters.  http://doi.org/10.21832/9781853597060-020

Best, C. T. (1995). A direct realist view of cross-language speech perception. In W. Strange (Ed.), Speech perception and linguistic experience: Issues in cross-language research (pp. 171–204). York Press.

Boersma, P., & Weenink, D. (2021). Praat: Doing phonetics by computer (Version 6.1.42) [Computer Software]. praat.org

Borzone de Manrique, A. M. (1979). Acoustic analysis of the Spanish diphthongs. Phonetica, 36(3), 194–206.  http://doi.org/10.1159/000259958

Brand, J., Hay, J., Clark, L., Watson, K., & Sóskuthy, M. (2021). Systematic co-variation of monophthongs across speakers of New Zealand English. Journal of Phonetics, 88, 101096.  http://doi.org/10.1016/j.wocn.2021.101096

Chitoran, I. (2002). A perception-production study of Romanian diphthongs and glide-vowel sequences. Journal of the International Phonetic Association, 32(2), 203–222.  http://doi.org/10.1017/s0025100302001044

Chitoran, I. (2003). Inter-gestural timing between vocalic gestures as a function of syllable position: Acoustic evidence from Romanian. In S. Palethorpe & M. Tabain (Eds.), Proceedings of the 6th International Seminar on Speech Production (pp. 25–30). Macquarie Centre for Cognitive Science.

Chitoran, I., & Hualde, J. I. (2007). From hiatus to diphthong: The evolution of vowel sequences in Romance. Phonology, 24(1), 37–75.  http://doi.org/10.1017/s095267570700111x

Creutz, M., & Lagus, K. (2007). Unsupervised models for morpheme segmentation and morphology learning. ACM Transactions on Speech and Language Processing, 4(1), 3:1–3:34.  http://doi.org/10.1145/1217098.1217101

Cronenberg, J., Chitoran, I., Lamel, L., & Vasilescu, I. (2024). Crosslinguistic comparison of acoustic variation in the vowel sequences /ia/ and /io/ in four Romance languages. In Proceedings of Interspeech 2024 (pp. 3689–3693). International Speech Communication Association (ISCA).  http://doi.org/10.21437/Interspeech.2024-2090

Culhane, K. (2024). Acoustic correlates of word stress in te reo Māori: Historical speakers. Proceedings of Speech Prosody 2024, 1070–1074.  http://doi.org/10.21437/SpeechProsody.2024-216

De Lacy, P. (2004). Maximal words and the Maori passive. In J. J. McCarthy (Ed.), Optimality theory in phonology: A reader (pp. 495–512). Blackwell.  http://doi.org/10.1002/9780470756171.ch27

Edwards, O. (2016). Amarasi. Journal of the International Phonetic Association, 46(1), 113–125.  http://doi.org/10.1017/s0025100315000377

Face, T. L., & Alvord, S. M. (2004). Lexical and acoustic factors in the perception of the Spanish diphthong vs. hiatus contrast. Hispania, 87(3), 553–564.  http://doi.org/10.2307/20063061

Fasiolo, M., Wood, S. N., Zaffran, M., Nedellec, R., & Goude, Y. (2021). Qgam: Bayesian nonparametric quantile regression modeling in R. Journal of Statistical Software, 100(9), 1–31.  http://doi.org/10.18637/jss.v100.i09

Flege, J. E. (1995). Second-language speech learning: Theory, findings, and problems. In W. Strange (Ed.), Speech perception and linguistic experience: Issues in cross-language research (pp. 229–273). York Press.

Flege, J. E., & Bohn, O.-S. (2021). The revised speech learning model (SLM-r). In R. Wayland (Ed.), Second language speech learning: Theoretical and empirical progress (pp. 3–83). Cambridge University Press.  http://doi.org/10.1017/9781108886901.002

Friedman, V. A., & Joseph, B. D. (2025). Phonology. In V. A. Friedman & B. D. Joseph (Eds.), The Balkan languages (pp. 359–494). Cambridge University Press.  http://doi.org/10.1017/9781139019095.007

Fromont, R. (2019). Forced alignment of different language varieties using LaBB-CAT. In S. Calhoun, P. Escudero, M. Tabain, & P. Warren (Eds.), Proceedings of the 19th International Congress of Phonetic Sciences (ICPhS 2019). Australian Speech Science & Technology Association.

Fromont, R. (2023). nzilbb.labbcat: Accessing data stored in ‘LaBB-CAT’ instances (R package version 1.3-0) [Computer software]. https://CRAN.R-project.org/package=nzilbb.labbcat

Garrido, M. (2008). Diphthongization of non-high vowel sequences in Latin American Spanish [Doctoral dissertation, University of Illinois, Urbana-Champaign].

Gubian, M., Harrington, J., Stevens, M., Schiel, F., & Warren, P. (2019). Tracking the New Zealand English NEAR/SQUARE merger using functional principal components analysis. In Proceedings of Interspeech 2019 (pp. 296–300). International Speech Communication Association (ISCA).  http://doi.org/10.26686/wgtn.12319358.v1

Gubian, M., Torreira, F., & Boves, L. (2015). Using functional data analysis for investigating multidimensional dynamic phonetic contrasts. Journal of Phonetics, 49, 16–40.  http://doi.org/10.1016/j.wocn.2014.10.001

Harlow, R. (2007). Maori: A linguistic introduction. Cambridge University Press.  http://doi.org/10.1017/CBO9780511618697

Harlow, R., & Barbour, J. (2013). Maori in the 21st century: Climate change for a minority language? In W. Vandenbussche, E. H. Jahr, & P. Trudgill (Eds.), Language ecology for the 21st century: Linguistic conflicts and social environments (pp. 241–266). Novus Press.

Harlow, R., Keegan, P. J., King, J., Maclagan, M., & Watson, C. I. (2009). The changing sound of the Māori language. In J. N. Stanford & D. R. Preston (Eds.), Variation in indigenous minority languages (pp. 129–152). John Benjamins.  http://doi.org/10.1075/impact.25.07har

Heston, T. (2015). The segmental and suprasegmental phonology of Fataluku [Doctoral dissertation, University of Hawai‘i at Mānoa].

Hualde, J. I., & Prieto, M. (2002). On the diphthong/hiatus contrast in Spanish: Some experimental results. Linguistics, 40, 217–234.  http://doi.org/10.1515/ling.2002.010

Hurring, G., Black, J. W., Hay, J., & Clark, L. (2025). How stable are patterns of covariation across time? Language Variation and Change, 1–25.  http://doi.org/10.1017/s0954394525000043

Jared, D., & Kroll, J. F. (2001). Do bilinguals activate phonological representations in one or both of their languages when naming words? Journal of Memory and Language, 44(1), 2–31.  http://doi.org/10.1006/jmla.2000.2747

Jenkins, D. L. (1999). Hiatus resolution in Spanish: Phonetic aspects and phonological implications from northern New Mexican data [Doctoral dissertation, University of New Mexico].

Johnson, K., & Martin, J. (2001). Acoustic vowel reduction in Creek: Effects of distinctive length and position in the word. Phonetica, 58(1–2), 81–102.  http://doi.org/10.1159/000028489

Keegan, P. J. (1996). Reduplication in Maori [Doctoral dissertation, University of Waikato].

Keegan, P. J., Watson, C. I., Maclagan, M., & King, J. (2014). Sound change in Māori and the formation of the MAONZE project. In A. Onysko, M. Degani, & J. King (Eds.), He hiringa, he pūmanawa: Studies on the Māori language (pp. 33–54). Huia.

King, J., Maclagan, M., Harlow, R., Keegan, P. J., & Watson, C. I. (2010). The MAONZE corpus: Establishing a corpus of Māori speech. New Zealand Studies in Applied Linguistics, 16(2), 1–16.

King, J., Maclagan, M., Harlow, R., Keegan, P. J., & Watson, C. I. (2011). The MAONZE project: Changing uses of an indigenous language database. Corpus Linguistics and Linguistic Theory, 7(1), 37–57.  http://doi.org/10.1515/cllt.2011.003

King, J., Watson, C. I., Maclagan, M., Harlow, R., & Keegan, P. J. (2010). Māori women’s role in sound change. In J. Holmes & M. Marra (Eds.), Femininity, feminism and gendered discourse (pp. 191–211). Cambridge Scholars.

Krupa, V. (1966). Morpheme and word in Maori (Vol. 46). Mouton.

Labov, W. (2001). Principles of linguistic change: Social factors (Vol. 2). Blackwell.

Lenth, R. V. (2024). emmeans: Estimated marginal means, aka least-squares means (R package version 1.10.6) [Computer software]. https://rvlenth.github.io/emmeans/

Limanni, A. (2014). Production and perception of vocalic sequences in Mexican Spanish [Doctoral dissertation, University of Toronto].

Lobanov, B. M. (1971). Classification of Russian vowels spoken by different speakers. Journal of the Acoustical Society of America, 49(2B), 606–608.  http://doi.org/10.1121/1.1912396

Lüdecke, D. (2024). sjplot: Data visualization for statistics in social science (R package version 2.8.16) [Computer software].

Maclagan, M., & Gordon, E. (2000). The NEAR/SQUARE merger in New Zealand English. Asia Pacific Journal of Speech, Language and Hearing, 5(3), 201–207.  http://doi.org/10.1179/136132800805576951

Maclagan, M., Watson, C. I., Harlow, R., King, J., & Keegan, P. J. (2009). /u/ fronting and /t/ aspiration in Māori and New Zealand English. Language Variation and Change, 21(2), 175–192.  http://doi.org/10.1017/s095439450999007x

Maclagan, M., Watson, C. I., King, J., Harlow, R., Thompson, L., & Keegan, P. J. (2009). Investigating changes in the rhythm of Maori over time. In M. Uther, R. Moore, & S. Cox (Eds.), Proceedings of Interspeech 2009, 10th annual conference of the International Speech Communication (pp. 1535–1538). International Speech Communication Association (ISCA).  http://doi.org/10.21437/Interspeech.2009-465

MacLeod, B. (2007). Spanish dialects and variation in vocalic sequences. [Master’s thesis, University of Toronto].

Marin, S. (2014). Romanian diphthongs /ea/ and /oa/: An articulatory comparison with /ja/-/wa/ and with hiatus sequences. Revista de Filología Románica, 31, 83–97.

McElhanon, K. (1970). Selepet phonology. Pacific Linguistics.

Moorfield, J. (2011). Te Aka: Māori-English, English-Māori dictionary (3rd). Pearson. https://maoridictionary.co.nz/

Nakayama, M., Sears, C. R., Hino, Y., & Lupker, S. J. (2012). Cross-script phonological priming for Japanese-English bilinguals: Evidence for integrated phonological representations. Language and Cognitive Processes, 27(10), 1563–1583.  http://doi.org/10.1080/01690965.2011.606669

Panther, F., Mattingley, W., Hay, J., Todd, S., King, J., & Keegan, P. J. (2024). Morphological segmentations of non-Māori speaking New Zealanders match proficient speakers. Bilingualism: Language and Cognition, 27(1), 1–15.  http://doi.org/10.1017/S1366728923000329

Plag, I., Homann, J., & Kunter, G. (2017). Homophony and morphology: The acoustics of word-final S in English. Journal of Linguistics, 53(1), 181–216.  http://doi.org/10.1017/s0022226715000183

R Core Team. (2021). R: A language and environment for statistical computing [Computer software]. R Foundation for Statistical Computing. Vienna, Austria. https://www.R-project.org/

Rehg, K. (2007). Does Hawaiian have diphthongs? And how can you tell? In J. Siegel, J. Lynch, & D. Eades (Eds.), Linguistic indulgence in memory of Terry Crowley (pp. 119–131). John Benjamins.  http://doi.org/10.1075/cll.30.15reh

Roelofs, A. (2003). Shared phonological encoding processes and representations of languages in bilingual speakers. Language and Cognitive Processes, 18(2), 175–204.  http://doi.org/10.1080/01690960143000515

Sands, K. L. (2004). Patternings of vocalic sequences in the world’s languages [Doctoral dissertation, University of California, Santa Barbara].

Schütz, A. J. (1985). Accent and accent units in Māori: The evidence from English borrowings. The Polynesian Society, 94(1), 5–26.

Seyfarth, S., Garellek, M., Gillingham, G., Ackerman, F., & Malouf, R. (2018). Acoustic differences in morphologically-distinct homophones. Language, Cognition and Neuroscience, 33(1), 32–49.  http://doi.org/10.1080/23273798.2017.1359634

Simonet, M. (2011). Production of a Catalan-specific vowel contrast by early Spanish-Catalan bilinguals. Phonetica, 68(1–2), 88–110.  http://doi.org/10.1159/000328847

Smith, J. M., Flores, T. L., & Gradoville, M. S. (2008). An analysis of vowels across word boundaries in Veracruz, Mexican Spanish. IULC Working Papers, 8(1).

Stoakes, H. M., Watson, C. I., Keegan, P. J., Maclagan, M. A., King, J., & Harlow, R. (2019). The dynamics of closing diphthong formant trajectories in te reo Māori. In Proceedings of the 19th International Congress of Phonetic Sciences (pp. 989–993).

Tilsen, S., & Cohn, A. C. (2016). Shared representations underlie metaphonological judgments and speech motor control. Laboratory Phonology, 7(1).  http://doi.org/10.5334/labphon.52

Todd, S., Huang, A., Needle, J., Hay, J., & King, J. (2022). Unsupervised morphological segmentation in a language with reduplication. Proceedings of the 19th SIGMORPHON Workshop on Computational Research in Phonetics, Phonology, and Morphology, 12–22.  http://doi.org/10.18653/v1/2022.sigmorphon-1.2

Tomaschek, F., Tucker, B. V., Fasiolo, M., & Baayen, R. H. (2018). Practice makes perfect: The consequences of lexical proficiency for articulation. Linguistics Vanguard, 4(s2), 20170018.  http://doi.org/10.1515/lingvan-2017-0018

Torreira, F. (2007). Tonal realization of syllabic affiliation in Spanish. In J. Trouvain & W. J. Barry (Eds.), Proceedings of the 16th International Congress of Phonetic Sciences (ICPHS XVI) (pp. 1073–1076). Universität des Saarlandes.

Tyler, M. D., & Best, C. T. (2024). The perceptual assimilation model: Early bilingual adults and developmental foundations. In M. Amengual (Ed.), The Cambridge handbook of bilingual phonetics and phonology (pp. 148–172). Cambridge University Press.  http://doi.org/10.1017/9781009105767.008

van Santen, J. P. H. (1992). Contextual effects on vowel duration. Speech communication, 11(6), 513–546.  http://doi.org/10.1016/0167-6393(92)90027-5

Virpioja, S., Turunen, V. T., Spiegler, S., Kohonen, O., & Kurimo, M. (2011). Empirical comparison of evaluation methods for unsupervised learning of morphology. Traitement Automatique des Langues, 52(2), 45–90.

Watson, C. I., Maclagan, M. A., King, J., Harlow, R., & Keegan, P. J. (2016). Sound change in Māori and the influence of New Zealand English. Journal of the International Phonetic Association, 46(2), 185–218.  http://doi.org/10.1017/s0025100316000025

Wilson Black, J., & Brand, J. (2022). nzilbb.vowels: Functions for vowel covariation studies. https://github.com/JoshuaWilsonBlack/nzilbb_vowels/tree/main

Wood, S. N. (2017). Generalized additive models: An introduction with R (2nd). Chapman; Taylor & Francis.  http://doi.org/10.1201/9781315370279

Zhou, H., Chen, B., Yang, M., & Dunlap, S. (2010). Language nonselective access to phonological representations: Evidence from Chinese–English bilinguals. Quarterly Journal of Experimental Psychology, 63(10), 2051–2066.  http://doi.org/10.1080/17470211003718705

Zhou, Y., Bhattacharjee, S., Carroll, C., Chen, Y., Dai, X., Fan, J., Gajardo, A., Hadjipantelis, P. Z., Han, K., Ji, H., Zhu, C., Müller, H.-G., & Wang, J.-L. (2022). fdapace: Functional data analysis and empirical dynamics (R package version 0.5.9) [Computer software]. https://CRAN.R-project.org/package=fdapace