1. Introduction

Talkers imitate aspects of stimulus talkers’ speech in non-interactive word shadowing tasks when they are explicitly asked to imitate (e.g., Clopper & Dossey, 2020; Dufour & Nguyen, 2013; Schertz et al., 2023), when they are simply asked to repeat the words (Goldinger, 1998; Shockley et al., 2004), and even when they are asked to avoid imitation (Walker & Campbell-Kibler, 2015). The apparent ubiquity of phonetic imitation in word shadowing tasks suggests an automatic process, in which speech perception impacts the immediately following speech production. At the same time, however, the magnitude of imitation varies across instruction conditions, with greater evidence of imitation with instructions to imitate than without (e.g., Clopper & Dossey, 2020; Dufour & Nguyen, 2013; Schertz et al., 2023). The magnitude of imitation also varies considerably as a function of phonological (e.g., Babel et al., 2013; Mitterer & Ernestus, 2008; Nguyen et al., 2012; Nielsen, 2011), phonetic (e.g., Babel, 2010, 2012), and social factors (e.g., Clopper & Dossey, 2020; Ross et al., 2021; Walker & Campbell-Kibler, 2015). Together, these findings suggest that phonetic imitation is at least partially under the control of the talker.

The goal of the current study was to further explore talker control of phonetic imitation. We examined imitation of Australian English dialect features by American English participants in a shadowing task with explicit instructions to imitate, with instructions to avoid imitation, and without instructions to either imitate or not. As in previous work (e.g., Clopper & Dossey, 2020; Dufour & Nguyen, 2013; Schertz et al., 2023), we observed imitation without instructions to imitate or not, as well as with explicit instructions to imitate, with a greater magnitude of imitation in the latter condition. Contrary to previous work (e.g., Walker & Campbell-Kibler, 2015), we did not observe evidence of imitation when shadowers were instructed to avoid imitation. We also observed different degrees of imitation for different variables across dialects, with greater imitation of indexically meaningful variables than less indexically meaningful variables. These findings provide further evidence for the role of talker control and indexical meaning in phonetic imitation.

2. Background

2.1 Imitation and control

Imitation in the word shadowing paradigm, in which participants simply repeat words aloud after hearing them produced by the stimulus talkers, provides strong evidence for the automaticity of imitation through a tight speech perception-production link (Goldinger, 1998; Pickering & Garrod, 2004). Whereas phonetic imitation has also been argued to reflect social goals (e.g., Communication Accommodation Theory; Giles et al., 1987), the non-interactive nature of the shadowing task suggests that at least some of the cognitive mechanisms underlying phonetic imitation bypass participants’ explicit understanding of the speech context. Moreover, some participants imitate model talkers in a shadowing task even when they are explicitly asked to avoid imitation (Walker & Campbell-Kibler, 2015).

These observations are typically taken to support a characterization of lab-based phonetic imitation as automatic, although the precise sense of that term is often left open (Goldinger, 1998). At the same time, imitation is clearly not an across-the-board effect uninfluenced by attitudes or context. The magnitude of phonetic imitation in word shadowing tasks varies as a function of phonetic (e.g., Babel, 2010, 2012), phonological (e.g., Nielsen, 2011), and social (e.g., Clopper & Dossey, 2020; Ross et al., 2021; Walker & Campbell-Kibler, 2015) factors. More imitation is observed when the participants’ baseline productions are further from the model talkers’ productions (e.g., Babel, 2010), when imitation does not impinge on a phonological contrast (e.g., Nielsen, 2011), and when the model talker is attractive or likeable (e.g., Michalsky & Schoormann, 2017). This variation in the magnitude of imitation is typically characterized as linguistic and social selectivity (e.g., Babel, 2010, 2012), implying talker control over imitation to avoid collapsing a phonological contrast (e.g., Nielsen, 2011) or producing a negatively stereotyped linguistic feature (Clopper & Dossey, 2020).

Cognitive control remains a poorly understood phenomenon (De Neys, 2023 and responses; Buehler, 2025) but is often framed as an aspect of cognition which allows us to pursue our goals (e.g., Badre, 2025). A narrower and more operationalized conceptualization is that of deliberative control, which we use here to refer to behavior prompted by explicit instruction or a verbally reported experience of decision-making. The manipulation of explicit instructions has already shown effects in lab-based imitation: Participants typically exhibit more imitation when they are explicitly asked to imitate than when they are simply asked to repeat the model talker (e.g., Clopper & Dossey, 2020; Dufour & Nguyen, 2013; Sato et al., 2013; cf. Clopper et al., 2024).

Dufour and Nguyen (2013) proposed that this difference in magnitude of imitation reflects automatic imitation regardless of instructions, coupled with attention to indexical properties of the model talker under instructions to imitate, leading to greater imitation overall with instructions to imitate relative to without such instructions. Schertz et al. (2023) examined the correlations between by-participant imitation in tasks with and without instructions to imitate and argued for effects of instructions to imitate both in terms of overall amount of imitation (i.e., more imitation with instructions to imitate) and in other characteristics, including greater variability in the magnitude of imitation in explicit tasks as a function of phonetic, phonological, and social constraints. As in the current study, Schertz et al. characterized explicit imitation as involving both automatic processes and “additional, controlled processes” (pp. 3–4) related not only to attention, as proposed by Dufour and Nguyen, but also decisions about which aspects of the model speech to imitate. Thus, although phonetic imitation in shadowing tasks in the absence of instructions to imitate is generally agreed to reflect automatic processes, the observation of quantitative and qualitative changes in the magnitude of imitation in explicit imitation tasks suggests that talkers have some deliberative control over imitation, through attentional and other mechanisms.

In the current study, we combined our own previous approaches to examining the effects of instruction condition on phonetic imitation (Clopper & Dossey, 2020; Clopper et al., 2024; Walker & Campbell-Kibler, 2015) and compared performance in a cross-dialect shadowing task under three different instruction conditions: No Instruction, Avoid Imitate, and Imitate. The goal of this manipulation was to better understand the automatic vs. controlled nature of phonetic imitation in word shadowing and its interaction with the indexical meaning of the features involved.

2.2 Cross-dialect imitation

Participants imitate features of other dialects in shadowing tasks (e.g., Babel, 2012; Clopper & Dossey, 2020; Clopper et al., 2024; Ross et al., 2021; Walker & Campbell-Kibler, 2015), including under instructions to avoid imitation (Walker & Campbell-Kibler, 2015) and under explicit instructions to imitate (Clopper & Dossey, 2020; Clopper et al., 2024). However, the salience of dialect-specific features affects the magnitude of the observed imitation, reflecting the social meaning associated with the variables in the language. In particular, features that are associated with negatively stereotyped or low-prestige varieties are imitated less than more positively regarded or less meaningful features (Babel, 2010; Clopper & Dossey, 2020; Ross et al., 2021). For example, Babel (2010) observed more imitation for Australian dress by New Zealanders, relative to the more socially salient kit. Similarly, Clopper and Dossey (2020) observed more imitation for Southern American English fronted /u o/ by American Midwesterners, relative to the more socially salient /aj/. Likewise, Walker and Campbell-Kibler (2015) found that when Midland-dialect Americans and New Zealanders shadowed each other’s varieties alongside Inland North-dialect Americans and Australians with instructions not to imitate, the most convergence emerged across the largest dialectal gaps with the least indexical meaning. Multiple effects from the order of dialect presentation suggested that contextual social meaning might matter even in highly reduced social contexts.

In the current study, we considered imitation of Australian English by American English participants. Australian English is generally positively viewed by Americans (Garrett et al., 2005), allowing us to consider how the relative salience of high-regard cross-dialect variants are imitated in word shadowing across the instruction conditions.

2.3 The current study

The goal of the current study was to better understand the role of deliberative control in phonetic imitation. To address this question, we manipulated the instructions in a cross-dialect word shadowing task involving an Australian English model talker and American English participants, giving some participants no instructions about imitation and others instructions to either imitate or avoid imitation. We expected to observe more imitation in the Imitate condition than in the No Instruction condition, as in previous work (e.g., Clopper & Dossey, 2020; Dufour & Nguyen, 2013; Sato et al., 2013; cf. Clopper et al., 2024). We also expected to observe imitation in the Avoid Imitate condition, as in previous work (Walker & Campbell-Kibler, 2015), although potentially with a smaller magnitude relative to the No Instruction condition.

Australian English was selected for the model talker because it allowed us to explore imitation of dialect-specific variants that varied in their salience for the participants in a generally high-regard variety. In particular, the variants of interest were vowel quality in the word classes bath, trap, kit, and price, as well as rhoticity in near (Wells, 1982). Some speakers of Australian English, including our model talker, have the bath-trap split, a phonemic split in which bath is produced significantly further back and trap slightly front of the Midland American production of their combined phoneme, while the variety as a whole shows raised and fronted kit, backed price nucleus, and non-rhotic near, relative to American English (Horvath, 2004). For American English participants, the backing of bath (Austen, 2020; Boberg, 1999) and the lack of rhoticity in near are expected to be highly salient (Elliott, 2000; Labov, 1966; Walker, 2019), and so we expected to observe more imitation of these variants than for kit and price, as in previous work (e.g., Honorof et al., 2011; Mitterer & Müsseler, 2013; Podlipský & Šimáčková, 2015; Wade et al., 2023). Making a prediction for trap was more complex, since for the American shadowers it shares a phoneme with the more salient bath. To the extent that imitation is based on the immediately heard token, we would expect similar convergence patterns in trap as in kit. However, if either the implicit or the deliberative aspects of imitation are shaped by phonemic structure, we might expect convergence on trap to be limited by interference from backed bath (see Austen, 2020, for a discussion of outgroup perceptions of the split). If such interference exists and is at least partly driven by participants’ impression of the model talker’s variety, we might predict a trial effect whereby later trials show less convergence than earlier trials, as this impression develops. We do not know of any work examining variation in vowel duration across these two varieties, so we focus our analysis on imitation of vowel quality, as captured by formant frequencies.

3. Methods

3.1 Participants

For the study, 236 participants were recruited from among visitors to an American science museum. We excluded 37 for self-reported speech and/or hearing disorders, 40 for not completing the post-task questionnaire, 6 for starting to speak English after 5 years old and/or growing up outside the U.S., 10 because their questionnaires were subsequently lost, and 7 because of technical problems with the recording.

Data from the remaining 136 participants were included in the analysis: 43 in the No Instruction condition, 42 in Imitate, and 51 in Avoid Imitate. Of these, 68 were men, 64 were women, 1 was nonbinary, and 3 did not report their gender. Participant ethnicity was 115 white, 10 Black, 3 Asian or Pacific Islander, 1 Latino, 4 multiracial, and 3 did not report their ethnicity. Their mean age was 33 years, with a standard deviation of 13 years. Seventeen reported having advanced degrees, with 57 having completed a four-year degree, 43 some amount of college, including an associate’s degree, 14 finished high school, and 5 did not report their education experience.

3.2 Materials

The audio stimuli were taken from Walker and Campbell-Kibler (2015) and were recorded by a college-educated white Australian woman in her 20s from Perth. Eight tokens of (C)CVC(C) words were selected for each of the English vowel classes bath, trap, kit, price, and near. The recordings were made via a head-mounted microphone connected directly to a laptop running SoundForge in a quiet room at the University of Canterbury, New Zealand. The model talker’s productions for bath, trap, and kit at 50% of the way through the vowel and price at 35% of the way through the vowel are shown in Figure 1 (darker labels), along with the female participants’ baseline productions for comparison (lighter labels).

Relative to our American participants, the model talker has a higher and fronter kit, a slightly lower and backer nucleus of price, a backer bath, and slightly fronter trap. Not shown in Figure 1, her tokens of near were largely non-rhotic, with a mean F3 of 3043 Hz at the 80% point, whereas the mean baseline of the female participants was 2048 Hz. We assumed, following Clopper et al. (2024), that the American participants would converge to the relative positions of the model talker’s vowels in the vowel space and not to her absolute formant frequency values.

Figure 1: Model talker (darker labels) and female participants’ mean baseline productions (lighter labels). Gray trap label indicates mean of participants’ combined trap/bath phoneme.

3.3 Procedure

Participants were seated at a computer, facing away from the glass wall of the lab. They were given a set of over-ear headphones with an integrated microphone (Audio-Technica ATH-770COM). Once they had consented and were seated, Praat was opened and set to record from their microphone in a single file spanning the entire experiment. The stimulus presentation was managed in E-Prime.

In the first block, each word was presented as text in the center of the screen with no accompanying audio. Participants were asked to say each word aloud as it appeared. A single instance of every word represented in the audio stimuli was presented, along with five tokens each from the fleece and “pool” (goose with following /l/) classes. Words were presented in random order.

After a self-timed break, the second block presented the same words, without the fleece or “pool” words, again in random order. In this block, the model talker’s production of the word was played at the same time as the text was presented. Participants were asked to repeat the word after they heard it. The three conditions of the experiment differed only in the instructions given regarding this second block:

No Instruction: Participants were asked to say the word after they heard the model talker say it, with no further instructions.

Imitate: Participants were asked to say the word after they heard the model talker say it and, additionally, to “try to imitate the speaker as closely as possible.”

Avoid Imitate: Participants were asked to say the word after they heard the model talker say it and, additionally, to “try to say the word the way you would normally say it, not the way she says it.”

After the second block, participants were given a paper questionnaire to complete on their own. The first page asked them to give three words or phrases that described the person they just heard talking and then rate that person on four-point scales for educated, intelligent, articulate, friendly, trustworthy, and pleasant-sounding, in that order. They were also asked to name anything specific they noticed about how she said the words. On the second page, they were asked to guess where the talker was from and also answered those same attitudinal questions about the people from the place that were asked about the talker. The analysis discussed here focused only on the open-ended final questions regarding whether they noticed anything specific about the speech of the talker/the place. These responses were coded for mentions of the two most commonly named features, bath and near.

Specifically, we included any mention of “a” sounds as a mention of the bath class. Although logically this could also refer to trap, we felt that both the greater phonetic distance to the model talker and the broad awareness of bath (Austen, 2020; Boberg, 1999; Walker, 2019) supported this inclusion. We do not suggest that trap is explicitly excluded from these comments, but concentrate on bath as the clearer example.

Mentions coded as bath included:

  • rponounces (sic) “a” like “aww”

  • person did not emphasize “a” in words that I would such as clasp. Person would pronounce words with a short “a” sound

  • different “a” pronunciation from American English

Comments mentioning near included:

  • she didn’t pronounce r’s

  • Her r’s were not hard American R’s

  • “R”s changed to “A” at end of words

Comments coded as mentioning neither included:

  • nothing unusual besides accent

  • a lot of emphasis on vowels

  • very crisp, neat, proper

We found that 21 (15%, 4 in No Instruction, 8 in Imitate, 9 in Avoid Imitate) and 22 (16%, 3 in No Instruction, 10 in Imitate, 9 in Avoid Imitate) of participants mentioned bath and near, respectively, in at least one of the talker- or place-focused open-ended questions after the experiment. Three people, 2 in Imitate and 1 in Avoid Imitate, mentioned both features. The only other specific speech sound mentioned by more than one person was “s,” which was mentioned twice.

On the last two pages of the questionnaire, participants were asked for their age, highest level of education, ethnicity and what language they spoke, a regional history, and to report any speech or hearing difficulties.

3.4 Acoustic analysis

The recording for each participant was trimmed to include only the word lists and phonetically aligned with the words taken from the E-Prime record of the session using the Penn Forced Aligner (Yuan & Liberman, 2008). A Praat script automatically extracted the relevant formant values and then presented the proposed values and TextGrid alignment for each token to a research assistant who hand-corrected both the alignment and the formant estimates as necessary.

The dependent measure for bath, trap, kit, and price was the difference between the first and second formants. Formants for bath, trap, and kit were estimated at the midpoint. Since we were interested in the nucleus of price, formants for this vowel were estimated at the 35% timepoint. For near, in which rhoticity was the crucial feature, our dependent measure was the difference between the second and third formants, estimated at the 80% point of the combined span of the vowel and rhotic segment. These difference metrics were selected to minimize the effect of vocal tract differences across talkers and to allow for direct comparison across all of the variables.

3.5 Statistical analysis

Two tokens were excluded as misread, leaving 10878 tokens for analysis. A single model was fit to all the data, including baseline and shadowed tokens together. The crucial effect of imitation was tested with a three-way interaction between block (Baseline, Shadow), condition (No Instruction, Imitate, Avoid Imitate), and phone (bath, trap, kit, price, near). The contrasts for both block and condition were set as treatment, with reference levels of the Baseline block and the No Instruction condition, respectively. Phone was set to sum contrasts. The maximal random effects included intercepts for participant and item, a slope for the interaction between block and phone over participant, and a slope for the interaction between block and condition over item. The random effect structure was reduced as needed to allow for model convergence and to eliminate singular fit errors. The final model that converged included random intercepts for participant and item, random slopes for block over participant and item, and a random slope for phone over participant. To assess statistical significance of the main effects and interactions, degrees of freedom were estimated for F-tests using the Satterthwaite approximation in the lmerTest package in R (Kuznetsova et al., 2017). To assess statistical significance of pairwise comparisons, the estimated marginal means for each model were calculated using the emmeans package in R (Lenth, 2024). The data and R code for all analyses are available as supplementary material to this article and on the Open Science Framework (https://osf.io/4tqgh/).

4. Results

4.1 Effects of instruction condition on imitation

Figure 2 shows the participant means and grand means for the key formant difference in each variable, condition, and block. There was a strong effect of block in the Imitate condition for every variable, but little change across blocks in the other conditions. That is, the results indicate that participants converged when asked to imitate and refrained from convergence when asked to avoid imitation, even diverging from the model talker on kit when asked to avoid imitation. When given no instructions in either direction, they converged only for near.

Figure 2: Formant difference for each variable across condition and block. Dots are participant means and whiskers are standard error. Dashed lines show model talker means.

Table 1 shows the results of the main model, confirming that the three-way interaction between block, condition, and phone was significant. The emmeans analysis, summarized in Table 2, indicates a significant difference between the Shadow and Baseline blocks for all variables in the Imitate condition, confirming convergence, for near in the No Instruction condition, also consistent with convergence, and for kit in the Avoid Imitate condition, confirming divergence.

Table 1: F-test results of the main model. Random effects were intercepts for participant and item, slopes for block by participant and item, and a slope for phone by participant. Formula: FormantDiff ~ Phone * Block * Condition + (1 + Phone + Block | SoundFile) + (1 + Block | Word). Conditional R2 = 0.92, Marginal R2 = 0.78.

Predictor Sum Sq Mean Sq NumDF DenDF F value p
Phone 9,983,757 2,495,939 4 42.36 168.79 < .001
Block 329,737 329,738 1 79.65 22.30 < .001
Condition 18,220 9,110 2 133.00 0.62 0.542
Phone×Block 5,282,043 1,320,511 4 35.11 89.30 < .001
Phone×Condition 3,168,238 396,030 8 133.01 26.78 < .001
Block×Condition 1,228,371 614,185 2 133.01 41.53 < .001
Phone×Block× Condition 37,678,454 4,709,807 8 9,980.10 318.50 < .001

Table 2: Estimated marginal means for main model: shadow – baseline. Negative values indicate convergence for bath and price, positive values indicate convergence for kit, near, and trap. Bonferroni corrected α value is set to 0.003.

Phone Condition Estimate SE df t p
bath No Instruction –18.09 15.54 112.27 –1.16 0.247
bath Avoid Imitate 10.01 14.82 94.91 0.68 0.501
bath Imitate –227.08 15.65 114.97 –14.51 < .001
kit No Instruction –6.58 15.54 112.27 –0.42 0.673
kit Avoid Imitate –45.54 14.82 94.91 –3.07 0.003
kit Imitate 213.00 15.65 114.97 13.61 < .001
near No Instruction 66.53 15.54 112.27 4.28 < .001
near Avoid Imitate 10.15 14.82 94.91 0.68 0.495
near Imitate 522.65 15.65 114.97 33.41 < .001
price No Instruction –27.56 15.54 112.40 –1.77 0.079
price Avoid Imitate 13.42 14.82 94.91 0.91 0.368
price Imitate –114.27 15.65 114.97 –7.30 < .001
trap No Instruction 1.40 15.54 112.40 0.09 0.928
trap Avoid Imitate 0.07 14.82 94.91 0.01 0.996
trap Imitate –68.39 15.65 114.97 –4.37 < .001

In addition, a curious pattern emerged in the individual participant data for near, visible in Figure 2. The distribution of the F3-F2 values in the Shadow tokens for the Imitate condition was elongated relative to the other blocks and conditions, a pattern not seen for any of the other variables. Examination of the individual participant patterns revealed that this distribution is a result primarily of cross-participant differences in the Imitate condition for the near data. Rhoticity showed somewhat less cross-participant variation in the Baseline condition, relative to the vocalic variables, but in the Imitate condition it showed a great deal of variation. It appears that when instructed to imitate non-rhotic tokens, individual participants made different choices regarding how much non-rhoticity to imitate, more so than in the vocalic data.

At the suggestion of an anonymous reviewer, we additionally tested gender as a mediator of the convergence effects, excluding the four participants who were nonbinary or declined to state their gender. The model discussed above was fit with the addition of gender in interaction with the existing three-way phone x block x condition interaction and an additional random slope of gender by item. The overall pattern of results involving phone, block, and condition was the same as in the main model, confirming that the primary results were not driven by the four participants who were nonbinary or declined to state their gender. In addition, the resulting four-way interaction was significant, and post hoc emmeans analysis (provided in the supplementary materials) suggested relatively more convergence among women than men. In the Imitate condition, men and women both showed significant convergence, with larger estimates for women, across all variables except trap, for which women converged significantly but men did not. In the Avoid Imitate condition, the only significant effects were for kit, for which men and women both diverged, with women showing a larger effect. In the No Instruction condition, women showed the overall effect of convergence for near, while the same effect for the men was marginal. Women also showed a convergence effect for price, while men did not.

4.2 Exploratory analysis of social salience

The two most strongly imitated features in the Imitate condition were bath and near, while near was the only feature imitated in the No Instruction condition. These two variables were also the only two repeatedly referenced sounds in the post-experiment questionnaire, consistent with their being well-known indexically meaningful variables for American English speakers (Austen, 2020; Boberg, 1999; Elliott, 2000; Labov, 1966; Walker, 2019).

Based on the match between the social observations and the convergence behavior of these two features, an exploratory post hoc analysis was conducted, testing the participants’ comments on each of the two features as a predictor of their convergence behavior. A new model was fit to the data for bath and near only, taking the main model and adding commentary (whether the participant had commented or not on that feature) as a predictor in interaction with the existing three-way interaction. An additional random slope for the phone * block interaction across participant was also added. The four-way interaction between phone, condition, block, and commentary was not significant, but the three-way interaction between condition, block, and commentary was (F(2, 172.8) = 3.17; p = 0.044). Post hoc emmeans analysis showed the same effects as the main model for the Imitate condition, with larger effects for those who commented on the phone in question (No Comment: est. = 358; t(155) = 16.54, p < 0.001; Comment: est. = 434; t(259) = 13.60, p < 0.001). In the No Instruction condition, those who commented on the relevant variable showed a significant convergence effect (est. = 96; t(224) = 2.19, p = 0.030), while those who did not had only a marginal effect (est. = 38; t(140) = 1.81, p = 0.072). In the Avoid Imitate condition, no convergence was seen regardless of commentary.

One potential concern related to social salience is that instructions to imitate or avoid imitation might lead to increased attention to the features under study, thus changing the imitation behavior. A chi-square test showed no significant effect of condition on noticing either bath or near. However, the descriptive patterns (bath: 9% in No Instruction, 19% in Imitate, 18% in Avoid; near: 7% in No Instruction, 24% in Imitate, 18% in Avoid) suggest that this concern may not be unfounded.

Finally, we investigated whether imitation of trap was affected by sharing a phoneme with the more indexically meaningful bath in the linguistic systems of the participants, though not in that of the model talker. We hypothesized that the influence of each individual item would be stable across trials, while a grammatical model of the talker’s variety and/or identification of the talker as belonging to a familiar language variety would develop over time. Accordingly, another exploratory post hoc analysis examined the effect of time in the recording on the block effect for trap in the Imitate condition, testing an interaction between block and time in the recording. Random effects were intercepts and slopes for block by both participant and item, and a slope for time across participant. The block x time interaction showed a significant effect (est. = –0.308; t(567.3) = –2.11, p = 0.036). While participants did shift their tokens as instructed, converging toward the model talker’s trap, this effect was attenuated over time such that early tokens showed a roughly 100 Hz difference across blocks and the final tokens showed a difference of around 20 Hz. Examination of individual participant plots indicated a great deal of variability in slopes, but all with gradual effects, with no indication of a sudden change following a single exposure to the talker’s bath production. To confirm the specificity of the phenomenon, the same analysis was conducted on bath and kit, which found no significant effect of time. While tentative, this analysis suggests that future work into imitation and phonemic splits or mergers (Babel & McAuliffe, 2013; Lin et al., 2021) might be a fruitful place to explore issues of control.

5. Discussion

In this study, we asked how cross-dialect imitation changed as a result of instructions to imitate or avoid imitation and whether those patterns differed across dialect features. We found that participants converged toward the model talker when instructed to do so for all dialect features and refrained from convergence when instructed to do so for all dialect features, diverging in one case. Convergence in the absence of instructions to imitate or not and the amount of convergence in the Imitate condition both differed across dialect features. These differences suggest that features with established social meaning within the community are more subject to imitation, both with and without instructions to imitate.

5.1 Imitation and control

In the Imitate condition, the participants converged to all five dialect features, confirming that participants can and will imitate features of a high-regard dialect when asked to do so. This finding complements previous work demonstrating explicit imitation of some dialect features in non-prestigious dialects (Clopper & Dossey, 2020; Dufour & Nguyen, 2013). Participants also imitated non-rhoticity in near in the No Instruction condition, confirming that participants can and will imitate some dialect features when they are not asked to do so (Babel, 2010, 2012; Clopper & Dossey, 2020; Walker & Campbell-Kibler, 2015). The differences in convergence across the Imitate and No Instruction conditions are also consistent with previous work (Clopper & Dossey, 2020; Dufour & Nguyen, 2013; Sato et al., 2013), demonstrating both imitation of more dialect features in the Imitate condition than the No Instruction condition and a larger magnitude of convergence in the Imitate condition than the No Instruction condition for near. That is, explicit instructions to imitate led to greater evidence of convergence than no instructions about imitation.

In the Avoid Imitate condition, the participants did not converge to any of the five dialect features, in contrast to previous findings by Walker and Campbell-Kibler (2015). The differences between our findings and theirs are striking, as the two studies drew on virtually the same population and used the same model talker. Aside from minor differences in procedure and experimenters, there were two main differences between their study and our Avoid Imitate condition. First, their participants were told they were in a dialect identification study and instructed to observe differences between their own speech and the model talkers’. Second, the Australian talker was the only (and therefore first) model talker in our study and never first in theirs, appearing either after a New Zealand talker or two different American talkers. Order of presentation emerged as a mediator in their study, but differently across variables. It is difficult to pinpoint a cause for the apparent conflict in results, but one possible explanation is that the dialect identification instructions, supported by the exposure to multiple varieties, altered the attention participants paid to specific features, prompting a change in automatic convergence. Another possibility is that participants in the earlier study prioritized the identification task over the instructions not to imitate and used the shadowing task to explore and reflect on the dialect features via imitation, despite being asked not to.

Together, the results of the current study suggest that participants have considerable control over convergence to dialect features. When participants are asked to imitate, they do, and when they are asked not to imitate, they refrain. When they are not given instructions either way, they imitate some features, but not others. These results raise questions about the frequent characterization of imitation in shadowing tasks as “automatic,” in the sense of emerging without deliberative control. While shadowing provides no interlocutor with whom to engage, participants may choose to imitate for their own amusement, particularly absent instructions to the contrary (or possibly in the face of such instructions). While it’s possible that the lack of imitation in the Avoid Imitate condition is a result of active self-regulation, it is also possible that the imitation in the No Instruction condition stems from deliberate, though unrequested, choices to play with language. Future research exploring different kinds of instructions can help to distinguish these possible interpretations.

While not necessarily a matter of deliberative control, participants’ own phonological structure provides an additional top-down influence. Our post hoc analysis of trap suggests that, as predicted, participants’ developing understanding of the model talker as someone who backs bath had an influence on the trap tokens that, to them, belong the same phoneme.

5.2 Imitation and social meaning

In the current study, the model talker spoke a dialect that was likely to be in high regard for the participants (Garrett et al., 2005), allowing us to examine the imitation of different dialect features in a prestigious variety. Previous work has shown selective imitation of dialect features of non-prestigious varieties in explicit imitation tasks (Clopper & Dossey, 2020; Dufour & Nguyen, 2013), as well as variability in the magnitude of imitation of dialect features in implicit imitation tasks (Babel, 2010; Walker & Campbell-Kibler, 2015). We observed significant imitation of all five dialect features in the Imitate condition, suggesting a lack of selectivity in convergence to dialect features of a high-regard variety.

We also observed significant imitation in the No Instruction condition in one of the two most remarked-on variables, supporting previous work (Babel, 2010; Walker & Campbell-Kibler, 2015). The importance of sociolinguistic awareness is further supported by our exploratory analysis showing that participants who commented on one of the features imitated it more in the Imitate condition and even in the No Instruction condition. An additional observation across all the features is that even when explicitly trying to avoid imitation, only one instance of divergence was seen and, indeed, even the bulk of the nonsignificant marginal mean patterns were in the converging direction. The significant effect of divergence in the Avoid Imitate condition with kit is surprising, as divergence effects are unusual across the imitation and accommodation literatures. It is possible that kit represents a special case in which divergence is unusually promoted, perhaps due to low sociolinguistic meaning for the American shadowers and high cross-dialect phonetic difference. Without support from future work, however, the lack of evidence of divergence elsewhere in the literature leads us to view this result as possibly spurious.

Finally, the post hoc analysis of a trial effect for trap tentatively suggests that the interplay between specific tokens and the phonemic system might provide useful terrain for issues of control. When asked to imitate, participants started out converging toward the trap tokens they heard but, over time, were apparently influenced by the more meaningful bath tokens, which for the participants were instances of the same phoneme. If this effect is real, it suggests that as the experiment progressed, the participants formed an awareness that the model talker produced backed versions of what, for them, is a single bath/trap phoneme. This awareness counteracted but did not override the influence of the immediately preceding fronted trap token, producing tokens that converged only minimally by the end of the session. The fact that trap underwent this attenuation and bath did not may be due to the greater phonetic distance of the latter, its existing social meaning, or a combination of the two.

5.3 Imitation and gender

Our post hoc analysis of gender suggested that the women in our study tended to show greater convergence than the men, as well as greater divergence in the single case of divergence that we found. This finding echoes previous reports of gender effects in the literature (e.g. Namy et al., 2002; Pardo, 2006), but its implications are unclear. Broad gender categories do not translate straightforwardly to linguistic behavior (Eckert, 2000), and understanding their contribution to imitation in this and similar studies would require a deeper understanding of the social dynamics of lab-based shadowing tasks than has yet been offered.

6. Conclusions

This study set out to investigate whether and how American English speakers are able to adjust their patterns of phonetic imitation upon instruction when shadowing a talker from a high-regard geographically distant variety. We found strong support for a deliberative component, with participants increasing cross-dialect imitation across the board when asked to do so, as well as eliminating it completely when asked not to imitate. Some imitation remained when neither instruction was given. Imitation in the No Instruction condition may be due to automatic forces, deliberate choice (playful or task-oriented) on the part of the participants, or a combination of the two.

Sociolinguistic awareness was a key component as well. Deliberate imitation was greater for those features which were remarked on generally, and even stronger imitation was seen in a feature that a given participant noted post-task. An analogous effect is not present when participants were asked to avoid imitation, however; the only significant divergence effect was found in kit, which attracted no comment from any participant. However, further research is needed to more explicitly test these effects, given the exploratory nature of our sociolinguistic awareness analysis and the descriptive (non-significant) trend in noticing rates by condition. If future work supports this divergence as a real effect, it will raise the question why the patterns for deliberate imitation and deliberate avoidance of imitation differ.

To better understand the interacting roles of deliberative and automatic imitation, it would be useful in future work to more clearly tease apart participant deliberative choices from experimenter instruction. In particular, the role of language play or deliberate exploration of language difference needs to be better understood to clearly distinguish automatic from controlled aspects of imitation.

Acknowledgements

We would like to thank Nandi Sims, Nick Bednar, Sruti Parthasarathy, Matthew Woodward, Amanda Ciani, Jacquelyn Jarachovic, Katherine Melchioris, Elizabeth Lamar, and Michelle Bullock for their work on this project.

Competing Interests

The authors have no competing interests to declare.

References

Austen, M. (2020). The role of listener experience in perception of conditioned dialect variation. [Doctoral Dissertation, The Ohio State University]. http://rave.ohiolink.edu/etdc/view?acc_num=osu159532560325774

Babel, M. (2010). Dialect divergence and convergence in New Zealand. Language in Society, 39, 437–456.  http://doi.org/10.1017/S0047404510000400

Babel, M. (2012). Evidence for phonetic and social selectivity in spontaneous phonetic imitation. Journal of Phonetics, 40, 177–189.  http://doi.org/10.1016/j.wocn.2011.09.001

Babel, M., McAuliffe, M., & Haber, G. (2013). Can mergers-in-progress be unmerged in speech accommodation? Frontiers in Psychology, 4, 653.  http://doi.org/10.3389/fpsyg.2013.00653

Badre, D. (2025). Cognitive control. Annual Review of Psychology, 76(1), 167–195.  http://doi.org/10.1146/annurev-psych-022024-103901

Boberg, C. (1999). The attitudinal component of variation in American English foreign (a) nativization. Journal of Language and Social Psychology, 18(1), 49–61.  http://doi.org/10.1177/0261927X99018001004

Buehler, D. (2025). What is cognitive control? Wiley Interdisciplinary Reviews: Cognitive Science, 16(2), e70004.  http://doi.org/10.1002/wcs.70004

Clopper, C. G., & Dossey, E. (2020). Phonetic convergence to Southern American English: Acoustics and perception. Journal of the Acoustical Society of America, 147, 671–683.  http://doi.org/10.1121/10.0000555

Clopper, C. G., Dossey, E., & Gonzalez, R. (2024). Raw acoustic vs. normalized phonetic convergence: Imitation of the Northern Cities Shift in the American Midwest. Laboratory Phonology, 15, 1–34.  http://doi.org/10.16995/labphon.10893

De Neys, W. (2023). Advancing theorizing about fast-and-slow thinking. Behavioral and Brain Sciences, 46, e111.  http://doi.org/10.1017/S0140525X2200142X

Dufour, S., & Nguyen, N. (2013). How much imitation is there in a shadowing task? Frontiers in Psychology, 4(346), 1–7.  http://doi.org/10.3389/fpsyg.2013.00346

Eckert, P. (2000). Linguistic variation as social practice: The linguistic construction of identity in Belten High. Blackwell.

Elliott, N. C. (2000). Rhoticity in the accents of American film actors: A sociolinguistic study. Voice & Speech Review, 1(1), 103–130.  http://doi.org/10.1080/23268263.2000.10761390

Garrett, P., Williams, A., & Evans, B. (2005). Attitudinal data from New Zealand, Australia, the USA and UK about each other’s Englishes: Recent changes or consequences of methodologies? Multilingua, 24, 211–235.  http://doi.org/10.1515/mult.2005.24.3.211

Giles, H., Mulac, A., Bradac, J. J., & Johnson, P. (1987). Speech accommodation theory: The first decade and beyond. Communication Yearbook, 10, 13–48.  http://doi.org/10.1080/23808985.1987.11678638

Goldinger, S. D. (1998). Echoes of echoes? An episodic theory of lexical access. Psychological Review, 105, 251–279.  http://doi.org/10.1037/0033-295X.105.2.251

Honorof, D. N., Weihing, J., & Fowler, C. A. (2011). Articulatory events are imitated under rapid shadowing. Journal of Phonetics, 39, 18–38.  http://doi.org/10.1016/j.wocn.2010.10.007

Horvath, B. (2004). Australian English: Phonology. In B. Kortmann & E. Schneider (Ed.), A Handbook of varieties of English: A multimedia reference tool. Volume 1: Phonology. (pp. 625–644). De Gruyter Mouton.  http://doi.org/10.1515/9783110197181-041

Kuznetsova, A., Brockhoff, P. B., & Christensen, R. H. B. (2017). lmerTest package: Tests in linear mixed effects models. Journal of Statistical Software, 82(13), 1–26.  http://doi.org/10.18637/jss.v082.i13

Labov, W. (1966). The social stratification of English in New York City. Center For Applied Linguistics, Washington, DC.

Lenth, R. (2024). emmeans: Estimated marginal means, aka least-squares means [Version 1.10.0. Computer software, R package]. https://cran.r-project.org/package=emmeans

Lin, Y., Yao, Y., & Luo, J. (2021) Phonetic accommodation of tone: Reversing a tone merger-in-progress via imitation. Journal of Phonetics, 87:101060.  http://doi.org/10.1016/j.wocn.2021.101060

Michalsky, J., & Schoormann, H. (2017). Pitch convergence as an effect of perceived attractiveness and likability. Proceedings of Interspeech 2017, 2253–2256.  http://doi.org/10.21437/Interspeech.2017-1520

Mitterer, H., & Ernestus, M. (2008). The link between speech perception and production is phonological and abstract: Evidence from the shadowing task. Cognition, 109, 168–173.  http://doi.org/10.1016/j.cognition.2008.08.002

Namy, L. L., Nygaard, L. C., & Sauerteig, D. (2002). Gender differences in vocal accommodation: The role of perception. Journal of Language and Social Psychology, 21(4), 422–432.  http://doi.org/10.1177/026192702237958

Nguyen, N., Dufour, S., & Brunellière, A. (2012). Does imitation facilitate word recognition in a non-native regional accent? Frontiers in Psychology, 3, 480.  http://doi.org/10.3389/fpsyg.2012.00480

Nielsen, K. (2011). Specificity and abstractness of VOT imitation. Journal of Phonetics, 39, 132–142.  http://doi.org/10.1016/j.wocn.2010.12.007

Pardo, J. S. (2006). On phonetic convergence during conversational interaction. Journal of the Acoustical Society of America, 119(4), 2382–2393.  http://doi.org/10.1121/1.2178720

Pickering, M. J., & Garrod, S. (2004). Toward a mechanistic psychology of dialogue. Behavioral and Brain Sciences, 27, 169–226.  http://doi.org/10.1017/S0140525X04000056

Podlipský, V. J., & Šimáčková, S. (2015). Phonetic imitation is not conditioned by preservation of phonological contrast but by perceptual salience. Proceedings of the 18th International Congress of Phonetic Sciences. https://www.internationalphoneticassociation.org/icphs-proceedings/ICPhS2015/Papers/ICPHS0399.pdf

Ross, J. P., Lilley, K. D., Clopper, C. G., Pardo, J. S., & Levi, S. V. (2021). Effects of dialect-specific features and familiarity on cross-dialect phonetic convergence. Journal of Phonetics, 86(101041), 1–23.  http://doi.org/10.1016/j.wocn.2021.101041

Sato, M., Grabski, K., Garnier, M., Granjon, L., Schwartz, J., & Nguyen, N. (2013). Converging toward a common speech code: Imitative and perceptuo-motor recalibration processes in speech production. Frontiers in Psychology, 4, 422.  http://doi.org/10.3389/fpsyg.2013.00422

Schertz, J., Adil, F., & Kravchuk, A. (2023). Underpinnings of explicit phonetic imitation: Perception, production, and variability. Glossa Psycholinguistics, 2(1), 4.  http://doi.org/10.5070/G601123

Shockley, K., Sabadini, L., & Fowler, C. A. (2004). Imitation in shadowing words. Perception & Psychophysics, 66, 422–429.  http://doi.org/10.3758/BF03194890

Wade, L., Embick, D., & Tamminga, M. (2023). Dialect experience modulates cue reliance in sociolinguistic convergence. Glossa Psycholinguistics, 2(19), 1–30.  http://doi.org/10.5070/G6011187

Walker, A. (2019). The role of dialect experience in topic-based shifts in speech production. Language Variation and Change, 31(2), 135–163.  http://doi.org/10.1017/S0954394519000152

Walker, A., & Campbell-Kibler, K. (2015). Repeat what after whom? Exploring variable selectivity in a cross-dialectal shadowing task. Frontiers in Psychology, 6(546), 1–18.  http://doi.org/10.3389/fpsyg.2015.00546

Wells, J. C. (1982). Accents of English: Volume 1 (Vol. 1). Cambridge University Press.

Yuan, J., & M. Liberman. (2008). Speaker identification on the SCOTUS corpus. Proceedings of Meetings on Acoustics 2008, 5687–5690.