1. Introduction

In many languages speakers vary prosody to realise focus (see Baumann & Kügler, 2015; Kügler & Calhoun, 2020 for a review); listeners rely on such changes in speech to make sense of incoming information (see Cutler et al., 1997; Cole, 2015; Dahan, 2015 for a review). Focus refers to the predication on a topic in a sentence and typically contains new or contrastive information to the interlocutor (Lambrecht, 1994; Vallduví & Engdahl, 1996). The form-function relationship between prosody and focus has been a long-standing subject of investigation in linguistics and related disciplines. Over the past three decades, the field has undergone substantial development. Reach on prosodic focus marking has been extended to typologically diverse or previously understudied languages (e.g., Calhoun, 2024; Kügler & Calhoun, 2020 and references therein; Liu & Wang, 2024; Roessig et al., 2024). At the same time, studies have increasingly adopted a holistic approach by examining the role of prosody in relation to morphosyntactic means such as particles (e.g., Chen et al., 2014 on Chinese; Tomioka, 2014 on Japanese), word order (e.g., Arnhold, 2016 on Finnish; Işsever, 2003 on Turkish; Jepson, 2023 on Djambarrpuyŋu; Portes & Reyle, 2022 on French; Skopeteas, 2014 on Greek) and cleft constructions (Arnhold, 2024 on Mandarin Chinese, Portes & Reyle, 2022 on French), and discourse factors such as prior beliefs (Ahn et al., 2024). More recent work has approached focus marking from a multimodal perspective by taking gesture into account (e.g., Gregori et al., 2024 on Catalan and German; Türk & Calhoun, 2024 on Turkish). Finally, research has begun to examine the role of signal-extrinsic factors (e.g., pragmatic skills) in explaining individual differences in prosodic ability (Bishop et al., 2020, 2024).

The advancements in research on the focus-to-prosody mapping in adults have resulted in new theoretical insights (e.g., Baumann, 2014; Kügler & Calhoun, 2020; Myrberg & Riad, 2014; Repp, 2014; Roettger et al., 2019; Truckenbrodt, 2014) and encouraged more detailed study of children’s prosodic focus marking from a production perspective (see Chen, 2018; Chen et al., 2020 for a review). Developmental research over the past two decades has primarily addressed three key questions, i.e., A. what is the developmental trajectory of adult-like prosodic focus-marking in a given language, B. how do cross-linguistic similarities and differences in the use of prosody in focus marking relative to morpho-syntactic means (e.g., word order, cleft construction) and the broader prosodic system (e.g., tone languages vs. non-tone languages) shape the rate and route of the acquisition of prosodic focus marking in children acquiring different languages (Chen 2018 and references therein, Destruel et al., 2024; Kapia & Kleber, 2024; Romøren & Chen, 2022; Yang et al., 2024), and C. how do children use gesture in focus marking in relation to the use of prosody (Coego et al., 2025; Esteve-Gibert et al., 2022). These lines of research are largely built on or align with the progress in research on adults’ prosodic focus marking in either using findings on adults to establish the target model in the ambient language or seeking cross-linguistic insights.

The present study investigates individual differences in children’s prosodic focus marking. Unlike previous developmental research on prosody, it does not adopt the perspectives found in adult individual differences research, which typically focuses on signal-extrinsic factors (Bishop et al., 2020, 2024) or uses individual variation in prosodic production to test predictions from hypothesised production mechanisms (e.g., Bishop & Intlekofer, 2020; Engelhardt et al., 2013). Instead, this study draws more closely on individual differences research in other areas of language acquisition, addressing the developmental interconnectedness between linguistic subsystems and between language and other cognitive systems (e.g., Kidd & Donnelly, 2020).

1.1. Individual differences in language acquisition

Following the pioneering work by Bates and associates over three decades ago (Bates et al., 1988; Bates et al., 1995), research on individual differences in language acquisition has gained momentum in recent years (see Kidd et al., 2018; Kidd et al., 2020; Kidd & Donnelly, 2020; Norbury, 2019 for a review). Kidd and Donnelly (2020) argued that adopting a correlational approach, the study of individual differences ‘provides a methodology for theory testing and development and points us toward mechanisms and architectural constraints on the acquisition process’ (321). More specifically, its significance for the field of language acquisition can be seen in the following three ways (Bates et al., 1995; Kidd & Donnelly, 2020). First, at an empirical level, such research informs the field on the scale (and degree) of differences in language development among typically developing children. This knowledge can have potentially useful implications for atypical language development, for example by providing developmental benchmarks and helping to quantify the extent of delays or deficits. Second, it can shed light on the interconnectedness of linguistic sub-systems. If variation in the development in one subsystem (e.g., vocabulary) can explain variation or can predict variation in the development in another subsystem of language (e.g., grammar), we may have evidence for a common underlying mechanism supporting the acquisition of both. Third, individual differences research can provide valuable information on whether the development of a linguistic subsystem is related to the development in other cognitive systems. Evidence for a connection between the development of the linguistic subsystem and non-linguistic abilities would suggest that language is not learned in isolation of development in other areas.

Substantial individual differences have been reported for children’s linguistic input, early auditory and speech processing ability, early vocabulary acquisition, and grammatical development. Individual differences appear to be even more pronounced in the learning of complex linguistic constructions and their usage. Importantly, research has also shown that individual differences in one domain can explain or influence developmental variation in another. For example, variation in children’s vocabulary and grammatical knowledge can explain variation in their tendency to use passive voice after priming. In a picture description task, Kidd (2012) found that four- to six-year-old English-speaking children produced more passive constructions after being primed to use passive voice than before, but those with higher vocabulary and grammatical knowledge scores showed a larger priming effect. Similarly, variation in both the quantity and the quality (e.g., lexical diversity, syntactic complexity) of children’s linguistic input significantly predicts later vocabulary and grammar development, while variation in early auditory and speech processing ability relates to later vocabulary outcomes (see Kidd & Donnelly 2020 for a review).

Individual differences in prosodic development have, however, not been systematically documented in earlier work, even though deficits in prosodic abilities are frequently observed in children with development disorders such as autism spectrum disorder (ASD), developmental language disorder (DLD), cerebral palsy, and hearing loss (Paul et al., 2020). Also, it remains to be investigated whether prosodic development relates to the development in the other linguistic subsystems and cognitive systems. Addressing these notable gaps is essential for the construction of theories for prosodic development that can ‘allow for and predict the existence (and absence) of individual differences’ and ‘account for the observed developmental relationships between linguistic subsystems and between language and other cognitive systems’, two empirical observations for which any theory of acquisition should account according to Kidd and Donnelly (2020). Furthermore, it is strongly needed for validating theoretical insights into individual differences in language acquisition, which are currently based entirely on research on the non-prosodic aspects of language development.

1.2. The present study

As a first step towards narrowing the gap between individual differences research on prosodic development and similar research on the development in other linguistic subsystems, the current study had two main aims. First, it sought to map individual differences in the use of prosody for focus marking in production in Dutch-speaking school-aged children, compared to adults’ production. Second, it set out to establish whether acquisition of prosodic focus marking relates to vocabulary development, another linguistic subsystem, and to musicality, a non-linguistic cognitive ability, by examining whether observed individual differences in prosodic focus marking can be explained by variation in vocabulary and musicality. Prosodic focus marking is well studied for declarative sentences with different focus structures in both adult and child speakers of Dutch at the group level (Section 1.2.1). Our choice to focus on possible connections with the development in vocabulary and musicality was motivated by current understanding of the interconnectedness between these abilities and other linguistic abilities and existing research on their (possible) connection with prosodic development (Sections 1.2.2 and 1.2.3).

1.2.1. Prosodic realisation of focus in adult and child Dutch

Adult speakers of Dutch use both phonological and phonetic means to realise non-contrastive narrow focus (hereafter narrow focus) in different positions, i.e., sentence-initial, medial and final positions (Chen, 2009; Romøren, 2016), and to distinguish different types of focus, i.e., narrow focus, contrastive narrow focus (hereafter contrastive focus), and whole-sentence broad focus (hereafter broad focus) (Hanssen et al., 2008; Romøren, 2016). Phonological means refer to the use of coarse-grained changes in duration, pitch, and intensity that lead to a change in the phonological category of prosody, such as placement of pitch accent and choice of pitch accent type. Phonetic means refer to the use of phonetic implementation or realisation of a pitch accent in the dimensions of duration, pitch, and intensity for focus-marking purposes, without changing the type of the pitch accent (Chen, 2009, 2018). These variations make Dutch an interesting case for studying individual differences in children’s prosodic focus marking. Specifically, in narrow focus, the subject noun preceding the focal word is nearly always accented, typically with a falling pitch accent (H*L) or a high-rising pitch accent (H*); the content words following the focal word are frequently accented with a downstepped falling pitch accent (!H*L) sentence-medially, and are typically deaccented but sometimes accented with !H*L sentence-finally.1 The prosody of the focal word varies depending on its position in the sentence. If the focal word is sentence-initial or sentence-medial, it is frequently accented with H*L and !H*L respectively, with adjustments in the phonetic realisation of the pitch accent for increased prominence (e.g., a longer duration, a larger pitch span). If it is sentence-final, it is frequently accented with H*L and sometimes with !H*L, with adjustments in the phonetic realisation to make it prosodically more prominent than the !H*L in the postfocus constituent. Comparisons between the focal word in (sentence-medial) narrow and contrastive focus conditions and the same word in the broad focus condition show that the focal word is accented with the H*L accent in all the three conditions but the accent is realised with a longer duration, a lower minimum pitch, and an earlier peak alignment in narrow and contrastive focus than in broad focus. The focal word also appears to have a longer duration in narrow focus than in broad focus. These findings show that there is a more transparent form-meaning mapping between the phonological means and sentence-final focus. That is, if a sentence-final word is accented and with H*L, it has a high likelihood of being focal; if it is unaccented, it has a high likelihood of being non-focal. Further, there is stronger reliance on the use of phonetic means to distinguish focus types and differentiate focus from non-focus in the non-final narrow focus conditions.

Cross-linguistically, children tend to acquire phonological means faster in languages with more transparent phonological form-meaning mappings than in languages with less transparent ones, and acquire phonological means faster than phonetic means if the target language replies less on the latter for focus marking (Chen, 2018; Yang et al., 2024). Studies on Dutch-speaking children show that differences in transparency of the form-meaning mapping and reliance on phonetic means between different focus structures also lead to differences in ease of acquisition between different focus conditions in the same language in a similar way, making the use of prosody in sentence-initial and medial narrow focus harder to acquire than in sentence-final narrow focus (Chen, 2009, 2011a, 2011b; Romøren, 2016). Specifically, Dutch-speaking children start to use accent type to mark narrow focus in the two-word stage (at about 2 years of age when children can produce utterances with two words), become adultlike in the use of accent placement to mark sentence-final narrow focus at the age of 4 or 5 years, can use choice of accent type in sentence-final narrow focus in a largely adult-like manner at the age of 7 or 8 years. In contrast, they do not use phonetic means to mark sentence-initial narrow focus at the age of 4 or 5 years and can use pitch cues but not duration cues at the age of 7 or 8. They are not adult-like in the phonetic marking of sentence-medial narrow focus at the age of 10 or 11 years. However, it has been found that English-learning 4- to 5-year-olds use phonetic means to mark focus in sentence-initial position if it is contrastive (Wonnacott & Watson, 2008), suggesting the presence of contrast can make the use of phonetic means in sentence-medial contrastive-focus easier to acquire than in sentence-initial and medial narrow focus. It is not clear when Dutch-speaking children become adult-like in the use of phonetic means to distinguish broad focus and sentence-final focus, in both of which the sentence-final word is accented with the falling pitch accent H*L. But given the general difficulty children have with the learning of phonetic means, it is possible that this ability is not acquired before the age of 7 or 8.

1.2.2. Development in vocabulary size and other linguistic subsystems

Developmental research consistently shows that early vocabulary size, an aspect of lexical development, is a strong predictor of later language abilities, literacy skills, or verbal intelligence (e.g., Verhoef et al., 2024; Bleses et al., 2016; Duff et al., 2015; Hao et al., 2008; Marchman & Fernald, 2008). For example, Duff et al. (2015) found in their five-year longitudinal study on 16- to 24-month English-learning children that early expressive vocabulary significantly predicted school-age vocabulary, phonological awareness, reading accuracy, and reading comprehension. Research on individual differences has also shown that development in vocabulary is coupled with development in grammar, although current evidence does not suggest a directional dependence between them (Kidd & Donnelly 2020).

The relationship between prosodic and lexical development has been widely discussed in the theory of prosodic bootstrapping, which holds that prosodic cues help children uncover basic lexical and grammatical properties of the ambient language (Gleitman & Wanner, 1982; Jusczyk, 1997; Morgan & Demuth, 1996; see Gervain et al., 2020; Höhle, 2009 for a review). The prosodic ability of interest is perceptual rather than productive: early perceptual sensitivity to rhythmic differences related to syllabic structures (e.g., trochee or the strong-weak pattern as in the noun ‘permit’ versus iamb or the weak-strong pattern as in the verb ‘permit’) (Nazzi et al., 1998; Ramus et al., 2000; Ramus, 2002) helps infants segment words from continuous speech streams, thereby supporting lexical development (e.g., Höhle et al., 2001; Houston et al., 2000; Nazzi et al., 2006; Ramus et al., 1999).

Productive prosodic development has also been linked to vocabulary growth. Past research based on a holistic approach to intonation (i.e., utterance-level pitch characteristics) emphasised a connection between significant development in intonation and the emergence of multiword utterances, an important milestone of grammatical development (Snow, 2000, 2006; Snow & Balog, 2002). However, more recent research conducted within the autosegmental metrical framework (see Arvaniti & Fletcher, 2020 for a review) shows that children’s ability to produce language-specific intonation contours emerges after reaching certain vocabulary thresholds. For example, children aged 1 to 2 years can produce an adult-like inventory of nuclear contours (i.e., pitch accent + boundary tone combinations) at or around the 25-word point in Catalan and Spanish (Prieto et al., 2012), after reaching the 20-word point in European Portuguese (Frota et al., 2016), between the 100-word point and the 160-word point in Dutch (Chen & Fikkert, 2007).2 There is also evidence that the phonetic realisation of nuclear pitch accents becomes adult-like at around the 25-word point in some languages but not earlier. For example, European Portuguese-learning children can realise the nuclear contours with adult-like pitch range and peak alignment (in the pitch accent) after reaching the 20-word point (Frota et al., 2016). English- and French-learning children begin to use pitch to realise utterance-level stress (equivalent to nuclear pitch accent in the autosegmental metrical framework) in an adult-like way after reaching the 25-word point (Vihman & DePaolis, 1998; Vihman DePaolis & Davis, 1998). These findings suggest a need for acquiring some words to exercise and improve articulatory control of pitch to the precision needed for adult-like production (DePaolis et al., 2008; Prieto et al., 2012; Frota et al., 2016).

Whether this also applies to the use of prosody for focus marking in older children remains unclear. It is possible that vocabulary size is less important than grammatical development in the context of prosodic focus marking. Children frequently produce the topic-focus structure in successive one-word utterances in the one-word stage, i.e., the stage when children primarily produce one-word utterances, (e.g., Finger. Touch. said when the child was reaching out to touch the microphone used in the recording’) (Scollon, 1979) and in two-word utterances in the two-word stage (e.g., Federica acqua ‘Federica water’, said when the child was pretending to drink) (D’Odorico & Carubbi, 2003). It is not clear whether children already use prosody to distinguish topic and focus in the one-word stage. But limited studies on Dutch- and English-learning children show that children can use accentuation in words carrying new or contrastive information, but downstepped accents, devoicing or deaccentuation in words carrying given information in the late two-word stage (Chen, 2011b; Wieman, 1976). These findings suggest that the ability to vary accent placement, choice of pitch accent and phonetic realisation of pitch accent in focus marking may be more closely coupled with grammatical development.

A possible connection with vocabulary size may, however, arise via parental wh-questions. These types of questions are known to play an important role in young children’s vocabulary development (Oshima-Takane & Titova, 2021 and references therein). According to Rowe et al. (2017), a wh-question draws the child’s attention toward a particular object or event and gives the child the opportunity to make the connection between a word and its referent or action and practise the use of the word. If the child does not use the right word, she or he often receives feedback from the parent, who provides the correct word. This question-answer-feedback loop can thus be an efficient way for the child to build his or her vocabulary. Importantly, wh-questions also elicit responses with different focus structures. If certain questions are asked more often in the input than other questions, the child has more opportunities to practice the use of prosody in the corresponding focus structures. If the child does not respond in a prosodically appropriate manner, his or her interlocutor may rephrase the answer to provide a prosodically appropriate version of the answer or ask clarification questions, giving the child a chance to reflect or repair. In the input of English-speaking 1–3 year olds, wh-questions follow a rough frequency hierarchy: what, where, and who are most frequent; followed by when, why, and how; with which and whose being least frequent (Rowland et al., 2003). Broad-focus what-happens questions appear to be relatively rare in the input. This distribution suggests that vocabulary development and the ability to use prosody for focus marking may become concurrent in the focus structures elicited by the most common wh-questions. In addition, neurolinguistic research shows that when wh-questions put the word carrying the requested information in the response in focus, they also trigger deeper processing of the meaning of the word in the brain than when the word is not in focus (Wang et al., 2011). This enhanced processing advantage of focal words may in turn facilitate word learning.

Taken together, these findings reflect a well established link between vocabulary development and other aspects of language, such as grammar and phonological awareness. There is also some evidence linking vocabulary size to the emergence of adult-like production of prosodic patterns in the one- and two-word stages, but little is known about its role in prosodic focus marking. Parental wh-questions in the input may provide a pathway by jointly supporting word learning and eliciting focus structures that require appropriate use of prosody.

1.2.3. Musicality and other linguistic subsystems

Musicality can be broadly defined as an individual’s cognitive capability for music, shaped by both musical aptitude, i.e., innate potential for music learning, and formal instruction in music (Rhee et al. 2022). It can improve with age during childhood (Welch, 1998) and can be further developed through musical instructions and exposure (Besson et al., 2007). Musicality has been positively associated with various aspects of language development in typically developing children. Prior research has generally examined one dimension of musicality at a time, e.g., musical aptitude, musical expertise (obtained via years of formal instruction in music apart from music lessons through compulsory school education), musical training, i.e., experimentally administered musical training for a certain period of time to children with no prior musical expertise, following the definitions of Chobert et al. (2011).

Several studies have found that musical aptitude is positively linked to phonological skills in children. For example, Anvari et al. (2002) reported that musical aptitude in 4- to 5-year-olds was significantly associated with both phonological awareness and early reading ability. Milovanov et al. (2009) further found that children aged 10 to 12 with strong musical aptitude and good pronunciation skills showed enhanced brain responses to changes in speech sound duration, compared to peers with lower aptitude. These findings suggest that musical aptitude may contribute meaningfully to speech processing and early literacy development in children.

In addition to aptitude, musical training has been found to enhance several core language skills in children. Benefits that have been reported across studies uitlising a pretest-training-posttest paradigm include improvements in verbal memory (Ho et al., 2003), verbal intelligence, vocabulary (Linnavalli et al., 2018), syntactic processing (Jentschke & Koelsch, 2009), phonological awareness (Vidal et al., 2020), and phonological processing (Bolduc et al., 2020; Chobert et al., 2014). These findings are further supported by recent reviews. For example, a review by Pino and Giancola (2023) highlighted the positive effects of rhythm- and melody-based activities on children’s phonological, grammatical, and syntactic development. A meta-analysis by Barakat et al. (2024) confirmed consistent gains in phonological skills following preschool music interventions involving rhythm and pitch.

While research remains limited, an increasing number of studies suggest that musicality contributes to children’s prosodic development, in accordance with shared reliance on pitch and rhythm between speech prosody and music (e.g., Hausen et al., 2013; Patel, 2007). In terms of perception, several studies show that musical training enhances sensitivity to prosodic features. For example, Magne et al. (2006) and Moreno et al. (2009) found that children with extended musical training were better at detecting pitch violations in speech, both behaviorally and neurally. Nan et al. (2018) reported that Mandarin-speaking children who received piano training exhibited enhanced neural pitch tracking and improved speech perception relative to untrained peers. Similarly, a meta-analysis by Jansen et al. (2023) confirmed a moderate correlation between musical ability, especially pitch and rhythm perception, and prosodic perception across ages. Evidence also points to a relationship between musicality and prosodic production. Cardillo (2008) found that musical aptitude, measured through pitch and rhythm discrimination tasks, significantly predicted 5-year-old children’s ability to imitate speech prosody, specifically, to replicate stress placement in reiterative /ma/ sequences of varying syllabic lengths. This ability, in turn, predicted phonological awareness, suggesting a developmental link from musical aptitude to prosodic production to metalinguistic skill. Similarly, Rhee et al. (2022) showed that Mandarin-speaking 4- and 5-year-olds with higher musical aptitude produced acoustically more distinct lexical tones than their peers.

Taken together, these findings reflect a consistent and growing body of evidence that musicality, whether defined as aptitude, musical expertise or musical training, supports a range of language skills, including prosodic development. However, most studies to date have focused on general prosodic sensitivity or lexical tone production, rather than children’s ability to use prosody for focus marking specifically.

1.2.4. Research questions and hypotheses

Three research questions were addressed in this study: (1) what are the individual differences in the use of prosody for focus marking in production among Dutch-speaking school-aged children? (2) Is children’s ability in prosodic focus marking related to their development in vocabulary, another linguistic subsystem? (3) Is this ability related to their development in musicality, another non-linguistic system? To investigate these questions, we conducted a longitudinal study on Dutch-speaking 4- to 7-year-olds’ use of prosody in marking narrow focus in sentence-initial, -medial and -final positions, contrastive focus in sentence-medial position, and broad focus in SVO sentences, produced as responses to wh-questions or statements (in the case of contrastive focus). Their vocabulary size and musicality were assessed at multiple longitudinal points.

It has been argued that individual differences are ‘large and notably stable across development … observed early and across all domains’ in first language acquisition (Kidd et al., 2018) and they appear to be even more pronounced in the learning of complex linguistic constructions and their usage (Vasilyeva et al., 2008). Wells et al. (2004) found that individual differences could occur at both younger and older ages and in different aspects of prosodic development in their study with 120 English-speaking 7- to 13-year-olds performing multiple production and comprehension tasks. But they did not assess the degree of individual differences at different ages or in specific aspects of prosodic development. Production data from a sample of Dutch-speaking 4- to 5-year-olds (N = 22) and 7- to 8-year-olds (N = 18) showed that there were noticeable individual differences in prosodic focus marking in the younger children but not in the older children (Chen, 2011a), contra Kidd et al. (2018). We thus hypothesised that there would be more individual differences in more demanding focus conditions than in less demanding focus conditions, especially at a younger age (Hypothesis I). We predicted to observe more individual differences in sentence-initial narrow focus, sentence-medial narrow focus and possibly broad focus than in sentence-final focus and sentence-medial contrastive focus, and more pronounced individual differences at younger ages than at older ages. Sentence-initial and -medial narrow focus conditions and broad focus were considered to be more demanding than sentence-final focus and sentence-medial contrastive focus because focus marking in the former entails the use of phonetic means and Dutch-speaking children become adult-like in their prosody in some of the former conditions at a later age than in sentence-final focus (Section 1.2.1).

Based on the evidence for the relationship between vocabulary development and prosodic development reviewed in Section 1.2.2, we hypothesised that children’s vocabulary size would be connected with their ability in prosodic focus marking, especially at younger ages (Hypothesis II). Our predictions were that variation in children’s vocabulary size could explain variation in children’s ability in prosodic focus marking in a statistically significant way and that this pattern would be stronger at an younger age than at an older age. No hypothesis was formulated regarding differences across focus structures, as no empirical data is available on the frequency of wh-question types in the input of Dutch-learning children. Potential differences that emerged in our analysis will instead be discussed post hoc in Section 4, informed by findings on parental wh-questions in the input of English-learning children.

In the light of the evidence for the relationship between musicality and prosodic development reviewed in Section 1.2.3, we hypothesised that children’s musicality would be connected with their ability in prosodic focus marking to different degrees at different ages (Hypothesis III). Musicality was operationalised as musical aptitude. Our predictions were that variation in children’s musicality could explain variation in children’s ability in prosodic focus marking in a statistically meaningful way and that this pattern would be stronger at an younger age than at an older age.

2. The methodology

SVO sentences were elicited in different focus conditions from monolingual Dutch-speaking children at different ages over a three-year period and from monolingual Dutch-speaking adults (as the adult control group) via a picture-matching game adapted from Chen (2011a), preceded by a picture-naming task, in a natural and interactive setting, similar to recent studies on children acquiring Mandarin Chinese (e.g., Yang & Chen, 2018), Dutch (Romøren & Chen, 2015; Romøren, 2016) Korean (e.g., Yang et al., 2024) and Swedish (e.g., Romøren & Chen, 2022). Focus conditions were varied by means of the question-answer paradigm (Roberts, 1996). The picture-naming task was conducted to familiarise the children with the characters and actions that appeared in the picture matching game and to make sure they would use the intended words in the picture matching game. The adult participants did not do the picture-naming task as it was deemed unnecessary. The SVO sentences elicited in the pitch-matching game were subsequently rated individually for prosodic appropriateness in the corresponding context on a five-point scale by trained adult listeners. The scores for prosodic appropriateness reflected the integrated use of prosody, instead of the use of a specific prosodic parameter (e.g., pitch) or a specific prosodic strategy (e.g., accent placement), capturing the perceptual relevance of production details. This makes perceptual rating an effective alternative for assessing children’s prosodic focus marking in production on a large scale. Additionally, vocabulary development and musicality were assessed for the children at different time points. The research was performed in accordance with the Declaration of Helsinki.

2.1. Participants and test timeline

One hundred and eight Dutch-speaking children from four age groups (51 male, 58 female ) were enrolled in this study. The children were on average 4;7 (n = 30; male = 12), 5;6 (n = 33; male: 14), 6;4 (n = 24; male: 13), and 7;5 (n = 21; male: 12) in their respective age group at the start of the study. See supplementary file 1 for age and gender statistics of the children throughout the study. The children were from monolingual Dutch-speaking families and were recruited from four primary schools in the province Utrecht in the Netherlands. They had no hearing loss and speech or language disabilities according to parents’ reports. We received written informed consent from the parents for the children to participate in the longitudinal study and for their experimental sessions to be recorded.

Forty-six adult native speakers of Dutch (13 male, 33 female, mean age = 21;8, SD = 4;8, range: 18;8 – 48;0) participated in the production study as the control group to provide a measure of adultlike production.3 They were university students at the time of testing and were tested on the same tasks following the same procedure as the children. Prior to the experiment, they were informed that the tasks were of a simple nature as they were also used on child participants.

The children were tested four times in a period of three years with an interval of three months between phase I and phase II, and five to six months between phase II and phase III, and 16~19 months between phase III and phase IV. In each phase, they did multiple tests. The production test was conducted in each test phase to gather dense data on development in production. The vocabulary and musicality tests were distributed over two or three phases to generate enough data to model the development for the entire age range from 4;1 to 10;3: the vocabulary test in phases I, III and IV; the musicality test in phases II and IV. Actual participation varied across phases due to missing sessions and the longitudinal design. A total of 108, 103, 95, and 105 children participated in Phases I–IV, respectively (i.e., contributing data to at least one task in the respective phase).

2.2. The production experiment

Fifteen question-answer dialogues were embedded in the picture-matching game to elicit 15 SVO sentences in five focus conditions: narrow focus in sentence-initial position (NFi), responding to who-questions; narrow focus in sentence-final position (NFf), responding to what-questions; narrow focus in sentence-medial position (NFm), responding to what-does-X-do-with-Y questions; contrastive focus in sentence-medial position, correcting the experimenter’s statement about the action (CFm); and broad focus over the whole sentence (BF), responding to what-happens questions.4 The target SVO sentences were unique combinations of five subject-nouns (baker, ‘baker’, hond, ‘dog’, leeuw, ‘lion’, meisje, ‘girl’, poes, ‘cat’), three verbs (tekenen, ‘draw’, koken ‘cook’, toveren ‘conjure’), and three object-nouns (wortel, ‘carrot’, laars, ‘boot’, lepel, ‘spoon’) (see Supplementary file 2 for the full list of stimuli). All words were highly familiar to Dutch 4-year-olds.

Three female student assistants, who were native speakers of Dutch and students of linguistics or literature, conducted the production experiment with the children individually at their schools during school time and with the adults in a sound-attenuated booth in the Institute for Language Sciences Labs at Utrecht University. A fourth female student assistant, who was also a native speaker of Dutch, assisted with testing at the schools during the final phase of the study to alleviate the workload of the primary experimenters. To ensure consistency of the highest degree between the experiment sessions, the experimenters were trained on both the procedure and the use of prosody using an experiment protocol before the testing. To help the children to feel at ease with the experimenters, every experimenter picked up the child that she was to test from the classroom and had a chat with the child prior to the experiment.

2.2.1. The picture-naming task

In the picture-naming task, the participants first named the nouns illustrated in individual pictures. The experimenter showed one picture a time to the participant, and invited the participant to name it by saying ‘Dit is een … ’ (‘This is a …’). In the case of incorrect naming (e.g., calling a baker a cook), the experimenter first acknowledged that the entity might look like what the participant had in mind, and then drew the participant’s attention to distinctive features in the picture (e.g., the dough on the table) and suggested to the participant the intended label for the entity. Second, the experimenter showed the participant pictures of the actions occurring in the game, and explained to him which action each picture depicted. Finally, to check whether the participant has got all the labels right, the experimenter went through the pictures and asked the participant to name them once more. If the participant did not have difficulty with naming the pictures using the intended nouns during the first part of the picture-naming task, he only named the actions during the checking round.

2.2.2. The picture-matching game

In this game, the child was supposed to help the experimenter to put pictures in matched pairs. The experimenter first explained to the participant how the game worked, and introduced two rules of the game: A. Never reveal his pictures to the experimenter; B. Always say everything he sees happening in the picture in his response. The first rule helped to justify the purpose of the game and the use of prosody to mark focus in the child’s response; the second rule made sure that the child answered in full sentences. The game consisted of fifteen trials, corresponding to the fifteen question-answer dialogues. Each trial was carried out in a fixed number of steps (Figure 1). First, the experimenter took a picture from her set of pictures (e.g., a picture of a girl drawing something on a piece of paper), drew the participant’s attention to the picture, and briefly described it by saying “Kijk! Het meisje. Het lijkt alsof het meisje iets tekent.” ‘Look! The girl. It seems that the girl is drawing something.’ She then asked the participant a question about the picture (e.g., “Wat tekent het meisje?” ‘What is the girl drawing?’) or made a guess about the missing information in the case of contrastive focus. The participant then took a picture from his own set of pictures, and identified the information requested by the experimenter. The experimenter repeated her question when possible. The participant then answered the question in an SVO sentence (e.g., “Het meisje tekent de lepel.” ‘The girl is drawing the spoon.’). The experimenter thanked the participant for his information, looked for the matching picture (e.g., the picture of the spoon) in the box, and handed over both pictures to the participant for his approval.

Figure 1
Figure 1

An example trial of sentence-final focus in the picture-matching game (reproduced from Chen & van den Bergh, 2025).

The game proper was preceded by five practice trials. If the participant gave elided answers or full-sentence answers containing non-target words during the practice trials, the experimenter reminded the participant of the rules of the game or the intended words. This intervention turned out to be necessary only in the case of a small number of the 4- to 5-year-olds. Most children provided full-sentence answers using the intended words right from the start. Each session lasted about 15–25 minutes, and was recorded at a sampling rate of 44.l kHz with 16 bits resolution, and filmed. The film recordings were used to check consistency in how the game was carried out.

2.2.3. Data annotation

The audio recording of each participant was first orthographically annotated using Praat (Boersma, 2001). Second, full-sentence responses were selected as usable responses if they were not plagued by any of the following factors: self-correction, use of pronouns, use of non-target words, detectable hesitation-induced silences, responding to a non-target question, elided responses, overlap with the experimenter’s speech, and poor recording quality. Third, the usable full-sentence responses and the corresponding questions or statements were selected and extracted as individual .wav files.

2.2.4. Perceptual rating

The usable full-sentence responses and corresponding questions or statements were combined into context-response dialogues with a 300 ms interval between the question or statement and the response in each dialogue and a 1000 ms interval between dialogues. The three native speakers of Dutch who administered the production experiment as the primary experimenters served as the raters. Their familiarity with both child speech and the specific speech material to be evaluated could enable consistent and targeted assessments of prosodic features, minimising the influence of irrelevant variation in child speech, such as slower speaking rate, non-adult-like articulation, or immature voice quality. The raters rated each response in each dialogue on how well its prosody fitted in the context on a five-point Equal Appearing Interval scale, with 1 standing for ‘does not fit’ and 5 standing for ‘fits perfectly’. Prior to the rating, the raters were given sound examples illustrating what the prosody typically sounded like in each focus condition and written instructions on how to do the rating. They did the rating at least ten months after the completion of the production experiment at each testing moment, minimising the influence of the raters’ familiarity with individual children on their ratings.

To reduce variation in the scores due to comparisons between children, the dialogues were presented to the raters per child and per age group. The rating experiment was conducted in seven 30- to 40-minute sessions. The raters could listen to each dialogue maximally three times before finalising the score.

2.3. Assessment of vocabulary development

The children’s vocabulary development was assessed as the size of their receptive vocabulary, using the Peabody Picture Vocabulary Test – III (Dutch) (Schlichting, 2005). This test was standardised on a national sample of 1746 children ( 2;3 to 15;11) and 1164 adults (17–90 years) in the Netherlands and can be used on children and adults aged between 2;3 and 90. It includes in total 204 words and corresponding pictures (i.e., four colour pictures for each word arranged in a two-by-two grid, including one target picture and three distractor pictures), divided in 17 sets of 12 words each. The test is administered individually. The test-taker’s task is to identify the picture that corresponds to the word spoken by the examiner following a scoring sheet, either by pointing to it or saying its number aloud. Based on a child’s age, the test begins with an age-appropriate start set. If the child makes five or more errors in the start set, the tester switches to an earlier set until the child makes less than five errors. The final set is the set in which the child makes nine or more errors. In our study, we used a digitised version of the test materials. That is, each word and its corresponding pictures were presented on a PowerPoint slide; the slides were ordered following the scoring sheet. All the words were prerecorded by a female speaker of standard Dutch. To familiarise the children with the procedure, three practice items were administered prior to the formal testing phase, which consisted of presenting test words in increasing order of difficulty while responses were recorded by the tester. It took about 25 to 30 minutes to administer the test on average. The raw score, i.e., the total number of correctly identified words from the start set to the final set of words, was converted to a standard score using age-based normative data provided in the PPVT-III (Dutch) manual.

2.4. Assessment of musicality

The participants’ musicality was assessed individually using the Primary Measures of Music Audiation (PMMA), a standardised test designed to measure the musical aptitude of children in primary grades (Kindergarten to Grade 3 or aged approximately between 5 and 9 years), independent of musical training (Gordon, 1979). It consists of a tonality subset and a rhythm subset. The tonality subtest assesses the ability to perceive pitch-based differences between tunes; the rhythm subtest evaluates the ability to perceive differences in temporal duration between sequences of tones of the same pitch. Only the tonality subset was used in our study. While administering both subtests would provide a broader measure of musicality, the tonality subtest was judged most suitable for examining the hypothesised connection between musicality and the ability in prosodic marking of focus for the following considerations. First, pitch is the primary cue for marking focus in Dutch declarative sentences, whereas durational cues such as lengthening play a secondary role (Section 1.2.1), making pitch perception more relevant to our study than sensitivity to duration. Second, pitch-based musical ability seems to be related to pitch processing in language, as pitch processing in music and language relies on overlapping neural systems (Patel, 2007). For example, Wong et al. (2007) found that individuals with no prior experience with a tone language but with better musical pitch perception due to their musical expertise showed enhanced sensitivity to lexical tones in Mandarin Chinese. Jansen et al. (2023) four cross-study evidence in their meta-analysis that music perception was more strongly correlated with pitch perception than with timing perception in speech. In addition, practical considerations such as limited testing time and attention span in young participants necessitated a selective approach.

The tonality subset contains 40 items. For each item, the test-takers’ task is to listen to two short melodies and indicate whether the melodies are identical by selecting the corresponding picture pair in their answer sheets. If the two musical melodies are perceived as the same, they circle the picture pair with identical images (always fruits). If the phrases are perceived as different, they circle the picture pair with different images. In our study, the children were tested individually. The music melodies were played via loudspeakers at an adequate volume; a beep could be heard before each test item. To familiarise the children with the task and the response format, six practice items were administered before formal testing. It took about 20 to 25 minutes to administer the test on average. Raw scores for the test were calculated by summing the number of correct responses. These raw scores were then converted to age-normed percentile ranks using normative data provided in the PMMA manual.

3. Analyses and results

Of the 108 child participants enrolled in this study, 108, 102, 93, and 105 participants provided usable production data in phases I-IV, respectively. Vocabulary scores were available for 108 participants in Phase I, 90 participants in phase III, and 80 participants in phase IV. Musicality scores were available for 102 participants in phase II and 104 participants in phase IV. See supplementary file 3 for descriptive statistics for the production scores (Table 3-1), the vocabulary scores (Table 3-2), and the musicality scores (Table 3-3).

3.1. Production

We computed the production score for each test item of each participant in each focus condition by averaging the scores from the three raters. The production scores were nested both within participants (i.e., measurements from the same child are more alike than those from different children) and within items (i.e., measurements from the same stimulus are more alike than those from different stimuli). We thus adopted a cross-classified multilevel modelling approach to analyse the production scores (Goldstein, 2011; Quené & van den Bergh, 2008; Snijders & Bosker, 1999). The multilevel modelling was conducted using the software MLwiN (version 2.32) (Rasbash et al., 2015).

To model the development of children’s prosodic focus marking and examine individual differences, we used a structured, stepwise model-building approach. A series of seven models was fitted (see Supplementary file 4), starting with a baseline model including only the intercept and random effects, and progressively adding fixed effects for group (child vs. adult), focus condition, age_child, and their interactions. The final model (Model 7) is represented in Equation (1) with Y i(jk) standing for the outcome variable, i.e., the production score of child j (j = 1, 2, …, J) at age i (i = 1, 2, …, I (jk)) rated by rater k (k = 1, 2, …, K).

    1. (1)
    1. Yi(jk)= β0 + β1 * BFi(jk) + β2 * CFmi(jk) + β3 * NFfi(jk) + β4 * NFii(jk) + β5 * NFmi(jk) + β6 * BFi(jk) * Adulti(jk) + β7 * CFmi(jk) * Adulti(jk) + β8 * NFfi(jk) * Adulti(jk) + β9 * NFii(jk) * Adulti(jk) + β10 * NFmi(jk) * Adulti(jk) + β11 * BFi(jk) * Age_childi(jk) + β12 * CFmi(jk) * Age_childi(jk) + β13 * NFfi(jk) * Age_childi(jk) + β14 * NFii(jk) * Age_childi(jk) + β15 * NFmi(jk) * Age_childi(jk) + [ u 0j + u 1j * Age_childi(jk)+ v 0k+ e i(jk)]

It consisted of two parts: a fixed part and a random part. In the fixed part, the following were estimated: the intercept (β0), the main effect of focus condition (β1 * focus condition to β5 * focus condition), the two-way interaction between the average change in scores in the adults and each focus condition (β6 * focus condition to β10 * focus condition), and the three-way interaction of the average change in scores in the children, focus condition and age (at the test moment) (β11 * focus condition * Age_child to β15 * focus condition * Age_child). The intercept was the average score of a child at the mean age, which was calculated as the average age of all children at the four test moments (83 months or 6 years and 11 months, abbreviated as 6;11) and scaled in units of 10 months. The factor Age_child was centred around the mean age, such that the model’s slope estimated the change in the production scores per 10 months of age with each focus condition.

In the random part (between square brackets), four residual scores were estimated: u 0j, u 1j, v i0 and e (ij)t. The residual scores were assumed to be normally distributed with an expected value of 0.0 and a variance of S2u0j, S2u1j, S2vi0, and S2e(ij)t, respectively (S2 stands for estimated variance). The variance of the first residual (S2u0j or S2child) indicated the variance in scores between children at the mean age. The variance of the second residual (S2u1j or S2age_child) showed the variance in change in scores with age between children; this variance was assumed to depend on their age. These two variance components quantified individual differences in production. The variance of the third residual (S2v0k or S2rater) quantified differences between raters. The variance of the fourth residual (S2e(ij)k) indicated the interaction between child and rater. Note that the variable Age_child was modelled both as a fixed effect (average trend) and as a random slope (individual variation in age-related change). This was necessary because the children belonged to different age groups and were tested at different ages.

The model revealed significant main effects of focus condition (p < .05), group (adults vs. children) (p < .05), and Age_child (p < .05), and a significant interaction between focus condition and Age_child (p < .05). However, no significant group × focus condition interaction was observed, indicating that the relative differences between focus conditions were consistent across groups. Visual inspection of Figure 2 suggests that by the oldest age tested, children’s production scores approached or even exceeded adult performance in all focus conditions. The significant interaction between focus condition and age_child indicated that the children’s production scores differed across conditions at specific ages, and that developmental trajectories varied by condition. As shown in Figure 2, the children scored lower in the NFi and NFm conditions than in the NFf, CFm, and BF conditions at both the youngest and mean ages. This pattern aligns largely with the greater prosodic complexity of focus realisation in non-sentence-final positions in Dutch, where both phonological and phonetic cues are involved. Moreover, the production scores increased more steeply with age in the NFi, NFm, and CFm conditions than in NFf and BF. However, it was unexpected that the development in the BF condition was different from the NFi and NFm conditions but was similar to that in the NFf and CFm conditions, as the marking of BF requires also both phonological and phonetic means, similar to NFi and NFm.

Figure 2
Figure 2

Production scores of the children at each age of test and the adults (the first five panels). The black solid line is the regression line for the children’s data in each focus condition: narrow focus in sentence-initial position (NFi), narrow focus in sentence-final position (NFf), narrow focus in sentence-medial position (NFm), contrastive focus in sentence-medial position (CFm), and broad focus over the whole sentence (BF). The adults’ production scores were displayed for each speaker, as indicated by the abbreviation of each focus condition printed in light grey. The abbreviation of the focus condition in bold represents the mean score of the adults in that focus condition. The regression lines for the children’s data are reproduced in the lower-right panel, together with the adults’ mean in each focus condition.

With respect to individual differences, we found significant variation in the ratings between raters (p < .05), but the variation was the same across focus conditions and children of different ages, indicating high correlations between raters in their ratings. We also found significant variation between children (p < .001), but this variation was similar across focus conditions. The variation between children’s production scores was 0.548 at the mean age. This variance decreased by 0.032 for every 10-month increase in age (p < .005), suggesting that individual differences in prosodic focus marking were larger at younger ages and became more uniform over time, as can be seen in Figure 2.

3.2. Relation between vocabulary size and production scores

To examine the relationship between vocabulary size and the use of prosody in focus marking across age, we used multilevel modelling rather than simple correlation analysis. While correlations can identify associations between two variables, they do not account for potential confounding variables or allow for hypothesis-driven testing of fixed effects. In contrast, multilevel modelling allowed us to test whether variation in children’s PPVT scores explained variation in their production scores across different focus conditions, while accounting for age as a continuous variable and addressing the nested structure of the data through random effects.

We used the same random structure as in our earlier production model, which included four variance components: variance between children in their average production scores at the mean age, variance between children in their age-related change, variance between raters, and residual variance for the interaction between child and rater. For the fixed part of the model, we started with the model with the main effect of focus condition and the interaction between focus condition and Age_child, and extended it stepwise to included vocabulary-related predictors (see Table 1). The outcome variable was the production score each child received on individual test items at each test moment. Model comparisons were based on the –2 log-likelihood (–2LL) statistic, with likelihood ratio tests used to assess whether added predictors significantly improved model fit. The multilevel modelling was conducted using the software MLwiN (version 2.32) (Rasbash et al., 2015).

Table 1

Overview of model comparisons for multilevel modelling on the relation between vocabulary size (PPVT scores) and production scores.

Comparison
Model –2 loglikelihood Models χ2 df p
1. Focus condition + Focus condition * Age_child 14486.21
2. model 1 + PPVT score 14460.30 1 vs. 2 25.92 1 <.001
3. model 2 + PPVT score * Focus condition 14369.46 2 vs. 3 90.84 1 .627
4. model 3 + PPVT score * Age_child 14366.09 3 vs. 4 3.37 1 .001
5. model 4 + PPVT score * Focus condition * Age_child 13956.56 4 vs. 5 409.53 3 <.001

The model that best fit the data included the fixed effect of focus condition, its interaction with Age_child, and a three-way interaction between focus condition, PPVT score, and Age_child (Table 1). This model revealed a significant three-way interaction, indicating that the relationship between vocabulary size and production scores varied by focus condition and age. Specifically, higher PPVT scores were associated with higher production scores in the NFi condition across age, and to a lesser extent in the NFf and CFm conditions at older ages, but not in the NFm and BF conditions.

To estimate how much of the variation in the production scores could be explained by the PPVT scores, we compared the variance components in the random part of the model before and after including the PPVT-related predictors. Since PPVT was a child-level variable, we focused on the variance between children, specifically, the variance between children in their average production scores at the mean age, and the variance between children in their age-related change. We calculated the proportion of variance explained by computing the relative reduction in these components across models. Based on this comparison, the PPVT scores accounted for 9.2% of the variance in the production scores in the NFi condition, 8.6% in NFf, and 6.6% in CFm. These findings suggest that individual differences in vocabulary size are selectively related to prosodic focus marking in specific focus conditions and age ranges.

3.3. Relation between musicality and production scores

We examined the relationship between musicality, as measured by the children’s PMMA scores, and their production scores, using the same multilevel modeling framework described in Section 3.2. The random structure of the model was kept identical, and the fixed part was adjusted by replacing vocabulary_realted predictors with musicality-related predictors. The model that best fit the data included the fixed effect of PMMA score, its interaction with focus condition, an interaction between focus condition and Age_child, and a three-way interaction between focus condition, PMMA score, and Age_child (Table 2). The three-way interaction was significant and indicated that the relationship between musicality and prosodic production varied by both focus condition and developmental stage. Higher PMMA scores were associated with higher production scores in the NFf and CFm conditions, particularly at younger ages, but no such relationship was observed in the NFi, NFm, or BF conditions.

Table 2

Overview of model comparisons for multilevel modelling on the relation between musicality and production scores.

Comparison
Model –2 loglikelihood Models χ2 df p
1. Focus condition * Age_child 16041.523
2. model 1 + PMMA score * Focus condition 16022.037 1 vs. 2 19.486 5 .002
3. model 2 + PMMA score * Focus condition * Age_child 16009.777 2 vs. 3 12.260 5 .031

As with the analysis in Section 3.2, we estimated the proportion of variance in the production scores explained by PMMA by comparing child-level variance components before and after including musicality predictors. The PMMA scores accounted for 8.8% of the variance in the NFf condition and 7.1% in the CFm condition. These findings suggest that individual differences in musicality are selectively related to prosodic focus marking in specific focus conditions and age ranges.

4. Discussion

We conducted a longitudinal study on Dutch-speaking children spanning from 4;1 to 10;3 to address three research questions: (1) what are the individual differences in the use of prosody for focus marking in production among Dutch-speaking school-aged children? (2) Is children’s ability in prosodic focus marking related to their development in vocabulary, another linguistic subsystem? (3) Is this ability related to their development in musicality, another nonlinguistic system? Based on previous work in relevant research areas, we hypothesised that there would be more individual differences in more demanding focus conditions than in less demanding focus conditions, especially at a younger age (Hypothesis I), that children’s ability in prosodic focus marking would be related to some degree to their development in vocabulary, especially at younger ages (Hypothesis II), and that children’s musicality would be related to their ability in prosodic focus marking to different degrees at different ages (Hypothesis III).

The children reached adult-like production in all focus conditions by the age of 10;3, but at different rates in different focus conditions. As found in previous studies on the prosodic focus marking in children acquiring English and Dutch (Chen, 2009, 2011a, 2011b; Romøren, 2016; Wonnacott & Watson, 2008), they attained adult-like production faster and earlier in contrastive focus and sentence-final narrow focus than in sentence-initial narrow focus and sentence-medial narrow focus. The current data also showed that the children acquired adult-like production in broad focus at a similar rate to sentence-final narrow focus. This was an unexpected finding given that the prosodic marking of BF requires also both phonological and phonetic means, similar to NFi and NFm.

Based on differences in prosodic complexity in different focus conditions, we predicted most individual differences in sentence-initial narrow focus, sentence-medial narrow focus, broad focus followed by sentence-final narrow focus, and lastly, contrastive focus. We have found evidence for substantial individual differences in children’s ability to use prosody for focus marking in all focus conditions and a decrease in individual differences with age. But we have observed no evidence for an effect of focus condition on the extent of individual differences. Individual differences were present to a similar degree across focus conditions, consistent with Kidd et al.’s (2018) claim that individual differences are ‘stable across development … and across all domains’. Hypothesis I was thus supported regarding the presence of individual differences, the effect of age but not regarding the effect of focus condition.

Furthermore, our results have brought forth evidence for interconnectedness between vocabulary size and prosodic focus marking in both the contrastive focus condition and the sentence-final narrow focus condition, especially at a younger age, and in sentence-initial narrow focus across age, partially supporting Hypothesis II. Similarly, we have found evidence for interconnectedness between musicality and prosodic focus marking in the contrastive focus condition and the sentence-final narrow focus condition, especially at a younger age, partially supporting Hypothesis III.

It was unexpected that differences in the complexity in the use of prosody between focus conditions did not have an influence on individual differences in production between children, especially considering that the children as a group appeared to become adult-like faster or at an earlier age in some focus conditions than in other focus conditions. This finding may suggest that development in prosodic focus marking is consistent within the same child regardless of prosodic complexity and a complex form-meaning mapping seems to pose similar challenges to different children.

Nevertheless, it raises questions on whether prosodic complexity is a suitable determinant for degree of individual differences in acquiring adult-like prosody. Studies on the acquisition of syntax and morphology have found strong evidence that the relative frequency of different constructions in the input is consistent with the relative order of acquisition of those constructions in the child’s speech (Lieven, 2010; Ambridge et al., 2015). The question arises as to whether the frequency of different focus constructions in the child’s linguistic input can better predict degrees of individual differences in their ability to prosodically mark focus. For example, children may show more variation in their production in the use of prosody that is uncommon in the input, because they hear fewer exemplars of appropriate use of prosody and have less chance to ‘practise’ their use of prosody in those contexts. Future corpus-based studies combined with computational modelling that can both determine the relative frequencies of different focus structures in children’s input and the production of these focus structures in the same children are needed to test out the hypothesised effect of relative frequency of focus structures on degree of individual differences in children’s use of prosody for focus marking.

Furthermore, vocabulary size was connected with prosodic focus marking only in sentence-initial narrow focus, sentence-final narrow focus and contrastive focus conditions. This is another unexpected finding. It would seem to suggest that if there is a common underlying mechanism that facilitates both word learning and prosodic focus marking, this mechanism only operates in these focus conditions. This may sound improbable but not impossible if we relate it to the relative frequency of parental wh-questions in the input children hear.

As discussed in Section 1.2.2, parental wh-questions not only play an important role in young children’s vocabulary development (Oshima-Takane & Titova, 2021 and references therein) but may also facilitate the learning of using appropriate prosody in answers with different focus structures. Frequent exposure to certain question types may give children more opportunities to practice the corresponding prosodic forms. When children respond inappropriately, caregivers may rephrase or ask clarifying questions, providing further feedback and modelling. Research on the acquisition of wh-questions in English-learning 1- to 3-year-olds shows children first acquire what-, where-, and who- questions, then when-, why- and how-questions, and finally which- and whose-questions, corresponding with the relative frequency of wh-questions in the input (Rowland et al., 2003). In our study, the two probably most common types of questions in the input, i.e., who- and what-questions, were used to elicit sentence-initial and final narrow focus respectively. What-happens questions were used to elicit broad focus. They seem to occur infrequently in parental input, as least in the input heard by English-speaking toddlers (Rowland et al., 2003). The what-do-X-do-with-Y type of questions were used to elicit responses with a focus on the verb in sentence-medial narrow focus. They are semantically more similar to how-questions than to what-questions typically asking about sentence object, but are expected to be less common than typical what-questions. Besides asking questions, parents also draw contrasts and make corrections in their interactions with children. Studies qualifying the frequency of these behaviours are still lacking. It remains to be tested whether they occur more frequently than the what-do-X-do-with-Y type of questions in parent-child interactions. Assuming similar relative frequency of wh-questions in Dutch-learning children’s input to that in English-learning children’s input, and a higher relative frequency of corrections-making and contrast-drawing speech acts than the type of wh-question used to elicit sentence-medial narrow focus, the relatively high frequency of who- and what-questions and the speech act of making contrasts can explain why they may facilitate both word learning and prosodic focus marking in sentence-initial and final narrow focus and contrastive focus.

The same input frequency-based account may also explain the coupled development of musicality and prosodic focus marking in some focus conditions but not in other focus conditions. We have found that musicality was connected with prosodic focus marking only in contrastive focus and sentence-final narrow focus. Musicality has been shown to affect adults’ and children’s pitch perception in speech. There is limited evidence for a facilitating effect of musicality on pitch production, i.e., production of lexical tones in Mandarin-speaking 4- and 5-year-olds (Rhee et al., 2022). We may then speculate that musicality on its own may not suffice to facilitate the use of prosody in focus marking in production but it may give children an edge in frequently produced focus structures where they presumably get more opportunities to practise their use of prosody, such as sentence-final narrow focus and contrastive focus. However, this account does not explain the lack of a musicality effect in sentence-initial narrow focus, another possibly common focus structure in the input.

Future research on the frequency of wh-questions and speech acts such as making corrections and drawing contrasts in Dutch-learning children’s linguistic input, especially at the age of 3- to 5-years, is strongly needed to validate the assumptions made in these lines of reasoning and ultimately to test frequency effects in the coupled development between prosodic focus marking and vocabulary and between prosodic focus marking and musicality in specific focus conditions.

Finally, this study has a number of limitations. First, it focused on the interconnectedness between the acquisition of prosodic focus marking and development in other linguistic subsystems and cognitive abilities. While we observed modest evidence for such interconnectedness, vocabulary size and musicality only explained a small amount of variation in the children’s prosodic production. This suggests that there may be other linguistic or non-linguistic factors that contribute more substantially to individual differences in the children’s prosodic focus marking ability. Bishop et al. (2020, 2024) found that variation in pragmatic skills is related to individual differences in the focus-marking related prosodic ability of adult native speakers of English. It can thus be useful to examine the influence of signal-extrinsic factors such as pragmatic skills on individual differences in children’s prosodic focus marking. The current findings may also suggest that our correlational approach is not sensitive enough to investigate a directional influence of vocabulary size and musicality on children’s prosodic focus marking ability. To more rigorously investigate developmental directionality, future research should consider using lagged modelling to assess how changes in vocabulary size and musicality may predict subsequent changes in production over time, or using a pretest-training-posttest paradigm to find out whether musical training can improve children’s ability to use prosody for focus marking.

Second, the children’s use of prosody was rated by trained listeners with expertise in Dutch prosodic focus marking and child language. This approach was chosen to ensure consistent and linguistically grounded evaluations of prosodic function, reducing the influence of irrelevant characteristics of child speech. Ratings were also carried out at a considerable delay after data collection to minimise potential bias from familiarity with individual children. However, it is possible that naïve listeners, lacking both prosodic expertise and experience with child speech, might evaluate the same productions differently. Further research comparing trained and naïve listeners’ assessments could provide insight into how linguistic expertise and familiarity with child speech shape the perception of prosodic competence and communicative effectiveness. Such work might also have implications beyond language development, e.g., in clinical assessment and speech therapies.

Additionally, the current study addressed the research questions using data from Dutch-speaking children. It remains to be investigated whether our findings are generalisable to children acquiring other languages, considering that cross-linguistic differences in prosody and prosodic focus marking are linked with differences in rate and route of acquisition in children acquiring different languages. It is possible that similar findings will be found regarding individual differences in the other languages. The reason is that we found no evidence for an effect of focus condition on the degree of individual differences in spite that children achieve adult-like production at different rates in different focus conditions in Dutch, reminiscent of the cross-linguistic developmental differences reported for the same focus conditions. However, there may be cross-linguistic differences regarding the connection with vocabulary development and musicality under certain circumstances. That is, if the coupled development in prosodic focus marking on the one hand and vocabulary and musicality on the other hand proves to be driven by the relative frequency of focus structures in the input children hear, children acquiring different languages may show different effects of vocabulary size and musicality if the relative frequency of focus structures is different in their input. Future research is needed to examine individual differences in prosodic focus marking by children acquiring typologically different languages.

5. Conclusions

Research on individual differences in morphosyntactic development has advanced by leaps and bounds in recent years. But it is still rare in the area of prosodic development, making it impossible to test theoretical insights into individual differences and shared underlying development mechanisms that have emerged from research in other areas of language development. In the current longitudinal study on Dutch-speaking school-aged children’s prosodic focus marking, we have uncovered evidence for stable individual differences across age and focus conditions, in line with existing findings on morphosyntactic development. We have also identified first evidence for connections between development in prosodic focus marking and development in vocabulary and musicality in the focus conditions that may occur relatively frequently in children’s linguistic input. Future research is strongly needed to understand the postulated frequency effects-related common mechanism underlying the coupled development between prosodic focus marking and vocabulary, and between prosodic focus marking and musicality.

Notes

  1. The pitch accents are transcribed following ToDI (Transcription of Dutch Intonation), a notation developed within the framework of the autosegmental-metrical (AM) theory (Gussenhoven, 2005). In the AM theory (see Arvaniti & Fletcher 2020 for a review), the pitch contour of a sentence is described as a sequence of high (H) and low (L) tones. The domain of transcription is the intonational phrase. The pitch levels at phrase boundaries are transcribed in terms of boundary tones (e.g., %L and %H at the initial boundary, L% and H% at the final boundary). Between the boundary tones, certain words bear a pitch accent and thus sound more prominent than the other words. In ToDI, the pitch accent is described in terms of the tones in the stressed syllable and post-stress syllable(s) of the accented word. H*L (read as ‘high-star-low’) represents a falling pitch contour; !H*L (read as downstepped ‘high-star-low’) indicates a downstepped falling pitch contour (i.e., a falling pitch accent with a lower peak than the preceding pitch accent). The ‘*’ sign indicates which tone is primarily associated with the stressed syllable of the accented word. [^]
  2. The x-word point was defined as the first recording session in which the child’s cumulative number of spontaneously produced, unique, identifiable adult-based words reached x, based on the method used by Vihman et al. (1998) and DePaolis et al. (2008). [^]
  3. Age could not be calculated for three female participants because testing dates were unavailable. All age-related statistics are therefore based on the remaining sample (n = 43). One participant was older (48;0) than the other participants, whose age ranged from 18;8 to 28;8 (n = 42, mean age = 21;0, SD = 2;3). [^]
  4. Section 2.2 is a slightly adapted version of Section 2.2 in Chen and van den Bergh (2025). [^]

Supplementary files

Supplementary file 1. Age and gender statistics of child participants. https://doi.org/10.16995/labphon.19010.s1

Supplementary file 2. Stimuli used in the picture-matching game. https://doi.org/10.16995/labphon.19010.s2

Supplementary file 3. Descriptive statistics for production scores (across focus conditions), vocabulary scores, and musicality scores across testing phases and age groups. https://doi.org/10.16995/labphon.19010.s3

Supplementary file 4. The models built for the analysis on children’s production in Section 3.1. https://doi.org/10.16995/labphon.19010.s4

Data accessibility statement

The material used in the picture-matching game is available at https://doi.org/10.17605/OSF.IO/JHR29. The data supporting the conclusions of this article will be made available upon request by the authors.

Ethics and consent

Ethical review and approval were not required for the study on human participants in accordance with the local legislation and institutional requirements at the time of testing. The participants or the legal guardians provided their written informed consent to participate in this study.

Acknowledgements

A big thank you goes to the children and teaching staff from Houten Montessori Primary School, De Ontdekkingsreis Primary School (Doorn), Soest Montessori Primary School, De Wegwijzer Primary School (Soest) for their indispensable cooperation in this research. We thank Paula Cox (in memoriam) for drawing the pictures and administering the tests, Martine Veenendaal, Saskia Verstegen and Laura Smorenburg for administering the tests, Frank Bijlsma, Alex Manus, Sjef Pieters and Theo Veenker for technical support.

Funding information

The authors declare financial support was received for the research, authorship, and/or publication of this article. This work was supported by a VIDI grant (grant number 276-89-001) and a VICI grant (grant number VI.C.201.109) awarded to AC by the Dutch Research Council (NWO).

Competing interests

The authors have no competing interests to declare.

Author contributions

AC: Conceptualisation, Data curation, Funding acquisition, Investigation, Methodology, Project administration, Resources, Supervision, Validation, Visualisation, Writing – original draft, Writing – review & editing. HB: Formal analysis, Methodology, Resources, Software, Visualisation, Writing – original draft, Writing – review & editing.

References

Ahn, B., Brugos, A., Jeong, S., Shattuck-Hufnagel, S., & Veilleux, N. (2024). Cues to a speaker’s previous beliefs in English intonation. In T. Cho, S. Kim, J. Holliday, & S. Lee-Kim (Eds.), Proceedings of the 19th Conference on Laboratory Phonology (pp. 175–176). https://labphon.org/labphon19/proceedings

Ambridge, B., Kidd, E., Rowland, C. F., & Theakston, A. L. (2015). The ubiquity of frequency effects in first language acquisition. Journal of Child Language, 42(2), 239–273.  http://doi.org/10.1017/S030500091400049X

Arnhold, A. (2016). Complex prosodic focus marking in Finnish: Expanding the data landscape. Journal of Phonetics, 56, 85–109.  http://doi.org/10.1016/j.wocn.2016.02.002

Arnhold, A. (2024). No prosody-syntax trade-offs: Prosody marks focus in Mandarin cleft constructions. Laboratory Phonology, 15(1).  http://doi.org/10.16995/labphon.11515

Arvaniti, A., & Fletcher, J. (2020). The autosegmental-metrical theory of intonational phonology. In C. Gussenhoven & A. Chen (Eds.), The Oxford handbook of language prosody (pp. 78–95). Oxford University Press.  http://doi.org/10.1093/oxfordhb/9780198832232.013.4

Barakat, A. M. M., Mahmoud, B. A. A., & Elmaghraby, R. M. M. (2024). The neural mechanism underlying the effect of musical training on phonological awareness of preschoolers: A meta-analysis. Psycholinguistics, 36(1), 42–69.  http://doi.org/10.31470/2309-1797-2024-36-1-42-69

Bates, E., Bretherton, I., & Snyder, L. (1988). From first words to grammar: Individual differences and dissociable mechanisms. Cambridge University Press.

Bates, E., Dale, P., & Thal, D. (1995). Individual differences and their implications for theories of language development. In P. Fletcher & B. MacWhinney (Eds.), The handbook of child language (pp. 96–151). Basil Blackwell.  http://doi.org/10.1111/b.9780631203124.1996.00005.x

Baumann, S. (2014). Second occurrence focus. In C. Féry & S. Ishihara (Eds.), The Oxford handbook of information structure. Oxford University Press.  http://doi.org/10.1093/oxfordhb/9780199642670.013.38

Baumann, S., & Kügler, F. (2015). Prosody and information status in typological perspective – Introduction to the Special Issue. Lingua, 165(Part B), 179–182.  http://doi.org/10.1016/j.lingua.2015.08.001

Besson, M., Schön, D., Moreno, S., Santos, A., & Magne, C. (2007). Influence of musical expertise and musical training on pitch processing in music and language. Restorative Neurology and Neuroscience, 25, 399–410.  http://doi.org/10.3233/RNN-2007-253423

Bishop, J., & Intlekofer, D. (2020). Lower working memory capacity is associated with shorter prosodic phrases: Implications for speech production planning. In Proceedings of Speech Prosody 2010 (pp. 191–195). https://www.isca-archive.org/speechprosody_2010/

Bishop, J., Kuo, G., & Kim, B. (2020). Phonology, phonetics, and signal-extrinsic factors in the perception of prosodic prominence: Evidence from rapid prosody transcription. Journal of Phonetics, 82, 100977.  http://doi.org/10.1016/j.wocn.2020.100977

Bishop, J., Zhou, C., & Ki, M.-Y. (2024). The perception of prosodic prominence: Continuous or categorical—and for whom? In T. Cho, S. Kim, J. Holliday, & S. Lee-Kim (Eds.), Proceedings of the 19th Conference on Laboratory Phonology (pp. 173–174). https://labphon.org/labphon19/proceedings

Bleses, D., Makransky, G., Dale, P. S., Hølen, A., & Ari, B. A. (2016). Early productive vocabulary predicts academic achievement 10 years later. Applied Psycholinguistics, 37(6), 1461–1476.  http://doi.org/10.1017/S0142716416000060

Boersma, P. (2001). Praat, a system for doing phonetics by computer. Glot International, 5(9–10), 341–345.

Bolduc, J., Gosselin, N., Chevrette, T., & Peretz, I. (2020). The impact of music training on inhibition control, phonological processing, and motor skills in kindergarteners: A randomized controlled trial. Early Child Development and Care, 190(7), 1058–1070.  http://doi.org/10.1080/03004430.2020.1781841

Calhoun, S. (2024). Revisiting how contrast and prominence link: A laboratory phonologist’s view. In T. Cho, S. Kim, J. Holliday, & S. Lee-Kim (Eds.), Proceedings of the 19th Conference on Laboratory Phonology (pp. 169–172). https://labphon.org/labphon19/proceedings

Cardillo, G. C. (2008). Relationships among prosodic sensitivity, musical processing, and phonological awareness in pre-readers. In Proceedings of Speech Prosody 2008 (pp. 595–598).  http://doi.org/10.21437/SpeechProsody.2008-135

Chen, A. (2009). The phonetics of sentence-initial topic and focus in adult and child Dutch. In M. C. Vigário, S. Frota, & M. J. Freitas (Eds.), Phonetics and phonology: Interactions and interrelations (pp. 91–106). John Benjamins.  http://doi.org/10.1075/cilt.306.05che

Chen, A. (2011a). Tuning information packaging: Intonational realization of topic and focus in child Dutch. Journal of Child Language, 38(5), 1055–1083.  http://doi.org/10.1017/S0305000910000541

Chen, A. (2011b). The developmental path to phonological encoding of focus in Dutch. In S. Frota, P. Prieto, & G. Elordieta (Eds.), Prosodic production, perception and comprehension (pp. 93–109). Springer.  http://doi.org/10.1007/978-94-007-0137-3_5

Chen, A. (2018). Get the focus right across languages: Acquisition of prosodic focus-marking in production. In P. Prieto & N. Esteve-Gibert (Eds.), Prosodic development in first language acquisition (pp. 295–314). John Benjamins.  http://doi.org/10.1075/tilar.23.15che

Chen, A., Esteve-Gibert, N., Prieto, P., & Redford, M. (2020). Development in phrase-level prosody from infancy to late childhood. In C. Gussenhoven & A. Chen (Eds.), The Oxford handbook of language prosody. Oxford University Press.  http://doi.org/10.1093/oxfordhb/9780198832232.013.35

Chen, A., & Fikkert, P. (2007). Intonation of early two-word utterances in Dutch. In Proceedings of the 16th International Congress of Phonetic Sciences (pp. 315–320). Saarbrücken, Germany: Universität des Saarlandes. https://www.icphs2007.de/conference/Papers/1757/1757.pdf

Chen, A., & van den Bergh, H. (2025). The production-comprehension relationship in the acquisition of prosodic focus marking: The role of age and individual differences. Languages, 10(9), 234.  http://doi.org/10.3390/languages10090234

Chen, Y., Lee, P.-L., & Pan, H. (2014). Topic and focus marking in Chinese. In The Oxford handbook of information structure (pp. 733–752). Oxford University Press.  http://doi.org/10.1093/oxfordhb/9780199642670.013.34

Chobert, J., François, C., Velay, J. L., & Besson, M. (2014). Twelve months of active musical training in 8- to 10-year-old children enhances the preattentive processing of syllabic duration and voice onset time. Cerebral Cortex, 24(4), 956–967.  http://doi.org/10.1093/cercor/bhs377

Chobert, J., Marie, C., François, C., Schön, D., & Besson, M. (2011). Enhanced passive and active processing of syllables in musician children. Journal of Cognitive Neuroscience, 23, 3874–3887. http://mitprc.silverchair.com/jocn/article-pdf/23/12/3874/1777151/jocn_a_00088.pdf

Coego, S., Esteve-Gibert, N., & Prieto, P. (2025). Preschoolers mark focus types through multimodal prominence: Further evidence for the precursor role of gestures. Languages, 10(5).  http://doi.org/10.3390/LANGUAGES10050092

Cole, J. (2015). Prosody in context: A review. Language, Cognition and Neuroscience, 30(1–2), 1–31.  http://doi.org/10.1080/23273798.2014.963130

Cutler, A., Dahan, D., & van Donselaar, W. (1997). Prosody in the comprehension of spoken language: A literature review. Language and Speech, 40(2), 141–201.  http://doi.org/10.1177/002383099704000203

Dahan, D. (2015). Prosody and language comprehension. Wiley Interdisciplinary Reviews: Cognitive Science, 6(5), 441–452.  http://doi.org/10.1002/wcs.1355

DePaolis, R. A., Vihman, M. M., & Kunnari, S. (2008). Prosody in production at the onset of word use: A cross-linguistic study. Journal of Phonetics, 36(2), 406–422.  http://doi.org/10.1016/j.wocn.2008.01.003

Destruel, E., Lalande, L., & Chen, A. (2024). The development of prosodic focus marking in French. Frontiers in Psychology, 15.  http://doi.org/10.3389/fpsyg.2024.1360308

D’Odorico, L., & Carubbi, S. (2003). Prosodic characteristics of early multi-word utterances in Italian children. First Language, 23(1), 97–116.  http://doi.org/10.1177/0142723703023001005

Duff, F. J., Reen, G., Plunkett, K., & Nation, K. (2015). Do infant vocabulary skills predict school-age language and literacy outcomes? Journal of Child Psychology and Psychiatry, 56(8), 848–856.  http://doi.org/10.1111/jcpp.12378

Engelhardt, P. E., Nigg, J. T., & Ferreira, F. (2013). Is the fluency of language outputs related to individual differences in intelligence and executive function? Acta Psychologica, 144, 424–432.  http://doi.org/10.1016/j.actpsy.2013.08.002

Esteve-Gibert, N., Lœvenbruck, H., Dohen, M., & D’Imperio, M. (2022). Pre-schoolers use head gestures rather than prosodic cues to highlight important information in speech. Developmental Science, 25(1).  http://doi.org/10.1111/desc.13154

Frota, S., Cruz, M., Matos, N., & Vigário, M. (2016). Early prosodic development: Emerging intonation and phrasing in European Portuguese. In M. Armstrong, N. C. Henriksen, & M. M. Vanrell (Eds.), Intonational grammar in Ibero-Romance: Approaches across linguistic subfields (pp. 295–324). John Benjamins.  http://doi.org/10.1075/ihll.6.14fro

Gervain, J., Christophe, A., & Mazuka, R. (2020). Prosodic bootstrapping. In C. Gussenhoven & A. Chen (Eds.), The Oxford handbook of language prosody. Oxford, England: Oxford University Press.  http://doi.org/10.1093/oxfordhb/9780198832232.001.0001

Gleitman, L. R., & Wanner, E. (1982). Language acquisition: The state of the state of the art. In E. Wanner & L. R. Gleitman (Eds.), Language acquisition: The state of the art (pp. 3–48). Cambridge University Press.

Goldstein, H. (2011). Multilevel statistical models (Vol. 922). John Wiley & Sons.  http://doi.org/10.1002/9780470973394

Gordon, E. E. (1979). Developmental music aptitude as measured by the primary measures of music audiation. Psychology of Music, 7, 42–49.  http://doi.org/10.1177/030573567971005

Gregori, A., Sánchez-Ramón, P. G., Prieto, P., & Kügler, F. (2024). Prosodic and gestural marking of focus types in Catalan and German. In Y. Chen, A. Chen, & A. Arvaniti (Eds.), Proceedings of Speech Prosody 2024 (pp. 891–895).  http://doi.org/10.21437/SpeechProsody.2024-180

Gussenhoven, C. (2005). Transcription of Dutch intonation. In S.-A. Jun (Ed.), Prosodic typology: The phonology of intonation and phrasing (pp. 118–145). Oxford University Press.  http://doi.org/10.1093/acprof:oso/9780199249633.003.0005

Hanssen, J., Peters, J., & Gussenhoven, C. (2008). Prosodic effects of focus in Dutch declaratives. In P. A. Barbosa, S. Madureira, & C. Reis (Eds.), Proceedings of Speech Prosody 2008 (pp. 609–612). https://www.isca-archive.org/speechprosody_2008/

Hao, M., Shu, H., Xing, A., & Li, P. (2008). Early vocabulary inventory for Mandarin Chinese. Behavior Research Methods, 40, 728–733.  http://doi.org/10.3758/BRM.40.3.728

Hausen, M., Torppa, R., Salmela, V. R., Vainio, M., & Särkämö, T. (2013). Music and speech prosody: A common rhythm. Frontiers in Psychology, 4, 566.  http://doi.org/10.3389/fpsyg.2013.00566

Ho, Y.-C., Cheung, M.-C., & Chan, A. S. (2003). Music training improves verbal but not visual memory: Cross-sectional and longitudinal explorations in children. Neuropsychology, 17(3), 439–450.  http://doi.org/10.1037/0894-4105.17.3.439

Höhle, B. (2009). Bootstrapping mechanisms in first language acquisition. Linguistics, 47(2), 359–382.  http://doi.org/10.1515/LING.2009.013

Höhle, B., Giesecke, D., & Jusczyk, P. W. (2001). Word-segmentation in a foreign language: Further evidence for crosslinguistic strategies. Poster presented at the Annual Meeting of the Acoustical Society of America, Fort Lauderdale, FL.  http://doi.org/10.1121/1.4777233

Houston, D. M., Jusczyk, P. W., Kuijpers, C., Coolen, R., & Cutler, A. (2000). Cross-language word segmentation by 9-month-olds. Psychonomic Bulletin & Review, 7, 504–509.  http://doi.org/10.3758/BF03214363

Işsever, S. (2003). Information structure in Turkish: The word order-prosody interface. Lingua, 113, 1025–1053.  http://doi.org/10.1016/S0024-3841(03)00012-3

Jansen, N., Harding, E. E., Loerts, H., Başkent, D., & Lowie, W. (2023). The relation between musical abilities and speech prosody perception: A meta-analysis. Journal of Phonetics, 101, 101278.  http://doi.org/10.1016/j.wocn.2023.101278

Jentschke, S., & Koelsch, S. (2009). Musical training modulates the development of syntax processing in children. NeuroImage, 47(2), 735–744.  http://doi.org/10.1016/j.neuroimage.2009.04.090

Jepson, K. (2023). Prosody and word order in marking focus within Djambarrpuyŋu noun phrases. In R. Skarnitzl & J. Volín (Eds.), Proceedings of the 20th International Congress of Phonetic Sciences – ICPhS 2023. https://www.internationalphoneticassociation.org/icphs-proceedings/ICPhS2023/full_papers/310.pdf.

Jusczyk, P. W. (1997). The discovery of spoken language. Cambridge, MA: MIT Press.

Kapia, E., & Kleber, F. (2024). Exploring prosodic marking of information structure in child and adult Albanian. In Y. Chen, A. Chen, & A. Arvaniti (Eds.), Proceedings of Speech Prosody 2024 (pp. 11–15).  http://doi.org/10.21437/SpeechProsody.2024-3

Kidd, E. (2012). Individual differences in syntactic priming in language acquisition. Applied Psycholinguistics, 33, 393–418.  http://doi.org/10.1017/S0142716411000415

Kidd, E., Bigood, A., Donnelly, S., Durrant, S., Peter, M. S., & Rowland, C. F. (2020). Individual differences in first language acquisition and their theoretical implications. In C. F. Rowland, A. L. Theakston, B. Ambridge, & K. E. Twomey (Eds.), Current perspectives on child language acquisition: How children use their environment to learn (pp. 189–219). John Benjamins.  http://doi.org/10.1075/tilar.27.09kid

Kidd, E., & Donnelly, S. (2020). Individual differences in first language acquisition. Annual Review of Linguistics, 6, 319–340.  http://doi.org/10.1146/annurev-linguistics-011619

Kidd, E., Donnelly, S., & Christiansen, M. H. (2018). Individual differences in language acquisition and processing. Trends in Cognitive Sciences, 22, 154–169.  http://doi.org/10.1016/j.tics.2017.11.006

Kügler, F., & Calhoun, S. (2020). Prosodic encoding of information structure. In C. Gussenhoven & A. Chen (Eds.), The Oxford handbook of language prosody (pp. 453–467). Oxford University Press.  http://doi.org/10.1093/oxfordhb/9780198832232.013.30

Lambrecht, K. (1994). Information structure and sentence form: Topics, focus, and the representations of discourse referents. Cambridge University Press.  http://doi.org/10.1017/CBO9780511620607

Lieven, E. (2010). Input and first language acquisition: Evaluating the role of frequency. Lingua, 120(11), 2546–2556.  http://doi.org/10.1016/j.lingua.2010.06.005

Linnavalli, T., Putkinen, V., Lipsanen, J., Huotilainen, M., & Tervaniemi, M. (2018). Music playschool enhances children’s linguistic skills. Scientific Reports, 8, 8767.  http://doi.org/10.1038/s41598-018-27126-5

Liu, Z., & Wang, M. (2024). Prosodic focus marking in Wa. In Y. Chen, A. Chen, & A. Arvaniti (Eds.), Proceedings of Speech Prosody 2024 (pp. 777–781).  http://doi.org/10.21437/SpeechProsody.2024-157

Magne, C., Schön, D., & Besson, M. (2006). Musician children detect pitch violations in both music and language better than nonmusician children: Behavioral and electrophysiological approaches. Journal of Cognitive Neuroscience, 18, 199–211.  http://doi.org/10.1162/jocn.2006.18.2.199

Marchman, V. A., & Fernald, A. (2008). Speed of word recognition and vocabulary knowledge in infancy predict cognitive and language outcomes in later childhood. Developmental Science, 11, 9–16.  http://doi.org/10.1111/j.1467-7687.2008.00671.x

Moreno, S., Marques, C., Santos, A., Santos, M., Castro, S. L., & Besson, M. (2009). Musical training influences linguistic abilities in 8-year-old children: More evidence for brain plasticity. Cerebral Cortex, 19, 712–723.  http://doi.org/10.1093/cercor/bhn120

Morgan, J., & Demuth, K. (1996). Signal to syntax: Bootstrapping from speech to grammar in early acquisition. Erlbaum.

Myrberg, S., & Riad, R. (2014). On the expression of focus in the metrical grid and in the prosodic hierarchy. In C. Féry & S. Ishihara (Eds.), The Oxford handbook of information structure (pp. 441–462). Oxford University Press.  http://doi.org/10.1093/oxfordhb/9780199642670.013.41

Nan, Y., Liu, L., Geiser, E., Shu, H., Gong, C. C., Dong, Q., Gabrieli, J. D. E., & Desimone, R. (2018). Piano training enhances the neural processing of pitch and improves speech perception in Mandarin-speaking children. Proceedings of the National Academy of Sciences, 115(28), E6630–E6639.  http://doi.org/10.1073/pnas.1808412115

Nazzi, T., Bertoncini, J., & Mehler, J. (1998). Language discrimination by newborns: Toward an understanding of the role of rhythm. Journal of Experimental Psychology: Human Perception and Performance, 24, 756–766.  http://doi.org/10.1037/0096-1523.24.3.756

Nazzi, T., Iakimova, G., Bertoncini, J., Frédonie, S., & Alcantara, C. (2006). Early segmentation of fluent speech by infants acquiring French: Emerging evidence for crosslinguistic differences. Journal of Memory and Language, 54, 283–299.  http://doi.org/10.1016/j.jml.2005.10.004

Norbury, C. F. (2019). Individual differences in language acquisition. In J. Horst & J. von Koss Torkildsen (Eds.), International handbook of language acquisition (1st ed., pp. 323–340). Routledge.  http://doi.org/10.4324/9781315110622-17

Oshima-Takane, Y., & Titova, P. (2021). Changes in parental input patterns of wh-questions. In D. Dionne & L. Vidal Covas (Eds.), Proceedings of the 45th Annual Boston University Conference on Language Development (pp. 598–611). Cascadilla Press.

Patel, A. D. (2007). Music, language, and the brain. Oxford University Press.  http://doi.org/10.1093/acprof:oso/9780195123753.001.0001

Paul, R., Simmons, E. S., & Mahshie, J. (2020). Prosody in children with atypical development. In C. Gussenhoven & A. Chen (Eds.), The Oxford handbook of language prosody (pp. 586–593). Oxford University Press.  http://doi.org/10.1093/oxfordhb/9780198832232.013.49

Pino, M. C., & Giancola, M. (2023). The association between music and language in children: A state of the art review. Children, 10(5), 801.  http://doi.org/10.3390/children10050801

Portes, C., & Reyle, U. (2022). Combining syntax and prosody to signal information structure: The case of French. In Proceedings of Speech Prosody 2022 (pp. 87–91).  http://doi.org/10.21437/SpeechProsody.2022-18

Prieto, P., Estrella, A., Thorson, J., & Vanrell, M. del M. (2012). Is prosodic development correlated with grammatical and lexical development? Evidence from emerging intonation in Catalan and Spanish. Journal of Child Language, 39(2), 221–257.  http://doi.org/10.1017/S030500091100002X

Quené, H., & van den Bergh, H. (2008). Examples of mixed-effects modeling with crossed random effects and with binomial data. Journal of Memory and Language, 59, 413–425.  http://doi.org/10.1016/j.jml.2008.02.002.

Ramus, F. (2002). Language discrimination by newborns: Teasing apart phonotactic, rhythmic, and intonational cues. Annual Review of Language Acquisition, 2, 85–115.  http://doi.org/10.1075/arla.2.05ram

Ramus, F., Hauser, M. D., Miller, C., Morris, D., & Mehler, J. (2000). Language discrimination by human newborns and by cotton-top tamarin monkeys. Science, 288, 349–351.  http://doi.org/10.1126/science.288.5464.349

Ramus, F., Nespor, M., & Mehler, J. (1999). Correlates of linguistic rhythm in the speech signal. Cognition, 73, 265–292.  http://doi.org/10.1016/S0010-0277(00)00101-3

Repp, S. (2014). Contrast: Dissecting an elusive information-structural notion and its role in grammar. In C. Féry & S. Ishihara (Eds.), The Oxford handbook of information structure (pp. 270–289). Oxford University Press.  http://doi.org/10.1093/oxfordhb/9780199642670.013.006

Rhee, N., Chen, A., & Kuang, J. (2022). Musicality and age interaction in tone development. Frontiers in Neuroscience, 16, 804042.  http://doi.org/10.3389/fnins.2022.804042

Roberts, C. (1996). Information structure in discourse: Towards an integrated formal theory of pragmatics. In J. H. Yoon & A. Kathol (Eds.), OSU Working Papers in Linguistics (pp. 91–136). Ohio State University.

Roessig, S., Taheri-Ardali, M., Pagel, L., & Mücke, D. (2024). Prosodic realization of different focus types in Persian. In Y. Chen, A. Chen, & A. Arvaniti (Eds.), Proceedings of Speech Prosody 2024 (pp. 762–766).  http://doi.org/10.21437/SpeechProsody.2024-154

Roettger, T. B., Mahrt, T., & Cole, J. (2019). Mapping prosody onto meaning—the case of information structure in American English. Language, Cognition and Neuroscience, 34(7), 841–860.  http://doi.org/10.1080/23273798.2019.1587482

Romøren, A. S. H. (2016). Hunting highs and lows: Acquisition of prosodic focus marking in Dutch and Swedish [Doctoral dissertation, Utrecht University]. Utrecht University Repository. https://dspace.library.uu.nl/handle/1874/338040

Romøren, A. S. H., & Chen, A. (2015). Quiet is the new loud: Pausing and focus in child and adult Dutch. Language and Speech, 58(1), 8–23.  http://doi.org/10.1177/0023830914563589

Romøren, A. S. H., & Chen, A. (2022). The acquisition of prosodic marking of narrow focus in Central Swedish. Journal of Child Language, 49(2), 213–238.  http://doi.org/10.1017/S0305000920000847

Rowe, M. L., Leech, K. A., & Cabrera, N. (2017). Going beyond input quantity: Wh-questions matter for toddlers’ language and cognitive development. Cognitive Science, 41, 162–179.  http://doi.org/10.1111/cogs.12349

Rowland, C. F., Pine, J. M., & Lieven, E. V. M. (2003). Determinants of acquisition order in wh-questions: Re-evaluating the role of caregiver speech. Journal of Child Language, 30, 609–635.  http://doi.org/10.1017/S0305000903005695

Schlichting, L. (2005). Peabody Picture Vocabulary Test–III–NL (PPVT–III–NL). Harcourt Test Publishers.

Scollon, R. (1979). A real early stage: An unzippered condensation of a dissertation on child language. In E. Ochs & E. Schieffelin (Eds.), Developmental pragmatics. Academic Press.

Skopeteas, S. (2014). Information structure in Modern Greek. In C. Féry & S. Ishihara (Eds.), The Oxford handbook of information structure (pp. 686–708). Oxford University Press.  http://doi.org/10.1093/oxfordhb/9780199642670.013.15

Snijders, T., & Bosker, R. (1999). Multilevel analysis: An introduction to basic and applied multilevel analysis. Sage.

Snow, D. (2000). The emotional basis of linguistic and nonlinguistic intonation: Implications for hemispheric specialization. Developmental Neuropsychology, 17(1), 1–28.  http://doi.org/10.1207/S15326942DN1701_01

Snow, D. (2006). Regression and reorganization of intonation between 6 and 23 months. Child Development, 77(2), 281–296.  http://doi.org/10.1111/j.1467-8624.2006.00870.x

Snow, D., & Balog, H. L. (2002). Do children produce the melody before the words? A review of developmental intonation research. Lingua, 112(12), 1025–1058.  http://doi.org/10.1016/S0024-3841(02)00060-8

Tomioka, S. (2014). Information structure in Japanese. In C. Féry & S. Ishihara (Eds.), The Oxford handbook of information structure (pp. 753–773). Oxford University Press.  http://doi.org/10.1093/oxfordhb/9780199642670.013.42

Truckenbrodt, H. (2014). Focus, intonation, and tonal height. In C. Féry & S. Ishihara (Eds.), The Oxford handbook of information structure (pp. 463–482). Oxford University Press.  http://doi.org/10.1093/oxfordhb/9780199642670.013.44

Türk, O., & Calhoun, S. (2024). Phrasal synchronization of gesture with prosody and information structure. Language and Speech, 67(3), 702–743.  http://doi.org/10.1177/00238309231185308

Vallduví, E., & Engdahl, E. (1996). The linguistic realization of information packaging. Linguistics, 34(3), 459–520.  http://doi.org/10.1515/ling.1996.34.3.459

Vasilyeva, M., Waterfall, H., & Huttenlocher, J. (2008). Emergence of syntax: Commonalities and differences across children. Developmental Science, 11, 84–97.  http://doi.org/10.1111/j.1467-7687.2007.00656.x

Verhoef, E., Allegrini, A. G., Jansen, P. R., Lange, K., Wang, C. A., Morgan, A. T., Ahluwalia, T. S., Symeonides, C., Andreassen, O. A., Bartels, M., Boomsma, D., Dale, P. S., Ehli, E., Fernandez-Orth, D., Guxens, M., Hakulinen, C., Harris, K. M., Haworth, S., de Hoyos, L., … St. Pourcain, B. (2024). Genome-wide analyses of vocabulary size in infancy and toddlerhood: Associations with attention-deficit/hyperactivity disorder, literacy, and cognition-related traits. Biological Psychiatry, 95(9), 859–869.  http://doi.org/10.1016/j.biopsych.2023.11.025

Vidal, M. M., Lousada, M., & Vigário, M. (2020). Music effects on phonological awareness development in 3-year-old children. Applied Psycholinguistics, 41, 299–318.  http://doi.org/10.1017/S0142716419000535

Vihman, M. M., DePaolis, R. A., & Davis, B. L. (1998). Is there a “trochaic bias” in early word learning? Evidence from infant production in English and French. Child Development, 69(4), 935–949.  http://doi.org/10.2307/1132354

Wang, L., Bastiaansen, M., Yang, Y., & Hagoort, P. (2011). The influence of information structure on the depth of semantic processing: How focus and pitch accent determine the size of the N400 effect. Neuropsychologia, 49, 813–820.  http://doi.org/10.1016/j.neuropsychologia.2010.12.035

Welch, G. F. (1998). Early childhood musical development. Research Studies in Music Education, 11, 27–41.  http://doi.org/10.1177/1321103X9801100104

Wells, B., Peppé, S., & Goulandris, N. (2004). Intonation development from five to thirteen. Journal of Child Language, 31, 749–778.  http://doi.org/10.1017/S0305000904006449

Wieman, L. A. (1976). Stress patterns of early child language. Journal of Child Language, 3(2), 283–286.  http://doi.org/10.1017/S0305000900001501

Wonnacott, E., & Watson, D. G. (2008). Acoustic emphasis in four year olds. Cognition, 107(3), 1093–1101.  http://doi.org/10.1016/j.cognition.2007.10.005

Yang, A., & Chen, A. (2018). The developmental path to adult-like prosodic focus-marking in Mandarin Chinese-speaking children. First Language, 38(1), 26–46.  http://doi.org/10.1177/0142723717733920

Yang, A., Cho, T., Kim, S., & Chen, A. (2024). Prosodic focus marking in Seoul Korean-speaking children: The use of prosodic phrasing. Frontiers in Psychology, 15.  http://doi.org/10.3389/fpsyg.2024.1352280