1. Introduction
1.1. Prosodic prominence
Prosodic prominence refers to the relative strength of a phonological constituent, a basic aspect of the structure and phonetic realization of speech as well as sign. On the speech side, the last twenty years have seen a flourishing of research into both how speakers encode prosodic prominence and how listeners perceive it. The present study is primarily about the latter, for which there are extensive relevant reviews available (e.g., Bishop et al., 2020; Cole et al, 2019; Bauman & Winter, 2018; and Turnbull et al., 2017; see also the recent discussions in Bruggeman et al., 2025, and in Lorenzen, 2024). The study we present here, which is a new analysis of a data set reported in previous work (Bishop et al., 2020), focuses more specifically on the role played by listeners themselves—i.e., individual differences. In what follows, we briefly review the aims and findings of Bishop et al. (2020) and related work (sections 1.2-1.5); we then describe the goals of the present study (Section 2.1) and then present the study itself and its findings (sections 2.2-2.3). Finally, we discuss the findings and their implications (Section 3) before making some concluding remarks (Section 4).
1.2. Phonological cues in prominence perception
Most work investigating how listeners perceive prosodic prominence has, unsurprisingly, addressed this phenomenon in languages for which prominence is a structural, i.e., phonological, aspect of the language (see Jun, 2005, 2014, 2025; and Ladd & Arvaniti, 2023, for the role of prominence in prosodic typology). But even more narrowly, the majority of this work has involved how listeners perceive prosodic prominence in English, German, and Dutch—prototypical “stress languages,” where prominence plays a central role in both word and phrasal prosody. While there have been a number of important outcomes of this work, we see two developments as driving the biggest advances in knowledge in recent years. The first is the use of sophisticated statistical modeling of speech data (data from both experimental tasks and from corpora) to identify the relevant signal-based (acoustic/bottom-up) cues that listeners respond to, such as duration, intensity/loudness, and F0, with a number of other factors contributing in smaller ways. The second major advance is the inclusion of phonological variables in the modeling of prominence perception, which, notably, reflects a kind of analysis of the signal that (for now) still relies largely on the work of human annotators whose decisions are based on a particular theory of how acoustic signals are linguistically structured.
Bishop et al. (2020) discuss a finding in the literature that demonstrates, in their view, the importance of this second development in particular. That finding is the one reported by Kochanski et al. (2005), who attempted to predict English-speaking listeners’ perceptual judgments based on purely acoustic cues. Their conclusion, summarized in the title of their widely cited paper, “Loudness predicts prominence; fundamental frequency adds little,” is bafflingly at odds with frameworks like Autosegmental Metrical theory that assume, for the relevant language types, F0 to be the primary correlate of phrase-level prominence (Beckman, 1986). While sophisticated in a technical sense, the problem with the “raw acoustics” approach exemplified by Kochanski and colleagues is that, in languages like English, some F0 values are prominence lending (namely, those associated with accentual targets, i.e., intonational pitch accents) and some are not (i.e., those reflecting interpolation between targets, or reflecting targets that are demarcative rather than accentual), and so a model that cannot distinguish between them will surely find F0 to be a weak predictor of prominence, at best. Moreover, even if we consider only prominence-lending F0s, those values also vary: Pitch accents can be (and in English are) specified for values that span the speaker’s range. That is, accentual targets for F0 include high F0s—reflecting the presence of tones like H* and L+H*—but also lower-level tones like !H* and L*, which are distinctive for their lowness or non-highness rather than their highness (Pierrehumbert, 1980), but they nonetheless correlate with a word’s status as structurally/metrically prominent. Reasonable assessment of F0’s role in perceiving prominence will therefore not be possible without first identifying the phonological status of the F0 value in question. And, to be sure, in analyses that do this (either by holding other cues constant or by including phonological information in the models), F0 is found to be a highly important or even the single most important predictor of listeners’ perception of prominence in various listening tasks (Bishop et al., 2020; Cole et al., 2019; Im et al., 2023, and references therein; see also Baumann & Winter, 2018, and Roessig et al., 2022, for German).
We assume, as did Bishop et al. (2020), that human listeners are similarly unlikely to perceive prominence in terms of raw acoustic values as such. We think a better assumption is that the listener first parses the raw signal into a prosodic structure and only after parsing this prosodic structure interprets the prominence of the residual acoustic phonetic variation. For example, if an English speaker produces a H* accent with an F0 that is significantly high in her range—higher than would be the minimum required to parse it as marking a H* rather than a !H*—the listener will then interpret the intended meaning of that heightened peak. However, this takes place only after identifying that peak as a prominence-lending one to begin with—i.e., one that is aligned with a lexically stressed syllable. Although surely more complex than our example of F0 here, presumably a similar sequence of events applies to other phonetic cues to prominence. Duration and loudness, for example, both contribute to an English-speaking listener’s parsing of an utterance’s intended metrical structure. But the primary point here is that the listener will perceive and assign meaning to any residual duration or loudness only in the context of a metrical structure.
1.3. Phonetic cues in prominence perception
Unlike an utterance’s accentual (i.e., pitch accent) structure, which largely conveys its information structure and syntactic structure, it is less clear how best to characterize the interpretation that listeners assign to gradient phonetic cues that are prominence-lending above and beyond what is used to establish a basic metrical/accentual structure—what we will refer to as “subaccentual” cues. But we know that such variation is under the speaker’s control. To take a case that has been of much interest to researchers over the years, consider the prosodic marking of so-called ambiguous focus structures, i.e., subject-verb-object (SVO) sentences in English (or German or Dutch, where the patterns are similar). When the nuclear pitch accent (i.e., the last accent) falls on the object in these constructions, such as in the utterance John bought a MOTORcycle, they are ambiguously felicitous for both intended narrow object focus (e.g., if uttered as an answer to the question, What did John buy?) or broader focus on the verb (if uttered in answer to What did John do?), a phenomenon long accounted for by theories of Focus Projection (Selkirk, 1984, 1995; see also Gussenhoven, 1983, 1999, and Büring, 2006). Phonetically, however, we know that speakers produce a number of phonetic cues above and beyond that needed to convey the intended accentual structure. For example, they may select different nuclear pitch accent types, they may increase the object’s duration and/or intensity, or they may adjust (relative to prenuclear material) the scaling of the nuclear accent’s peak (see Bishop, 2012b, and Roessig, 2024, for reviews). And we now know that focus also results in subtle, subaccentual articulatory adjustments of a number of kinds (for relevant data in German, see Baumann et al., 2006; Pagel et al., 2024; and Mücke & Grice, 2014. See also Jang & Katsika, 2024, for Korean).
On the listener’s side, it is also clear that the kind of subaccentual variation just described is perceptually relevant, as listeners seem to interpret such cues as conveying focus, though not in the clear-cut way that they interpret coarse-grained information about pitch accent placement. To continue with the example of phonologically ambiguous focus marking in SVOs (but see Mitterer et al., 2024, for recent discussion of another case), we know that listeners are more likely to recover intended narrow focus on an object when that object has prominence-lending phonetic enhancements, as shown in a study by Breen et al. (2010), who found that listeners could recover (above chance level, at least) the question contexts that a particular SVO sentence was intended to answer. Compelling evidence of the same sort was also reported for German by Grice et al. (2017), who emphasize the implications for phonological modeling. However, it is not only the case that listeners can use prominence-lending cues in this way when they are present—they also have expectations for them when they are absent, and these expectations are strong enough to form the basis of an illusion. This was shown by Bishop (2012b), who found English-speaking listeners to judge the verb in SVO sentences like John bought a MOTORcycle as less prominent, and the object as more prominent, when the sentence was presented in the context of a narrow focus rather than broad focus question context—even though the SVO sentence production that they judged was the same (via splicing) across question contexts.
At the same time, however, it is also apparent that this sensitivity to phonetic prominence is subject to considerable individual differences. For example, Bishop (2016) found the strength of the top-down illusory effect just described to vary significantly across listeners, suggesting that not all listeners have strong expectations for the cues in question. Similarly, Breen et al. (2010) found that a little more than a quarter of their listeners were unable to recover narrow object focus based on phonetic cues, and about half were unable to recover intended broad focus. Thus, while we have little evidence that listeners show significant variation in their recovery of focus from accentual structure, there appears to be significant variation in their ability to recover focus from subaccentual phonetic cues. Finally, while we have used the broad versus narrow focus ambiguity as an example here, the distinction between contrastive versus non-contrastive focus in English and other West Germanic languages seems to present a similar situation.
1.4. Variation in sensitivity to subaccentual cues to prominence
Apparent in the discussion so far is that modeling the systematic use of phonetic cues to prominence, while still regarding them as outside of the phonological system, has been an ongoing challenge in work on the relation between prosody and focus. One possibility is that mainstream models of the phonology/phonetics divide lack an appropriate mechanism for incorporating gradience, and we note that there are innovative and promising proposals for addressing this (e.g., Roessig et al., 2019; see also the important discussion in Grice et al., 2017, cited above). However, another approach, perhaps more conservative, is to assume that the phonetic adjustments in question do not actually reflect focus marking per se, but instead reflect emphasis, i.e., paralanguage. Arguments along these lines have been made a number of times over the years. Ladd (1996), for example, discussed the role of emphasis in cueing focus as one example of the major challenge facing the intonational analyst, namely teasing apart language from paralanguage. Hayes (1994), arguing against the idea of an emphatic pitch accent category in English, described how subaccentual cues could have a prominence-lending effect, but in a manner more akin to gesture—i.e., gesture that is vocal in nature and aligned with the accentual structure of an utterance, just as gesture can be aligned with other aspects of an utterance’s linguistic organization (Kendon, 1972, 1975). Gussenhoven (2015) offers a similar explanation, situated in a larger discussion about how listeners approach experimental tasks that involve making explicit (and subjective) prominence judgments. We understand his basic claim to be as follows:
Phonetic cues (indeed, any kind of cues) will increase perceived prominence if they increase perceived emphasis
Increased emphasis, in turn, will cue narrowness of focus because focus and emphasis are pragmatically linked
To us, the appeal of Gussenhoven’s claim (like Hayes’s) is that it offers a possible explanation for why phonetic cues to prominence, like the ones that seem to disambiguate broad and narrow focus in SVOs, are simultaneously very systematic yet also listener-specific: because decoding them is a matter of interpreting speaker intentions. Unlike genuine grammatical cues related to accentuation, for which sensitivity should be (by definition, we think) uniform in a language community, cues whose interpretation is pragmatically dependent are expected to vary considerably from listener to listener, based on factors like the strength of the context or—as is our primary interest here—properties of the listeners themselves, i.e., individual differences.
1.5. Individual differences related to pragmatic skill
Over the past several years, experimental work in psycholinguistics and in laboratory phonology has reported effects of individual differences related to what we tentatively refer to, after others, as “pragmatic skill.” Pragmatic skill, or the propensity to perceive and process linguistic information in relation to its context, and to attend closely to speaker intentions, varies in any population of neurotypical individuals—with measurable consequences for performance in experimental tasks. For example, there is mounting evidence that measures used to assess constructs like Theory of Mind and empathy (two constructs that, in our view, are not well distinguished in the literature) predict the extent to which pragmatic processing is relied upon during sentence comprehension (e.g., Nieuwland et al., 2010; Van den Brink et al., 2012; Xiang, et al., 2013; Ferguson et al., 2015; Kulakova & Nieuwland, 2016). Such measures include questionnaires like the Empathy Quotient (Baron-Cohen & Wheelwright, 2004) and the Autism Spectrum Quotient (Baron-Cohen, Wheelwright, Skinner, et al., 2001) or, more specifically, the communication subscale of the Autism Spectrum Quotient, which deals specifically with social communication skills rather than the broader set of cognitive traits associated with autism spectrum conditions (such as attention to detail or attention switching) and for which there is some evidence for a neural correlate in neurotypical groups (Jouravlev et al., 2020). Yet another measure is the Reading the Mind in the Eyes task (Baron-Cohen, Wheelwright, Hill, et al., 2001), which assesses emotion recognition from partially obscured photographs of faces and is presented by the authors as a measure of Theory of Mind/mentalizing.
Important to our purposes here, a number of studies have argued that pragmatic skill on one or another of these measures is associated with heightened sensitivity to prosodic prominence in perception and processing (Bishop, 2012a, 2016, 2017; Jun & Bishop, 2015a, 2015b; Bishop et al., 2020; Orrico et al., 2025; see also Turnbull, 2014, 2019, and Bishop et al., 2022, for some evidence from production). For example, Bishop (2016) found that listeners who scored higher on the Reading the Mind in the Eyes task (indicating better pragmatic skill) were more sensitive to durational cues when making prominence judgments, and listeners with stronger pragmatic skill according to the Autism Spectrum Quotient’s communication subscale were more susceptible to a discourse-based prominence illusion. Additionally, Bishop et al. (2020), in the context of a large model that included phonetic, phonological, and other factors, found that low pragmatic skill according to this measure was associated with weaker sensitivity to prominence for words of relatively weak metrical strength (unaccented and prenuclear accented words). And in their attempt to prime prosodic phrasing patterns in an implicit prosodic task (i.e., one involving silent reading), Jun and Bishop (2015b) argued that accentuation patterns were less primable in readers with low levels of pragmatic skill according to the communication subscale of the Autism Spectrum Quotient.
Most recently, Orrico et al. (2025) show that scores on such measures were predictive of effects of meaning and phonetic detail on prominence ratings by native-speaking listeners of Southern British English. In particular, they show that individuals with better pragmatic skill according to the Empathy Quotient were more influenced by top-down effects of meaning (namely, a word’s being used contrastively) than less pragmatically-skilled listeners—a fairly intuitive finding and one that is anticipated by the authors. Orrico and colleagues also found that words bearing a H* pitch accent and words bearing a L+H* accent were well separated in terms of their perceived prominence, suggesting that this phonological contrast in accent type, which is also associated with a phonetic difference in (pitch) prominence, is a predictor of prominence perception, at least at the group level. Looking at individual differences, however, Orrico and colleagues found that listeners with higher overall Autism Spectrum Quotient scores (indicating more autistic-like cognitive styles, based on all five traits measured by this instrument) showed a smaller perceptual difference in the perceived prominence of words bearing H* and L+H*, due mostly to increased perceptual prominence of H* for listeners with the most severe autistic traits.
Given that Orrico and colleagues (2025) used the composite Autism Spectrum Quotient scores, which encompass a number of dimensions, not just the communicative/pragmatic one, it is difficult to know whether the effects they find are driven by pragmatic skill, as we are interested in here, or by other dimensions assessed by the Autism Spectrum Quotient—such as attention to detail, attention switching, etc. But to the extent that they might reflect pragmatic skill in particular, their finding suggests that the phonetic differences between H* and L+H* might matter less for listeners with weaker pragmatic skill. While this does not preclude an influence from the kinds of meaning conveyed by these accents, it is consistent with the idea presented above, namely that phonetic cues to prominence convey information about emphasis, which requires sensitivity to speaker intentions to decode. This is what we explore in the study we now turn to.
2. Present study: Individual differences in Rapid Prosody Transcription
2.1. Overview
2.1.1. Bishop, Kuo, and Kim (2020)
The study we present below consists of a re-analysis of data from Bishop et al. (2020), a prominence perception experiment that used Rapid Prosody Transcription (RPT), a task in which linguistically naïve listeners must make speeded identifications of coarse prosodic events—prominence and juncture—from recordings of running speech (Cole, Mo, & Hasegawa-Johnson, 2010; Cole, Mo, & Baek, 2010). Important to point out here is that while Bishop et al. (2020) tested for a main effect of scores on the Autism Spectrum Quotient’s communication subscale, this was in the context of a larger model of prominence perception, and they did not carry out any analyses that allowed for interactions between measures of pragmatic skill and other predictor variables; this is the basic contribution of the present study. Given our reuse of this data set, this time with the primary goal of investigating individual differences, we provide a brief overview of Bishop et al. (2020) and its goals.
Bishop et al. (2020) used the RPT task to probe prominence perception in a large set of listeners (N = 158) who took part in both the prominence and juncture identification tasks, although, as in Bishop et al. (2020), only the prominence identification task is of interest here. Importantly, because that study, like the present one, was interested in the role that both phonetic and phonological cues play in prominence perception by English-speaking listeners, the speech materials presented to listeners as stimuli in the study were transcribed for their phonological structure using the Tones and Break Indices (ToBI) conventions for Mainstream American English (Beckman & Hirschberg, 1994; Beckman & Ayers Elam, 1997), which are based on the Autosegmental-Metrical (AM) model of the language’s intonational phonology, developed in Pierrehumbert (1980) and Beckman and Pierrehumbert (1986). The overarching goal of Bishop and colleagues’ (2020) study was to better understand how listeners use phonetic cues (for the purposes of prominence perception) in a phonologically structured way.
First, as also found earlier by Cole et al. (2019), phonology was, on its own, an important predictor. For example, Bishop and colleagues (2020) found that both pitch accent status and pitch accent type were important factors, and in the expected direction. That is, other things being equal, nuclear accented words were perceived as most prominent, prenuclear accented words as less prominent, and unaccented words as least prominent. Similarly, but with the effect sizes being smaller than for accent status, accent type also mattered: Words bearing L+H* were perceived as most prominent, followed by words bearing H*, and words bearing !H* were least prominent, at least numerically (Bishop and colleagues found the significance of the !H* ~ H* difference to depend on accent status, significant for nuclear accented words but not for prenuclear accented words).
Moreover, and in line with the criticisms of Kochanski and colleagues’ (2005) approach that we expressed above, Bishop and colleagues (2020) also found listeners’ sensitivity to phonetic cues to be quite dependent on phonology, a central theme of their study. For example, the extent to which phonetic F0 values were predictive of perceived prominence depended considerably on a word’s pitch accent status: Other things being equal, F0 was as much as twice as important to predicting prominence judgments for accented words as it was for unaccented words. And relevant to our purposes here, listeners’ sensitivity to phonetic cues like F0 and duration varied as a function of pitch accent type as well. One example of this kind of asymmetric sensitivity is summarized in Figure 1, based on a similar figure in Bishop et al. (2020), which illustrates the outsized attention listeners paid to duration and especially F0 height when the word being judged bore a L+H*. In the case of F0, this is also highly rational, since F0 peak height can vary significantly more (and thus should be potentially more informative) for L+H* and H* than it can for !H*. This is because F0 height for !H* is limited phonologically (i.e., it can only be so high before it becomes, by definition, a H*).
Figure 1: Effect sizes for increases in phonetic duration and F0 peak height on the perceived prominence of words bearing !H*, H*, and L+H* in Bishop et al.’s (2020) Rapid Prosody Transcription experiment. Effect sizes here are expressed as percent change in the odds ratio for a “prominent” response as the result of a one-standard-deviation increase in the acoustic variable.
2.1.2. Hypotheses and predictions
The overarching research question we ask in this study, a mostly exploratory rather than confirmatory one, is what it is about prominence perception that individual differences in pragmatic skill predict. Though exploratory, we frame the investigation in terms of two basic hypotheses, the first of which is derived from the discussion above about the divergent nature of phonetic and phonological cues to prosodic prominence:
Hypothesis 1: Pragmatic skill is a predictor of sensitivity to phonetic cues to prominence, but not phonological cues to prominence.
We can therefore formulate two specific predictions about the data that will allow us to evaluate this hypothesis:
Prediction 1: The influence of phonological/structural prominence on listeners’ behavior in the RPT task will be unrelated to measures of pragmatic skill. More specifically (and based on Bishop et al., 2020, and Cole et al., 2019), listeners’ prominence ratings will be influenced by a word’s accent status according to the following hierarchy (and do so regardless of pragmatic skill): unaccented < prenuclear accented < nuclear accented.
Prediction 2: The influence of phonetic prominence on listeners’ behavior in RPT will be related to measures of their pragmatic skill. We operationalize phonetic prominence in two ways, based on basic principles of the AM model of English intonational phonology discussed above. First, we assume that the phonetic prominence of the three high-toned pitch accent types varies systematically, at least in terms of F0/pitch, according to the following hierarchy: !H < H* < L+H*. This hierarchy should therefore apply primarily to listeners with higher levels of pragmatic skill. Second, within pitch accent types (i.e., considering words marked by !H*, H*, or L+H* separately), the extent to which listeners are sensitive to phonetic variation in duration, intensity and the phonetic height of F0 peaks, should also correlate positively with listeners’ pragmatic skill.1
We will also explore the extent to which pragmatic skill might influence listeners’ sensitivity to signal-extrinsic (i.e., top-down) cues. Signal-extrinsic variables come in different forms, from those related to lexical properties of words to their information status in a larger discourse. We frame this part of our study in terms of the following hypothesis:
Hypothesis 2: Pragmatic skill is a predictor of sensitivity to signal-extrinsic information about both lexical and discourse properties of words.
Though exploratory, previous literature does offer some basis for predictions here. For example, Cole, Mahrt, and Hualde (2014), Bishop (2016), Turnbull et al. (2017), and, as described above, Orrico et al. (2025), all find that aspects of meaning influence perceived prominence, particularly meaning derived from discourse or pragmatic context. In general, words that are newer in the discourse, or used in an explicitly contrastive/alternative-eliciting way in the discourse, are more likely to be perceived as prominent. This was also a finding by Im et al., (2023), who operationalized discourse meaning in terms of information status categories in the RefLex Scheme (Riester & Baumann, 2017). These include categories such as ‘given,’ ‘accessible,’ and ‘new.’ Tracking discourse and integrating expectations about discourse into percepts is likely to be something that requires pragmatic skill, leading to the following intuitive prediction:
Prediction 3: The effect of a word’s status as new or contrastive in the discourse (known to have a positive correlation with perceived prominence) will increase as the pragmatic skill of listeners increases. We define newness and contrastiveness in terms of categories in the RefLex Scheme just alluded to, which we describe further below.
Finally, Cole, Mo, and Hasegawa-Johnson (2010), Cole et al. (2019), Bishop et al. (2020), and others have found that English-speaking listeners’ performance in RPT is influenced by a word’s lexical frequency, with higher frequency being negatively associated with perceived prominence. We have only one, relatively tenuous, basis for a prediction about a relation between pragmatic skill and such frequency effects, which comes from a study with a clinical autism sample (autism is a disorder defined, in part, by deficits in pragmatic and social skills). Grice et al., (2016) asked a group of adults with autism to rate words in samples of speech in terms of their informativeness but based on their sound properties. Among the findings in their study was that, compared to a group of neurotypical control listeners, the judgments of listeners with autism were found to rely less on signal-based intonational measures (such as pitch accent type) and more on lexical properties—in particular, lexical frequency, such that the inverse relationship between frequency and prominence ratings was stronger for the autism group. This suggests the possibility that listeners who attend less to phonetic cues like F0 (which, by hypothesis here, chiefly encodes paralinguistic information about intended emphasis), may instead rely more on non-signal based lexical properties of words. This leads to our final prediction:
Prediction 4: The effect of a word’s lexical frequency (known to have an inverse relation to perceived prominence) will decrease as the pragmatic skill of listeners increases.
We now turn to the methods for this reanalysis of listeners’ performance in Bishop and colleagues’ (2020) RPT experiment, which includes an abbreviated description of the materials and procedures used in that study.
2.2. Methods
2.2.1. Materials
Materials presented as stimuli to listeners in the RPT task consisted of samples of connected speech from four editions of former U.S. President Barack Obama’s Your Weekly Address series, currently in the public domain and stored on a U.S. government web archive (Obama, 2013, 2014a, 2014b, 2014c). Although in monologue form, these addresses could be interpreted as having a political purpose (e.g., touting the achievements of Obama’s administration and speaking less favorably about the activities of the opposing political party). Only the brief salutations at the beginning and end of the addresses were edited out, resulting in four speech samples containing a total of 1,821 words, or approximately 10 minutes of speech (Sample A: 470 words/2.6 min; Sample B: 448 words/2.35 min; Sample C: 445 words/2.7 min; Sample D: 458 words/2.4 min). All four speech samples were phonologically annotated by two labelers with training in the MAE_ToBI conventions. Interrater reliability for the two key phonological contrasts was as follows: 92% agreement for accent status (or κ = .83) and 79% agreement for accent type (or κ = .70). Words on which the annotators disagreed were excluded from the relevant analyses. The result across the four speech samples was 899 unaccented words and 889 accented words (52.1% nuclear, 47.9% prenuclear) for the analyses of accent status. Of these accented words used for the accent type analyses, 123 were !H* (22.0% of which were prenuclear, 57.7% nuclear, and 20.3% were unspecified due to disagreements among the annotators about phrase boundaries), 280 were H* (51.1% prenuclear, 33.6% nuclear, 15.4% unspecified), and 167 were L+H* (34.7% prenuclear, 37.7% nuclear, and 27.5% unspecified). Acoustic measures were also collected from the speech samples, extracted automatically in Praat from forced-aligned phone tiers. Acoustic measures were taken from the vowel of the lexically-stressed syllable for each word in the materials by a Praat script. Measures included (a) Max F0 (the maximum F0 during the vowel, measured using autocorrelation and hand corrected where tracking errors clearly occurred); (b) RMS intensity (measured uniformly across the frequency spectrum); and (c) the acoustic duration of the vowel.
Several non-signal-based properties of the speech samples, of secondary interest to the present study, included (a) each word’s SUBTLEX frequency (Brysbaert & New, 2009), updated from Bishop and colleagues’ use of CELEX frequencies in the 2020 analysis; (b) a word’s number of previous repetitions in the speech sample, and (c) discourse annotations using the RefLex Scheme. The RefLex Scheme (Riester & Baumann, 2017) is a system for assigning words to referential and lexical information status categories (see Table 1, below, for examples), which were not part of Bishop and colleagues’ (2020) analysis. The annotations were carried out by a specialist in semantics and pragmatics (the fourth author). For the purposes of the present study, we collapsed the categories into either referentially new (r-unused or r-new) or referentially not new (r-bridging or r-given), and lexically new (l-new) or lexically not new (l-accessible or l-given). We also included the category “contrastive/alternative-eliciting,” which refers to words with an explicit alternative present in the discourse. Examples of RefLex categorizations and ToBI annotations of the speech samples used in the RPT task are shown in Figure 2, below. It is important to point out that the relation between such pragmatic/discourse meanings are not one-to-one with any particular intonational category (Hirschberg et al., 2007; Thorson & Burdin, 2024). This was evident for the pitch accents and meaning contrasts considered here. For example, of the 123 words bearing a !H*, 39% were referentially new, 57% were lexically new, and 3% had a contrastive alternative in the discourse; of the 280 words bearing H*, 33% were referential new, 50% were lexically new, and 4% had an alternative in the discourse; and of the 167 words bearing a L+H* accent, 41% were referentially new, 60% were lexically new, and 8% had an alternative in the discourse (note that these information status categories are not exclusive).
Table 1: Information status labels in the RefLex Scheme (Riester & Baumann, 2017), adapted from a similar table in Im et al. (2023). Note that the scheme’s information status categories resemble but are not to be confused with many common notions of information structure categories.
| Level | Label | Description | Example |
| Referential (Applies to all lexical material in referring expressions) | r-given | coreferring entity present in discourse | A man was standing outside the bar. We then saw a woman enter the bar. |
| r-bridging | accessible entity present in discourse | I put the key in the ignition but the car wouldn’t start. | |
| r-unused | unique (definite), new entity in discourse | President Barack Obama delivered a brilliant speech in Oklahoma. | |
| r-new | non-unique (indefinite), new entity in discourse | After the holidays, Andrew arrived in a new car and Jane had also bought a new car. | |
| Lexical (Applies to content words) | l-given | active expression in discourse | A rat makes a surprisingly good pet. Moreover, a rat is quite friendly and affectionate. |
| l-accessible | semi-active expression in discourse | I tried to open the door but the knob was broken. | |
| l-new | inactive expression in discourse | Abernathy was very optimistic. The polls all showed an overwhelming majority for the politician. | |
| Contrastive/Alternative-Eliciting | alt | clearly identifiable alternative entity present in discourse | Did you call Mary? No, I called John. |
Figure 2: Example of the intonational and information status annotations applied to the speech samples used in the RPT experiment, in this case for the sentence, “That’s why the budget I sent congress earlier this year is built on the idea of opportunity for all.” The top tier shows intonational labels (tones in the MAE_ToBI conventions) while the bottom two tiers show lexical and referential newness labels (information status categories in the RefLex Scheme). A “dis” label on the intonation tier reflects words excluded from analysis due to disagreement among the ToBI annotators about accent status or accent type. Labels on both the lexical and referential newness tiers are limited to referring expressions (i.e., words within a noun/determiner phrase) and lexical newness applies to content words only.
2.2.2. Participants
Participants in this data set were (after excluding two non-monolinguals) 158 monolingual American English speakers recruited from the Greater New York City area (49 male, 109 female), aged 18 to 48. “Monolingual” was defined as not having learned a language other than English before the age of ten, and not being (by their own estimation) a fluent speaker in any second language studied after that age. All participants confirmed that they were free of any history of hearing or communication disorders, and that they lacked any training in prosodic theory or transcription.
2.2.3. Procedures
The entire experimental session took approximately 45 minutes to complete and consisted of two basic tasks:
RPT task: As described in Bishop et al. (2020), participants served as listeners in an RPT experiment designed to elicit coarse prosodic “annotations” of both prominence and juncture (in separate tasks); however, as mentioned above, only the prominence identification part of the task is of interest here. The experiment was carried out in a laboratory setting, in a sound attenuated booth with paper and pen (as in Cole, Mo, & Hasegawa-Johnson, 2010, rather than the electronically administered version in subsequent work). Participants identified words as prominent by underlining them on the printed transcript provided, based on the following instructions: “This part of the study is about how people use their voice when pronouncing words in English. When people speak, they use things like ‘loudness’ and ‘tone of voice’ to make some words ‘stand out’ more than others. In this part of the experiment, your job is to listen to President Obama’s voice and underline any and all words that he makes stand out in this way. To do this, you will need to listen very carefully to how he pronounces words ‘in real time.’” Although these identifications had to be made in real time, without the ability to pause or rewind, participants were presented with the speech sample three times and were able to add (or make changes to) their identification of prominent words on each of the two subsequent passes (tracked with different color pens), although we limit our analysis to the initial, first-pass responses.
Individual differences measures: In addition to the RPT task, all participants in the study completed three measures of cognitive processing styles that are, arguably, related to pragmatic skill: the Autism Spectrum Quotient (Baron-Cohen, Wheelwright, Skinner, et al., 2001), the Broad Autism Phenotype Questionnaire (Hurley et al., 2007) and the Reading the Mind in the Eyes test (Baron-Cohen, Wheelwright, Hill, et al., 2001). The Autism Spectrum Quotient is a 50-item, self-report questionnaire measuring “autistic-like” traits along five dimensions: social skills, imagination, attention to detail, attention-switching, and communication. As discussed above, the communication subscale of the Autism Spectrum Quotient (henceforth AQ-Comm) is the measure associated with pragmatic skill in previous work, and so it is this subscale, rather than the whole AQ, that is used in analyses here (although participants completed the whole test). Moreover, because the relevant subscale of the Broad Autism Phenotype Questionnaire (namely, the pragmatic language subscale) has similar items and was highly correlated with AQ-Comm (r = .60), we use AQ-Comm rather than that measure in these analyses. Scoring for the 10-item AQ-Comm subscale was done using a 4-point Likert scale (leading to a possible score of 40 rather than the 10 if using binary scoring) and reversed so that better pragmatic skill was associated with higher scores on AQ-Comm. The Reading the Mind in the Eyes test (henceforth EYES) was also scored such that higher scores indicate better pragmatic skill. AQ-Comm and EYES scores were, strictly speaking, positively correlated. But the correlation was extremely weak (r= .06), with roughly similar distributions (see Figure 3).
2.3. Results
2.3.1. Analysis
Mixed-effects logistic regression was used to evaluate our hypotheses about the relation between pragmatic skill and prominence-judging behavior in the RPT task. Categorical predictors were contrast coded, and continuous predictors were adjusted so as to be on similar scales and then centered on their means. Our basic approach was, for each prediction, to test the contribution of the critical individual differences-related factor relative to a basic baseline model before interpreting a model that contained that factor. The baseline models were those comprised of basic linguistic predictors, such as accent status, accent type, or well-understood acoustic variables whose effects are either core to the theory behind the research questions or have been established empirically in previous research. A factor’s success or failure to improve model fit relative to the baseline model was determined by a log likelihood ratio test using the anova function in R (R Core Team, 2024) and, in each case, we rely on the final models (with only significantly contributing factors) for interpretation.
2.3.2. Phonology versus phonetics: Sensitivity to accent status and accent type
Accent Status: To identify an effect for accent status, we first tested a model consisting of only accent status as the fixed effect; random effects included intercepts for subject and word. This model served as our baseline model to explore the effects of pragmatic skill (which would be inferred here from interactions between accent status and listeners’ AQ-Comm scores and/or EYES scores). We first tested for a simple effect of AQ-Comm scores and found that a model that contained only accent status was not improved by adding a simple effect for AQ-Comm (χ2 = 2.06, p < .1). A model that contained an interaction between accent status and AQ-Comm (the more crucial comparison here) provided only a marginal improvement over a model that contained only accent status and a simple effect for AQ-Comm (χ2 = 5.29, p < .1). A similar set of comparisons was carried out to test for an effect for EYES scores. A model containing a term for both accent status and EYES scores provided only a marginally significant improvement over the baseline model that contained accent status only (χ2 = 3.77, p < .1), and a model that allowed for an interaction between accent status and EYES scores resulted in no improvement at all over one with accent status and a simple effect for EYES (χ2 = 2.45, p > .1).
We therefore retained the baseline model, the results of which are shown in Table 2, where accent status was treated as a discrete ordinal variable so that effects at each level, unaccented < prenuclear accented < nuclear accented, could be simultaneously compared with Tukey corrections applied using the glht function in the multcomp package (Sellers et al., 2018; see also Bretz et al., 2011) for R. Unsurprisingly, given previous findings, accent status had a highly significant simple effect on the probability of a prominence judgment: Unaccented words were least likely to be judged as prominent, followed by prenuclear accented words, and then nuclear accented words. Given the non-contribution of any simple or interaction terms related to AQ-Comm or EYES, we conclude that this basic effect is not subject to individual differences related to pragmatic skill (at least not based on the measures of pragmatic skill we have employed here). The patterns across the range of both pragmatic skill scores are illustrated in Figure 4.
Table 2: Results for fixed-effects factors for the final model exploring possible individual differences related to accent status, a phonological contrast in prominence. Neither measure of pragmatic skill contributed to model fit.
| B | SE | z | p | |
| (Intercept) | –2.1031 | 0.0785 | –26.78 | p < .001 |
| Accent Status (Unaccented vs. PPA) | –1.0187 | 0.0794 | –12.83 | p < .001 |
| Accent Status (NPA vs. PPA) | 0.6287 | 0.0613 | 10.25 | p < .001 |
| Accent Status (NPA vs. Unaccented) | 1.6474 | 0.0801 | 20.57 | p < .001 |
Figure 4: Violin plots showing the distributions of listeners’ p-scores (representing the proportion of words identified as prominent) for unaccented, prenuclear accented, and nuclear accented words in the RPT task. Listeners are divided into pragmatic skill-based quartiles in terms of AQ-Comm scores (top) and EYES scores (bottom), with Q1 reflecting listeners with the lowest pragmatic skill according to that measure and Q4 the highest. White circles in violins represent means.
Accent Type: The picture was somewhat different when probing for individual differences related to accent type. To establish an effect of accent type, we first constructed a baseline model consisting of only accent type as the fixed effect; random effects included intercepts for subject and word. Analogous to the process described above, we first tested for a simple effect of AQ-Comm scores, finding that a model that contained both accent type and a simple effect for AQ-Comm did not improve fit significantly, though it was marginal (χ2 = 0.900, p < .1). A model that contained an interaction between accent type and AQ-Comm (again, the more crucial comparison) also failed to fit the data significantly better than a model that contained only accent type and a simple effect for AQ-Comm (χ2 = 1.722, p < .1). A similar set of comparisons was then carried out to test for an effect for EYES scores. A model containing terms for both accent type and EYES scores had a significantly better fit to the data than the baseline model (χ2 = 4.83, p < .05). Moreover, this model’s fit was, in turn, improved by adding an interaction term between EYES and accent type (χ2 = 29.12, p < .001). The output of this model is shown in Table 3, where it can be seen that the simple effect for accent type was driven primarily by the difference between H* and L+H* rather than between H* and !H* (see Cole et al., 2019, for a similar pattern, which was also found by Bishop et al., 2020, for this data set, at least for prenuclear accented words). However, and crucial to our interests here, the significant interaction between accent type and EYES indicated that the effect for accent type was more pronounced for listeners with higher EYES scores (indicating better pragmatic skill according to that measure). This relationship can be seen clearly in Figure 5 (as can the lack of a similar interaction between accent type and AQ-Comm). We note also that higher EYES scores were not simply associated with increased identification of words bearing H* and L+H*; higher pragmatic skill on this measure was also associated with lower rates of identification of words bearing !H*.
Table 3: Results for fixed-effects factors for the final model exploring possible individual differences related to accent type, a phonetic contrast in prominence. The effect for accent type interacted with one measure of pragmatic skill, the Reading the Mind in the Eyes measure (EYES).
| B | SE | z | p | |
| (Intercept) | –1.8870 | 0.0941 | –20.06 | p < .001 |
| Accent Type (!H* vs. H*) | 0.0956 | 0.0901 | 1.06 | p > .1 |
| Accent Type (L+H* vs. H*) | 0.6211 | 0.0788 | 7.88 | p < .001 |
| Accent Type (L+H* vs. !H*) | 0.5254 | 0.1008 | 5.21 | p < .001 |
| EYES | 0.0288 | 0.0136 | 2.12 | p < .1 |
| Accent Type (!H* vs. H*) × EYES | –0.0522 | 0.0138 | –3.79 | p < .001 |
| Accent Type (L+H* vs. H*) × EYES | 0.0296 | 0.0121 | 2.45 | p < .1 |
| Accent Type (L+H* vs. !H*) × EYES | 0.0817 | 0.0150 | 5.46 | p < .001 |
Figure 5: Violin plots showing the distributions of listeners’ p-scores (representing the proportion of words identified as prominent) for accented words bearing !H*, H*, and L+H* pitch accents. Listeners are divided into pragmatic skill-based quartiles in terms of AQ-Comm scores (top) and EYES scores (bottom), with Q1 reflecting listeners with the lowest pragmatic skill according to that measure and Q4 the highest. White circles in violins represent means.
2.3.3. Phonetic effects on the perception of high-toned pitch accents
To probe for the possible influence of pragmatic skill on the perception of within-category phonetic cues, an exploratory analysis tested for differences in listeners’ sensitivity to duration, intensity, and F0 for words bearing the three accent type categories, !H*, H*, and L+H*. The !H* and H* accents are both defined specifically in terms of their F0 level, !H* being significantly lower in a speaker’s range relative to a preceding H*, a H* being high in the speaker’s overall range. While L+H* is distinguished in form from H* by the presence of a leading low target, it has often been associated with higher (and delayed) peaks, and some authors have claimed that L+H*, at least in some varieties of English (e.g., Arvaniti & Garding, 2007), may actually be an emphatic version of H* rather than a distinct category of its own (Ladd & Morton, 1997; see also Orrico et al., 2025, for recent discussion). Moreover, and as illustrated above in Figure 1, Bishop et al. (2020) found listeners as a group to be more sensitive to F0 variation for L+H* than for H* or !H* when judging prominence in this data set. Given that pragmatic skill (at least based on the EYES measure) was predictive of listeners’ prominence identifications for pitch accents of different types, we were particularly interested in whether sensitivity to such within-category F0 variation was also subject to individual differences.
To this end, we once again carried out a series of planned comparisons of nested models, doing so separately for words marked by !H*, H*, and L+H*. In addition to random intercepts for subject and word, each baseline model consisted of simple effects for the two measures of pragmatic skill (AQ-Comm and EYES) and for each of the acoustic measurements collected in Bishop et al. (2020) and described above, in Section 2.2.1. These were: the RMS intensity of the lexically stressed syllable’s vowel; the duration of the stressed syllable’s vowel; and the maximum F0 over the stressed syllable’s vowel. We then subjected each baseline model to two rounds of nested model comparison: one that tested for improvements to model fit based on an interaction term between AQ-Comm scores with each of the acoustic measures, and then between EYES scores and each of the acoustic measures—retaining these interaction terms in the models only when they resulted in significant improvements to fit. The results of these model comparisons are summarized in Table 4 and demonstrated the following: First, in no case did an interaction term between AQ-Comm and an acoustic predictor result in an improvement to model fit, indicating the lack of meaningful relationships between this measure of pragmatic skill and sensitivity to these phonetic cues to prominence. Second, however, some interaction terms between EYES scores and acoustic predictors did significantly improve fit for the models of words bearing H* and L+H* accents. For words bearing H*, an interaction between EYES and duration significantly improved fit (χ2 = 7.113, p < .01), as did an interaction between EYES and F0 max (χ2 = 6.264, p < .05). For words bearing L+H*, an interaction between EYES and F0 max also improved fit significantly (χ2 = 12.300, p < .001). We therefore included these interactions in the final models, shown in tables 5, 6, and 7, which indicated the following. First, for words bearing !H* (Table 5), phonetic variation in general did not have much of an influence on listeners’ identification of prominence, with increased intensity and duration being only marginally associated with perceived prominence and F0 having no effect at all. This finding for F0 is not particularly surprising: The !H* category is defined by an F0 peak significantly lower than a preceding H*, and so increased F0 is only even possible up to a limit, after which it encroaches on the region of the range occupied by H*. However, previous work has found that words bearing !H* may be phonetically reduced in other ways as well (Ayers, 1996; Thorson & Burdin, 2024). In any case, neither measure of pragmatic skill had an interactive or simple effect on the perceived prominence of words with !H*. For words with H* (Table 6), increased duration, intensity, and max F0 were all significantly associated with increases in perceived prominence, as indicated by simple effects for these factors. Better pragmatic skill according to the EYES measure was also associated with increased perceived prominence for these words, but this simple effect was only marginally significant. More interestingly, however, there were two significant interactions with the EYES measure, one for duration and one for F0, indicating that prominence judgments were more sensitive to increases in these two acoustic measures for listeners with better pragmatic skill (see Figure 6).
Table 4: Outcomes of model comparisons used to identify the contribution to model fit by interactions between acoustic measures and measures of pragmatic skill.
| Intensity | Duration | F0 | ||
| AQ-Comm | !H* | χ2 = 0.218, p > .1 | χ2 = 0.877, p > .1 | χ2 = 0.613, p > .1 |
| H* | χ2 = 3.627, p < .1 | χ2 = 0.038, p > .1 | χ2 = 0.077, p > .1 | |
| L+H* | χ2 = 1.372, p > .1 | χ2 = 0.708, p > .1 | χ2 = 0.092, p > .1 | |
| EYES | !H* | χ2 = 0.327, p > .1 | χ2 = 0.877, p > .1 | χ2 = 0.193, p > .1 |
| H* | χ2 = 0.394, p > .1 | χ2 = 7.113, p < .01 | χ2 = 6.264, p < .05 | |
| L+H* | χ2 = 0.015, p > .1 | χ2 = .949, p > .1 | χ2 = 12.300, p < .001 |
Table 5: Findings for !H*: Results for fixed-effects factors for the model testing for effects of pragmatic skill and acoustic factors for words bearing the !H* accent.
| B | SE | z | p | |
| (Intercept) | –2.1103 | 0.1525 | –13.84 | p < .001 |
| Intensity | 0.0931 | 0.0487 | 1.91 | p < .1 |
| Duration | 0.0515 | 0.0271 | 1.90 | p < .1 |
| F0 Max | 0.0226 | 0.0352 | 0.64 | p > .1 |
| AQ-Comm | 0.0256 | 0.0181 | 1.42 | p > .1 |
| EYES | –0.0550 | 0.0562 | –0.98 | p > .1 |
Table 6: Findings for H*: Results for fixed-effects factors for the model testing for effects of pragmatic skill and acoustic factors for words bearing the H* accent.
| B | SE | z | p | |
| (Intercept) | –1.8351 | 0.1012 | –18.13 | p < .001 |
| Intensity | 0.1635 | 0.0376 | 4.35 | p < .001 |
| Duration | 0.0810 | 0.0130 | 6.22 | p < .001 |
| F0 Max | 0.0957 | 0.0232 | 4.13 | p < .001 |
| AQ-Comm | 0.0052 | 0.0098 | 0.53 | p > .1 |
| EYES | 0.0245 | 0.0126 | 1.94 | p < .1 |
| Duration × EYES | 0.0037 | 0.0015 | 2.48 | p < .05 |
| F0 Max × EYES | 0.0076 | 0.0033 | 2.31 | p < .05 |
Table 7: Findings for L+H*: Results for fixed-effects factors for the model testing for effects of pragmatic skill and acoustic factors for words bearing the L+H* accent.
| B | SE | z | p | |
| (Intercept) | –1.2406 | 0.1228 | –10.10 | p < .001 |
| Intensity | –0.0668 | 0.0370 | –1.81 | p < .1 |
| Duration | 0.0981 | 0.0206 | 4.76 | p < .001 |
| F0 Max | 0.1243 | 0.0162 | 7.66 | p < .001 |
| AQ-Comm | 0.0189 | 0.0121 | 1.56 | p > .1 |
| EYES | 0.0728 | 0.0158 | 4.61 | p < .001 |
| F0 Max × EYES | 0.0068 | 0.0019 | 3.53 | p < .001 |
Finally, for words bearing L+H* (Table 7), increases in duration and F0 were both associated with increases in perceived prominence, as indicated by significant simple effects for those measures. Increases in intensity had an inverse relationship with perceived prominence (a pattern reported by Bishop et al., 2020, for which we still do not have an explanation), although the effect was only marginally significant. EYES scores were strongly associated with perceived prominence for L+H* words, as indicated by the significant simple effect. But again, more interesting was the highly significant interaction between EYES and max F0, indicating that prominence judgments were more sensitive to the phonetic height of F0 peaks for listeners with better pragmatic skill (see Figure 7).
Figure 6: Violin plots showing the distributions of listeners’ p-scores (representing the proportion of words identified as prominent) for H* words as a function of duration (top) and max F0 (bottom) in the RPT task. “Lower” and “Higher” bins refer to values below or above the mean value for each acoustic variable. Listeners are divided into pragmatic skill-based quartiles in terms of EYES scores, with Q1 reflecting listeners with the lowest pragmatic skill on this measure and Q4 the highest. White circles in violins represent means.
Figure 7: Violin plots showing the distributions of listeners’ p-scores (representing the proportion of words identified as prominent) for L+H* words as a function of max F0. “Lower” and “Higher” bins refer to values below or above the mean value for max F0. Listeners are divided into pragmatic skill-based quartiles in terms of EYES scores, with Q1 reflecting listeners with the lowest pragmatic skill on this measure and Q4 the highest. White circles in violins represent means.
2.3.4. Signal extrinsic effects
To explore whether effects of the signal-extrinsic factors of interest are subject to individual differences in pragmatic skill, we explored a subset of the data for which there were RefLex annotations, namely, words in noun phrases (i.e., material that is associated with referring expressions). This allowed us to test for possible interactions between our two measures of pragmatic skill and factors related to both lexical (frequency) and discourse (information status) information. We therefore subjected this subset of the data (amounting to 20,574 observations) to exploratory regression modeling. To remind, in addition to one lexical factor, namely frequency based on SUBTLEX counts, we tested the following discourse/information status factors: (a) whether or not a word was lexically new in the discourse (versus lexically given/accessible); (b) whether a word was referentially new in the discourse (versus referentially given/accessible); (c) the number of previous repetitions of a word in the speech materials; and (d) whether or not the word was used contrastively, i.e., was an alternative to another referent in the discourse. A process of model comparison was again carried out as a method to determine which factors to include in the model. The first baseline model contained simple effects for lexical frequency and the four discourse/information status variables, as well as two signal-based factors, intensity and duration (F0 being left out, given this measure’s strong dependence on accent type). Interaction terms were then added one at a time and, once again, a log likelihood ratio test was used to determine whether the interaction term contributed significantly to the model. A sequential comparison of nested models found only the following interactions to contribute to improving fit (see Table 8): an interaction between intensity and EYES scores (χ2 = 13.36, p < .001); an interaction between lexical frequency and EYES scores (χ2 = 7.33, p < .01); and an interaction between lexical newness and EYES scores (χ2 = 5.16, p < .05).
Table 8: Outcome of model comparisons used to identify the contribution to model fit by interactions between measures of pragmatic skill and signal-extrinsic (top-down) and signal-based (bottom-up) factors.
| AQ-Comm | EYES | |
| Intensity | χ2 = 1.335, p > .1 | χ2 = 13.357, p < .001 |
| Duration | χ2 = 0.232, p > .1 | χ2 = 1.410, p > .1 |
| Lexical Frequency | χ2 = 0.079, p > .1 | χ2 = 7.330, p < .01 |
| Repetition | χ2 = 1.919, p > .1 | χ2 = 0.523, p > .1 |
| Newness(Lexical) | χ2 = 2.808, p > .1 | χ2 = 5.158, p < .05 |
| Newness(Referential) | χ2 = 2.808, p > .1 | χ2 = 0.009, p > .1 |
| Contrastive | χ2 = 2.073, p > .1 | χ2 = 0.042, p >.1 |
The final model that we interpret is shown in Table 9. As can be seen, there were simple effects indicating that, for this subset of the data for which the range of accent types and accent status contrasts were included, greater intensity and longer duration were both significantly associated with a higher probability of perceived prominence. Words with higher lexical frequency were numerically associated with a lower probability of perceived prominence, though insignificantly. A greater number of previous occurrences in the materials was significantly associated with a lower probability of perceived prominence. A word’s status as lexically new was significantly associated with increased perceived prominence. Somewhat surprisingly, however, referential newness was significantly associated with lower probabilities of perceived prominence. A word with an alternative in the discourse did not have a significant effect on perceived prominence. Neither measure of pragmatic skill had a significant simple effect on perceived prominence. Some of these simple effects, however, are better understood in the context of significant interactions with EYES scores, of which there were three. First, a significant interaction between EYES and intensity indicated that the prominence-lending effect of intensity was stronger for listeners with better pragmatic skill according to EYES; similarly, the prominence-lending effect of a word’s being lexically new in the discourse was stronger for listeners with better pragmatic skill according to EYES; and finally, the (insignificant) simple effect of lexical frequency—namely that higher lexical frequency was associated with lower perceived prominence—was shown to apply primarily to listeners with lower-pragmatic skill, indicating that listeners with higher pragmatic skill were less influenced by a word’s lexical frequency. These patterns are illustrated in Figure 8. We note also that the interaction between pragmatic skill and lexical frequency appears to be driven primarily by divergent sensitivity to higher frequency rather than to lower frequency.
Table 9: Results for fixed-effects factors for the model testing signal-extrinsic (top-down) and signal-based (bottom-up) factors.
| B | SE | z | p | |
| (Intercept) | –1.8853 | 0.1099 | –17.151 | p < .001 |
| Intensity | 0.1308 | 0.0329 | 3.978 | p < .001 |
| Duration | 0.1718 | 0.0218 | 7.892 | p < .001 |
| Lexical Frequency | –0.0164 | 0.0383 | –0.428 | p > .1 |
| Repetition | –0.0511 | 0.0185 | –2.759 | p < .001 |
| Newness(+Lexically New) | 0.4997 | 0.0678 | 7.373 | p < .001 |
| Newness(+Referentially New) | –0.3993 | 0.0687 | –5.809 | p < .001 |
| Contrastive (+Contrastive) | 0.0743 | 0.1239 | 0.599 | p > .1 |
| AQ-Comm | 0.0126 | 0.0122 | 1.036 | p > .1 |
| EYES | 0.0005 | 0.0167 | 0.032 | p > .1 |
| Intensity × EYES | 0.0187 | 0.0050 | 3.754 | p < .001 |
| Lexical Frequency × EYES | 0.0127 | 0.0045 | 2.851 | p < .01 |
| Newness(+Lexically New) × EYES | 0.0216 | 0.0101 | 2.144 | p < .05 |
Figure 8: Distribution of listeners’ p-scores (representing the proportion of words identified as prominent) as a function of intensity (top), lexical frequency (middle), and discourse newness (bottom). “Lower” and “Higher” bins refer to values below or above the mean value for each variable. Listeners are divided into pragmatic skill-based quartiles in terms of EYES scores, with Q1 reflecting listeners with the lowest pragmatic skill on this measure and Q4 the highest. White circles in violins represent means.
3. Discussion
3.1. Pragmatic skill and phonetic versus phonological cues to prominence
The study presented above, a new analysis of data collected in Bishop et al.’s (2020) Rapid Prosody Transcription (RPT) experiment, was concerned with how the perception of prosodic prominence by English-speaking listeners might be influenced by individual differences in pragmatic skill. With respect to this basic question, one main hypothesis and one secondary hypothesis were tested. Our main hypothesis (Hypothesis 1) was that individual differences in pragmatic skill would predict sensitivity to phonetic cues to prominence but not to phonological cues to prominence. This is because, by hypothesis, gradient phonetic cues to prominence primarily convey information about a speaker’s intended level of emphasis, which should be a matter of interpreting speaker intentions, a skill that varies within a population. These phonetic cues differ from phonological cues to prominence, which primarily convey an utterance’s information structure and syntactic structure. We assume that sensitivity to these kinds of grammatically mediated relationships should be relatively uniform in a speech community (or at least not dependent on pragmatic skill). We operationalized phonological and phonetic prominence in line with the Autosegmental Metrical (AM) theory of English that we assumed (Beckman & Pierrehumbert, 1986), where only discrete contrasts in accentuation—i.e., whether a word is unaccented, prenuclear pitch accented, or nuclear pitch accented—is regarded as a phonological (i.e., metrical) contrast in prominence. Residual gradient variation in cues like duration, intensity, and especially F0 are regarded as lending prominence phonetically.
In fact, what we found was that neither of the two measures of pragmatic skill that we tested (namely, the communication subscale of the AQ-Comm and EYES) were predictors of cross-listener variation in sensitivity to accent status differences, consistent with our Prediction 1. Instead, for listeners of all levels of AQ-Comm and EYES scores, the expected pattern held: Nuclear accented words were most likely to be identified as prominent, followed by prenuclear accented words, and then unaccented words—as has been reported in previous studies of English (Bishop et al., 2020; Cole et al., 2019) and German (Baumann & Winter, 2018; Lorenzen, 2024). This is what would be expected if listeners’ tracking of structural prominence, i.e., discrete metrical distinctions in accent status, is independent from how pragmatically “tuned in” to the intentions of speakers they happen to be.
Things were different, however, when considering gradient phonetic cues to prominence. Consistent with Prediction 2, we found that pragmatic skill predicted sensitivity to longer duration, greater intensity, and higher F0, with the evidence for individual differences in F0’s effect being particularly robust. First, we tested a three-way distinction in accent type for the three simple high-toned pitch accents in English, namely !H*, H*, and L+H*. While these three pitch accents reflect non-metrical phonological differences (due to their paradigmatic function in a system of tonal contrasts), they are understood to also vary in terms of the height of their typical F0 peaks, with L+H* tending to be highest, followed by H*, and then !H*. In fact, better pragmatic skill, at least based on EYES scores, was associated with a perceived prominence hierarchy of !H* < H* < L+H*, which was not apparent in listeners with low pragmatic skill. Moreover, when looking at the effect for gradient differences in F0 height on prominence judgements within-accent type category, we found that better pragmatic skill (again, according to EYES) also predicted sensitivity to increases in F0 for words bearing H* and L+H*.
At this point, we think it is worth highlighting that our finding that pragmatic skill correlates with sensitivity to phonetic cues is consistent with our hypothesized characterization of phonetic prominence as cueing paralinguistic emphasis, but inconsistent with many findings related to low-level acoustic processing, especially of pitch, in populations with clinical deficits in pragmatic skill. For example, a frequent finding is that discrimination of subtle differences in F0 is enhanced in individuals with autism spectrum conditions (e.g., Bonnel et al., 2003; see also the predictions related to AQ in Orrico et al., 2025). It is therefore important to note that (in addition to many other factors that make it difficult to draw parallels with clinical groups) the task we used here, namely RPT, does not involve discrimination of F0s as such, but rather the identification of prominence—a very different kind of task from pitch perception per se. Moreover, we note that when Grice et al. (2016), whose work we described briefly above, used a task more similar to RPT, they found that listeners with clinical diagnoses for an autism spectrum condition were less sensitive to differences in pitch accent types (i.e., pitch-based differences) relative to a control group of neurotypical listeners. Taken together, these findings can be seen as consistent with our characterization of pragmatic skill as being about interpreting cues to prominence. That is, it is primarily a predictor of variation in higher-level mapping of sound to meaning—paralinguistic meaning, in this case—rather than a predictor of lower-level auditory perception. It is an open question as to why pragmatic skill should affect the perception/interpretation of pitch more than other prominence-lending phonetic cues, but we note that this relationship might also be driving the variation reported by Baumann and Winter (2018). In their RPT experiment in German, Baumann and Winter found (via cluster analysis) that their listeners fell into one of two groups: those whose prominence judgments were particularly sensitive to pitch-based variables (like accent type differences and gradient F0 differences) and those who were sensitive to metrical-type variables (like duration and intensity). Although we cannot know the pragmatic skill profile of Baumann and Winter’s listeners, our findings suggest individual differences/psychometric approaches might offer insight into the mechanisms underlying such patterns. We comment further on this matter, below, in Section 3.3.
3.2. Pragmatic skill and signal-extrinsic/top-down cues to prominence
Our secondary hypothesis (Hypothesis 2) stated that pragmatic skill would also be a predictor of sensitivity to signal-extrinsic factors related to the lexicon and to discourse meaning, and this hypothesis found some support as well. Here our study was particularly exploratory, but we predicted that top-down boosts in perceived prominence associated with a word’s discourse prominence would be positively correlated with pragmatic skill (Prediction 3) and that the effects of lexical frequency would be negatively correlated with pragmatic skill (Prediction 4). First, we operationalized discourse prominence in terms of the information status categories in the RefLex Scheme. We assumed, based on previous findings (e.g., Im et al., 2023), that being either lexically new or referentially new (the two kinds of discourse newness recognized in the RefLex system) should result in a top-down boost in prominence. However, our prediction was that this effect would apply primarily to pragmatically-skilled listeners, who, by hypothesis, track and process discourse more intensively. Consistent with the prediction, a boost in perceived prominence was found for lexically new words relative to lexically given ones, but this was driven by listeners with higher levels of pragmatic skill (again, when pragmatic skill was based on the EYES measure). Interestingly, we did not find the expected main effect for referential newness, and indeed we found the opposite: a significant simple effect showing that referentially new words were associated with decreased perceived prominence. While this simple effect for referential newness is somewhat mysterious, we note it was recently reported by Lorenzen (2024) as well, for a group of German-speaking listeners. It is also important to point out that the information status categories in the RefLex Scheme have some overlap with, but are distinct from, information structure. In any case, most crucially for us is the fact that referential newness did not enter into any interaction with a measure of pragmatic skill. Our results also showed the expected simple effect for a number of previous repetitions of a word, i.e., a negative relation to perceived prominence (Cole, Mo, & Hasegawa-Johnson, 2010), but no interaction with pragmatic skill. And unlike previous studies (Im et al., 2023; Orrico et al., 2025), we did not find contrastive/alternative-eliciting words to have a simple effect on prominence judgments or any that interacted with pragmatic skill.
As for lexical frequency, we did find the expected inverse relationship with perceived prominence and that it interacted with pragmatic skill. As we predicted based on Grice et al. (2016), this effect of frequency was weaker for more pragmatically skilled listeners (Prediction 4). As just described above, Grice and colleagues’ study found that ratings of informativeness of words by listeners with autism were less dependent on pitch accent type but more sensitive to lexical frequency. Our prediction here was a quite speculative one, as is our interpretation. But one possibility is that listeners with weaker pragmatic skill, either clinically or sub-clinically, attend less to phonetic cues (especially F0) and the meanings they encode, and rely more on their internal expectations based on the lexicon, a type of compensatory/alternative strategy.
3.3. Unresolved issues and limitations of the present study
To summarize, we found evidence sufficient to retain our two basic hypotheses, namely that pragmatic skill would predict sensitivity to phonetic but not phonological cues to prominence (Hypothesis 1) and that pragmatic skill would predict sensitivity to signal-extrinsic factors as well, such as those with their locus in the lexicon and discourse (Hypothesis 2). Before concluding, we think a few comments are in order regarding some of the details.
Recall that our Hypothesis 1 and the related predictions were based on the idea that phonetic prominence (i.e., residual phonetic variation not directly used to parse accentual structure) is primarily about conveying emphasis, which we regard as paralinguistic in nature. And while we found some evidence for duration and intensity having a bigger influence on higher pragmatic skill listeners than on lower pragmatic skill listeners, the most consistent and robust interactions were with F0. If the reasoning that motivated this part of the study is on the right track, the special connection between pragmatic skill and F0 cues in particular needs to be better understood. One possibility, noted by Baumann and Winter (2018)—who, as described above, found some listeners to be more sensitive to pitch-based cues—is that different listeners may be more influenced by some aspects of the task’s instructions than others. Baumann and Winter speculate, for example, that instructions emphasizing the “highlighting” of words may bias listeners towards attending to pitch-based cues, while instructions that emphasize the “importance” of words may bias listeners towards lexical or other factors. While something like this is certainly possible, on its own we do not find it particularly compelling, at least not for English- or German-speaking listeners. For one thing, it would imply that words like “highlighting” and “pitch,” or “stress” and “duration,” are associated with each other in ways that we think are unlikely for linguistically naïve listeners. In the present case, where instructions given to listeners in Bishop et al. (2020) included both the words “loudness” and “tone of voice,” (see Section 2.2.3., above), it is unclear why more pragmatically skilled listeners would latch on to one term rather than the other. Moreover, such an explanation based on interpretation of task instructions fails to explain similar patterns related to pragmatic skill in studies that did not use RPT-like tasks (e.g., Jun & Bishop, 2015a, 2015b; Hurley & Bishop, 2016), including one that used an online test of processing (Bishop, 2017). For the moment, we think it is quite plausible that some listeners are simply more “tuned in” to some kinds of cues than others as a sort of perceptual style (a possibility also entertained by Baumann and Winter). What is therefore left to understand is what mechanisms drive these differences in perceptual style and why measures of pragmatic skill correlate with them.
On that note, it must also be highlighted that pragmatic skill itself—i.e., its characterization as a construct and especially how to measure it—is not particularly well understood. As we discussed in the introductory sections, at this point there seems to be sufficient evidence that (a) speakers and listeners vary in socio-communicative ways relevant to perceiving and comprehending language in context, and (b) measures that assess perspective-taking are likely to tap into this variation. However, while multiple research groups have found AQ-Comm to predict performance in perceptual or comprehension tasks (and production tasks; Bishop et al., 2022), others—such as Lorenzen (2024) and the present study—have not. Instead, we found scores on EYES, a test of emotion recognition, that predicted prominence judgments, both signal-based and non-signal-based. One possible reason for the discrepancy is that the AQ (and therefore the AQ-Comm subscale) is less reliable due to its being a questionnaire that requires conscious self-assessment, while EYES scores reflect performance on a task that doesn’t require such introspection. However, we are unconvinced that this is the crucial factor and instead suspect that these two measures may tap into different constructs. First, we point to the lack of any real correlation between AQ-Comm and EYES in this group of listeners. Additionally, although the EYES test is often presented as a measure of Theory of Mind (including by its authors), it is, strictly speaking, a test of emotion recognition. That being the case, it may measure (or at least also partially rely on) skills more closely related to empathy. As mentioned above, some authors have suggested that empathy rather than pragmatic skill is the relevant construct for predicting variation in how listeners decode intonation (Esteve-Gibert et al., 2020; Orrico & D’Imperio, 2020; Orrico et al., 2025).
In fact, we suspect that AQ-Comm and EYES do tap into different constructs (and we note the possibility that “pragmatic skill” may actually be a more complex/higher-order construct consisting of two or more such component constructs). On this point, we highlight the findings of Bishop (2016), briefly alluded to above in Section 1.3, who also explored the role of AQ-Comm and EYES in prominence perception. In that study, a follow-up to Bishop (2012b), English-speaking listeners were asked to rate (on a 5-point Likert scale) the perceived prominence of verbs and objects in simple Subject-Verb-Object (SVO) constructions, such as I climbed a mountain. Bishop (2016) was interested in the interaction of two kinds of effects on listeners’ ratings of the verbs and objects in these sentences: (a) a top-down effect of focus, and (b) a bottom-up effect of increased duration on the verb. Bishop (2016) hypothesized that, if focus interpretation guides attentional resources in a top-down way (Kristensen et al., 2013), listeners should be more sensitive to durational (i.e., bottom-up) manipulation of words that are focused rather than unfocused/given. As in Bishop (2012b), splicing was used so that listeners always heard the same audio recording of SVO target sentences (e.g., I climbed a mountain) but the focus status of the verb and object varied based on context. For example, listeners heard the same recording of a sentence like I climbed a mountain spliced into mini dialogues so that they followed questions that induced either broad VP focus on the verb phrase (e.g., What did you do?) or narrow focus on the object (e.g., What did you climb?). Additionally, the duration of verbs in all target SVOs was manipulated by synthesizing a five-step continuum for (whole-word) duration. If focus guided attention/sensitivity to bottom-up cues to prominence as hypothesized, Bishop (2016) expected to find an interaction such that listeners’ prominence judgments were more sensitive to the durational manipulation when the SVO occurred in the broad VP focus context than in the narrow object focus context, since the former context placed the manipulated verb within the focus constituent and the latter outside of it. While this is not what Bishop (2016) found—there was no such interaction of top-down and bottom-up cues—he did find two simple effects in that study, one top-down and one bottom-up, that turned out to interact with listeners’ AQ-Comm and EYES scores. As in the present study, however, AQ-Comm and EYES were mostly uncorrelated with each other, and these two measures did not enter into interactions with the same variables. The patterns were as follows. First, there was a simple effect for the durational manipulation: In general, verbs were rated as more prominent if their (manipulated) duration was longer. However, and somewhat surprisingly, this simple effect was quite unstable due to an interaction with EYES (but not AQ-Comm) scores, such that listeners with stronger pragmatic skill on this measure were more sensitive to the durational manipulation. Especially when we consider exactly what the EYES task involves—recognition of emotion, which is usually associated with paralanguage in speech—it is possible that the EYES measure is less about sensitivity to discourse context and more about sensitivity to paralinguistic cues, including paralinguistic cues to emphasis. This would align with the present findings, since EYES scores were predictive of listeners’ responsiveness to phonetic (but not phonological) cues to prominence, which we characterized as reflecting paralinguistic information about emphasis (rather than about metrical/structural, i.e., phonological prominence).
The second effect in Bishop (2016) was a simple top-down effect of focus, though not the attention-orienting one described above. Instead, focus was found to induce a “boost” in perceived prominence, a replication of a finding first reported in Bishop (2012b), whereby verbs in SVOs are rated as more prominent by listeners when they are interpreted as focused rather than given. Bishop (2016) found that this top-down boost interacted with AQ-Comm scores (but not EYES scores), such that the “focus boost” was stronger for listeners with better pragmatic skill according to AQ-Comm. One possibility, therefore, is that AQ-Comm scores relate most closely to a listener’s integration of discourse/pragmatic context. However, this would be inconsistent with the findings of the present study, since AQ-Comm did not interact with referential or lexical givenness (see Table 8, above). Instead, we tentatively suggest a subtly different possibility: AQ-Comm scores predict the extent to which listeners project their expectations about prominence in a top-down way when the conditions are highly ambiguous. This would be consistent with Bishop’s (2016) setup with the focus ambiguity. It would likely also be consistent with the findings of Jun and Bishop (2015b), mentioned above in Section 1.5., who attempted to prime accentuation patterns in implicit prosody. Presumably, assigning implicit prosody to silently-read text is mostly about projecting expectations—based on meaning, syntactic structure, or, in their case, on a preceding auditory prime. In fact, Jun and Bishop argued that their results indicate that AQ-Comm was predictive of a reader’s primability for accentuation, such that readers with low pragmatic skill according to AQ-Comm were less readily primed.
Therefore, while we speculate that EYES and AQ-Comm might relate to two different skills (sensitivity to paralanguage in the former and the propensity to project expectations in the latter), for the present we think the safest assumption is that pragmatic processing styles, i.e., ones that tend towards decoding speakers’ intentions and quickly and deeply integrating communicative contexts into interpretation, are likely to correlate with performance on perspective-taking measures of various kinds. The consequence is that better measurement, characterization, and understanding of the relevant construct will require large sample sizes and tools like structural equation modelling (or other latent variable techniques). And this is a notable limitation of work of this kind: Most production and perception studies in the literature that have tested for effects of constructs like empathy, autistic traits, and Theory of Mind have, thus far, almost surely been dramatically underpowered in terms of number of participants.
Finally, we highlight two important limitations of the present study. The first is the rather limited modeling of cue interactions. While some bottom-cue interaction is built into our approach (necessarily, given that pitch accent categories are clusters of synchronized phonetic cues), our statistical modeling surely presents a simplified picture of how signal-based cues work and, especially, how they interact with top-down cues. While this was intentional, given our interest in testing the role of pragmatic skill in particular, future work will need to test hypotheses that require additional and more complex interactions. A second limitation of the present study is its focus on relatively local/proximal characterizations of prominence. While the cues we have tested are those found within individual words, and our analyses assume a particular Autosegmental-Metrical characterization of how those cues are organized, it is well known that listeners are also influenced by more global and distal properties of spoken utterances (e.g., Dilley & McAuley, 2008; Dilley & Pitt, 2010; Morrill, Dilley, McAuley, & Pitt, 2014; Morrill, Dilley, & McAuley, 2014; Brown et al., 2015; see also the review in McQueen & Dilley, 2021). Further insight about prominence perception in the RPT task, and especially individual differences in the perception of prominence-lending cues, may be gained by looking at distal effects, possibly using analytical tools that emphasize syntagmatic relationships among tonal events (e.g., Dilley, 2005; Dilley & Breen, 2022; Ladd, 1990; Bishop, 2013: Ch. 4).
4. Conclusion
The study presented above tested the idea that prominence perception by English-speaking listeners is subject to individual differences related to pragmatic skill. Pragmatic skill, or the propensity to perceive and process linguistic information in relation to its context, and to attend closely to speaker intentions, varies in any population of neurotypical individuals and can, arguably, be estimated using a number of available instruments. In a new analysis of the Rapid Prosody Transcription (RPT) experiment reported in Bishop et al. (2020), we tested two hypotheses about the effect that pragmatic skill should have on listeners’ prominence identifications in the task. First, we tested the hypothesis that pragmatic skill should predict differences in sensitivity to phonetic cues to prominence but not to phonological cues to prominence, a hypothesis driven by the idea that phonetic cues to prominence reflect emphasis, an aspect of paralanguage rather than language. This hypothesis was supported by the fact that pragmatic skill did not predict differences in the effect of accent status (i.e., accentual structure) on prominence judgments, but it did predict differences in the effect of pitch accent type (differences between !H*, H* and L+H*). Notably, in terms of phonetic cues to prominence, these phonological accent types vary primarily in terms of pitch, suggesting that pragmatic skill might predict mostly sensitivity to pitch prominence in particular. Consistent with this, individual differences in pragmatic skill also significantly predicted listeners’ sensitivity to the height of F0 peaks within accent type categories (for H* and L+H*), in addition to similar effects for other phonetic cues, such as duration. According to our second hypothesis, pragmatic skill should also predict sensitivity to factors outside the signal in the RPT task, i.e., top-down effects. This hypothesis was supported by the findings that (a) pragmatic skill was associated with stronger effects of meaning (discourse newness) on prominence judgments and weaker effects of lexical (frequency) effects. Notably, the present study provides further evidence that some of the cross-listener variation in prominence judgment tasks like RPT (and presumably many other tasks) is structured variation—not mere noise. Also important to acknowledge, however, is that a better understanding of pragmatic skill as a construct is still needed, including and especially regarding its measurement in neurotypical adults.
Notes
- An anonymous reviewer notes that pitch accent type contrasts such as !H* versus H* not only reflect phonetic differences like F0 height, they also convey differences in meaning. We, of course, do not dispute this fact, as mentioned at the end of Section 1.5. Rather, our descriptions here are intended to be maximally clear about our assumptions regarding the nature of different kinds of sound-based “prominence,” as this is our primary concern. We assume that, in English, only distinctions in accent status represent “prominence” in a phonological (i.e., metrical) sense. Distinctions such as !H* versus L+H* are phonological in that they reflect contrasts in a system of paradigmatically opposed tonal categories, but they are not phonological contrasts in “prominence” as such. However, and crucial for our purposes here, it is also known (e.g., Ayers, 1996) that the realizations of words bearing different paradigmatically opposed accent types, such as !H*, H*, and L+H*, do differ systematically along gradient phonetic dimensions. We regard such differences—i.e., higher F0 peaks, longer duration—as differences in phonetic prominence. [^]
Acknowledgements
The authors are grateful to two anonymous reviewers for their helpful comments and suggestions. We also thank Taehong Cho, who handled our manuscript as guest editor, as well as the other guest editors who worked on the special issue: Jeffrey Holliday, Sahyang Kim, and Sang-Im Lee-Kim. We are also grateful to Juliana Colon and Jessica Spensieri for their help with data collection during the original study, and to those at LabPhon 19 who provided us with useful feedback.
Competing interests
The authors have no competing interests to declare.
References
Arvaniti, A., & Garding, G. (2007). Dialectal variation in the rising accents of American English. In J. Cole & J. I. Hualde (Eds.), Laboratory Phonology 9 (pp. 547–576). De Gruyter Mouton.
Ayers, G. (1996). Nuclear accent types and prominence: Some psycholinguistic experiments [Doctoral dissertation, The Ohio State University.]
Baron-Cohen, S., & Wheelwright, S. (2004). The empathy quotient: An investigation of adults with Asperger syndrome or high functioning autism, and normal sex differences. Journal of Autism and Developmental Disorders, 34(2), 163–175. http://doi.org/10.1023/B:JADD.0000022607.19833.00
Baron-Cohen, S., Wheelwright, S., Hill, J., Raste, Y., & Plumb, I. (2001). The “Reading the Mind in the Eyes” test revised version: A study with normal adults, and adults with Asperger syndrome or high-functioning autism. Journal of Child Psychology and Psychiatry, 42(2), 241–251. http://doi.org/10.1111/1469-7610.00715
Baron-Cohen, S., Wheelwright, S., Skinner, R., Martin, J., & Clubley, E. (2001). The Autism-Spectrum Quotient (AQ): Evidence from Asperger syndrome/high-functioning autism, males and females, scientists and mathematicians. Journal of Autism and Developmental Disorders, 31(1), 5–17. http://doi.org/10.1023/A:1005653411471
Baumann, S., Grice, M., & Steindamm, S. (2006). Prosodic marking of focus domains – categorical or gradient? In R. Hoffmann, H. Mixdorff, & O. Jokisch (Eds.), Proceedings of Speech Prosody 2006 (Paper 065). http://doi.org/10.21437/SpeechProsody.2006-73
Baumann, S., & Winter, B. (2018). What makes a word prominent? Predicting untrained German listeners’ perceptual judgments. Journal of Phonetics, 70, 20–38. http://doi.org/10.1016/j.wocn.2018.05.004
Beckman, M. (1986). Stress and non-stress accent. Netherlands Phonetic Archives 7. Foris. http://doi.org/10.1515/9783110874020
Beckman, M., & Ayers Elam, G. (1997). Guidelines for ToBI labelling (Version 3. Original work published 1993). The Ohio State University Research Foundation. https://www.ling.ohio-state.edu/research/phonetics/E_ToBI/
Beckman, M., & Hirschberg, J. (1994). The ToBI annotation conventions. The Ohio State University. https://scholar.google.com/citations?view_op=view_citation&hl=en&user=n35k9mQAAAAJ&citation_for_view=n35k9mQAAAAJ:zYLM7Y9cAGgC
Beckman, M., & Pierrehumbert, J. (1986). Intonational structure in Japanese and English. Phonology Yearbook, 3, 15–70. http://doi.org/10.1017/S095267570000066X
Bishop, J. (2012a). Focus, prosody, and individual differences in “autistic” traits: Evidence from cross-modal semantic priming. UCLA Working Papers in Phonetics, 111, 1–26. https://escholarship.org/uc/item/1z6819t5
Bishop, J. (2012b). Information structural expectations in the perception of prosodic prominence. In G. Elordieta & P. Prieto (Eds.), Prosody and meaning (pp. 239–270). De Gruyter Mouton. http://doi.org/10.1515/9783110261790.239
Bishop, J. (2013). Prenuclear accentuation in English: Phonetics, phonology, information structure [Doctoral dissertation, University of California, Los Angeles.] https://escholarship.org/uc/item/48t4q7cx
Bishop, J. (2016). Individual differences in top-down and bottom-up prominence perception. In J. Barnes, A. Brugos, S. Shattuck-Hufnagel, & N. Veilleux (Eds.), Proceedings of Speech Prosody 2016 (pp. 668–672). http://doi.org/10.21437/SpeechProsody.2016-137
Bishop, J. (2017). Focus projection and prenuclear accents: Evidence from lexical processing. Language, Cognition and Neuroscience, 32(2), 236–253. http://doi.org/10.1080/23273798.2016.1246745
Bishop, J., Kuo, G., & Kim, B. (2020). Phonology, phonetics, and signal-extrinsic factors in the perception of prosodic prominence: Evidence from Rapid Prosody Transcription. Journal of Phonetics, 82, 100977. http://doi.org/10.1016/j.wocn.2020.100977
Bishop, J., Zhou, C., Antolovic, K., Grebe, L., Hwang, K.-H., Imaezue, G., Lee, K.-E., Kristanova, E., Paulino, K., & Sichen, Z. (2022). Autistic traits predict spectral correlates of vowel intelligibility for female speakers. Journal of Autism and Developmental Disorders, 52(5), 2344–2349. http://doi.org/10.1007/s10803-021-05087-5
Bonnel, A., Mottron, L., Peretz, I., Trudel, M., Gallun, E., & Bonnel, A.-M. (2003). Enhanced pitch sensitivity in individuals with autism: A signal detection analysis. Journal of Cognitive Neuroscience, 15(2), 226–235. http://doi.org/10.1162/089892903321208169
Breen, M., Fedorenko, E., Wagner, M., & Gibson, E. (2010). Acoustic correlates of information structure. Language and Cognitive Processes, 25(7), 1044–1098. http://doi.org/10.1080/01690965.2010.504378
Bretz, F., Westfall, P., & Hothorn, T. (2011). Multiple comparisons using R. Chapman & Hall/CRC Press.
Brown, M., Salverda, A., Dilley, L., & Tanenhaus, M. (2015). Metrical expectations from preceding prosody influence perception of lexical stress. Journal of Experimental Psychology: Human Perception and Performance, 41(2), 306–323. http://doi.org/10.1037/a0038689
Bruggeman, A., Włodarczak, M., & Wagner, P. (2025). A comparison of discrete and continuous prominence perception methods in German. Speech Communication, 168, 103165. http://doi.org/10.1016/j.specom.2024.103165
Brysbaert, M., & New, B. (2009). Moving beyond Kucera and Francis: A critical evaluation of current word frequency norms and the introduction of a new and improved word frequency measure for American English. Behavior Research Methods, 41, 977–990. http://doi.org/10.3758/BRM.41.4.977
Büring, D. (2006). Focus projection and default prominence. In V. Molnár & S. Winkler (Eds.), The architecture of focus (pp. 321–346). De Gruyter Mouton. http://doi.org/10.1515/9783110922011.321
Cole, J., Hualde, J., Smith, C., Eager, C., Mahrt, T., & Napoleão de Souza, R. (2019). Sound, structure and meaning: The bases of prominence ratings in English, French and Spanish. Journal of Phonetics, 75, 113–147. http://doi.org/10.1016/j.wocn.2019.05.002
Cole, J., Mahrt, T., & Hualde, J. (2014). Listening for sound, listening for meaning: Task effects on prosodic transcription. In N. Campbell, D. Gibbon, & D. Hirst (Eds.), Proceedings of Speech Prosody 2014 (pp. 859–863). http://doi.org/10.21437/SpeechProsody.2014-161
Cole, J., Mo, Y., & Baek, S. (2010). The role of syntactic structure in guiding prosody perception with ordinary listeners and everyday speech. Language and Cognitive Processes, 25(7), 1141–1177. http://doi.org/10.1080/01690960903525507
Cole, J., Mo, Y., & Hasegawa-Johnson, M. (2010). Signal-based and expectation-based factors in the perception of prosodic prominence. Laboratory Phonology, 1(2), 425–452. http://doi.org/10.1515/labphon.2010.022
Dilley, L. (2005). The phonetics and phonology of tonal systems [Doctoral dissertation, Harvard University/Massachusetts Institute of Technology.]
Dilley, L., & Breen, M. (2022). An enhanced autosegmental-metrical theory (AM+) facilitates phonetically transparent prosodic annotation: A reply to Jun. In J. Barnes & S. Shattuck-Hufnagel (Eds.), Prosodic theory and practice (pp. 182–203). MIT Press. http://doi.org/10.7551/mitpress/10413.003.0018
Dilley, L., & McAuley, J. D. (2008). Distal prosodic context affects word segmentation and lexical processing. Journal of Memory and Language, 59(3), 294–311. http://doi.org/10.1016/j.jml.2008.06.006
Dilley, L., & Pitt, M. (2010). Altering context speech rate can cause words to appear or disappear. Psychological Science, 21(11), 1664–1670. http://doi.org/10.1177/0956797610384743
Esteve-Gibert, N., Schafer, A., Hemforth, B., Portes, C., Pozniak, C., & D’Imperio, M. (2020). Empathy influences how listeners interpret intonation and meaning when words are ambiguous. Memory & Cognition, 48, 566–580. http://doi.org/10.3758/s13421-019-00990-w
Ferguson, H., Cane, J., Douchkov, M., & Wright, D. (2015). Empathy predicts false belief reasoning ability: Evidence from the N400. Social Cognitive and Affective Neuroscience, 10(6), 848–855. http://doi.org/10.1093/scan/nsu131
Grice, M., Krüger, M., & Vogeley, K. (2016). Adults with Asperger syndrome are less sensitive to intonation than control persons when listening to speech. Culture and Brain, 4, 38–50. http://doi.org/10.1007/s40167-016-0035-6
Grice, M., Ritter, S., Niemann, H., & Roettger, T. (2017). Integrating the discreteness and continuity of intonational categories. Journal of Phonetics, 64, 90–107. http://doi.org/10.1016/j.wocn.2017.03.003
Gussenhoven, C. (1983). Testing the reality of focus domains. Language and Speech, 26(1), 61–80. http://doi.org/10.1177/002383098302600104
Gussenhoven, C. (1999). On the limits of focus projection in English. In P. Bosch & R. van der Sandt (Eds.), Focus: Linguistic, cognitive, and computational perspectives (pp. 43–55). Cambridge University Press.
Gussenhoven, C. (2015). Does phonological prominence exist? Lingue e Linguaggio, 14(1), 7–24. http://doi.org/10.1418/80751
Hayes, B. (1994). “Gesture” in prosody: Comments on the paper by Ladd. In P. Keating (Ed.), Papers in laboratory phonology III: Phonological structure and phonetic form (pp. 64–75). Cambridge University Press. http://doi.org/10.1017/CBO9780511659461.005
Hirschberg, J., Gravano, A., Nenkova, A., Sneed, E., & Ward, G. (2007). Intonational overload: Uses of the downstepped (H* !H* L- L%) contour in read and spontaneous speech. In J. Cole & J. Hualde (Eds.), Laboratory Phonology 9 (pp. 455–482). De Gruyter Mouton. http://doi.org/10.7916/D8KW5QG4
Hurley, R., & Bishop, J. (2016). Interpretation of “only”: Prosodic influences and individual differences. In J. Barnes, A. Brugos, S. Shattuck-Hufnagel, & N. Veilleux (Eds.), Proceedings of Speech Prosody 2016 (pp. 193–197). http://doi.org/10.21437/SpeechProsody.2016-40
Hurley, R., Losh, M., Parlier, M., Reznick, J., & Piven, J. (2007). The broad autism phenotype questionnaire. Journal of Autism and Developmental Disorders, 37(9), 1679–1690. http://doi.org/10.1007/s10803-006-0299-3
Im, S., Cole, J., & Baumann, S. (2023). Standing out in context: Prominence in the production and perception of public speech. Laboratory Phonology, 14(1), 1–62. http://doi.org/10.16995/labphon.6417
Jang, J., & Katsika, A. (2024). Focus structure and articulatory strengthening in Seoul Korean. In Y. Chen, A. Chen, A. Arvaniti (Eds.), Proceedings of Speech Prosody 2024 (pp. 597–601). http://doi.org/10.21437/SpeechProsody.2024-121
Jouravlev, O., Kell, A. J., Mineroff, Z., Haskins, A. J., Ayyash, D., Kanwisher, N., & Fedorenko, E. (2020). Reduced language lateralization in autism and the broader autism phenotype as assessed with robust individual-subjects analyses. Autism Research, 13, 1746–1761. http://doi.org/10.1002/aur.2393
Jun, S.-A. (Ed.). (2005). Prosodic typology: The phonology of intonation and phrasing. Oxford University Press. http://doi.org/10.1093/acprof:oso/9780199249633.001.0001
Jun, S.-A. (Ed.). (2014). Prosodic typology II: The phonology of intonation and phrasing. Oxford University Press. http://doi.org/10.1093/acprof:oso/9780199567300.001.0001
Jun, S.-A. (2025). Prosodic typology: Intonational tone types and functions. In D. Bradley, K. Dziubalska-Kołaczyk, C. Hamans, I.-H. Lee, & F. Steurs (Eds.), Contemporary linguistics: Integrating languages, communities, and technologies (Brill’s Handbooks in Linguistics, Vol. 7, pp. 93–111). Brill Academic. http://doi.org/10.1163/9789004715608
Jun, S.-A., & Bishop, J. (2015a). Priming implicit prosody: Prosodic boundaries and individual differences. Language and Speech, 58(4), 459–473. http://doi.org/10.1177/0023830914563368
Jun, S.-A., & Bishop, J. (2015b). Prominence in relative clause attachment: Evidence from prosodic priming. In L. Frazier & E. Gibson (Eds.), Explicit and implicit prosody in sentence processing: Studies in honor of Janet Dean Fodor (Studies in Theoretical Psycholinguistics, Vol. 46, pp. 217–240). Springer. http://doi.org/10.1007/978-3-319-12961-7_12
Kendon, A. (1972). Some relationships between body motion and speech: An analysis of an example. In A. Siegman & B. Pope (Eds.), Studies in dyadic communication (pp. 177–210). Pergamon. http://doi.org/10.1016/B978-0-08-015867-9.50013-7
Kendon, A. (1975). Studies in behavior of social interaction. Peter De Ridder Press. http://doi.org/10.1515/9783110907643.1
Kochanski, G., Grabe, E., Coleman, J., & Rosner, B. (2005). Loudness predicts prominence: Fundamental frequency lends little. Journal of the Acoustical Society of America, 118(2), 1038–1054. http://doi.org/10.1121/1.1923349
Kristensen, L., Wang, L., Petersson, K., & Hagoort, P. (2013). The interface between language and attention: Prosodic focus marking recruits a general attention network in spoken language comprehension. Cerebral Cortex, 23(8), 1836–1848. http://doi.org/10.1093/cercor/bhs164
Kulakova, E., & Nieuwland, M. (2016). Pragmatic skills predict online counterfactual comprehension: Evidence from the N400. Cognitive, Affective, & Behavioral Neuroscience, 16, 814–824. http://doi.org/10.3758/s13415-016-0433-4
Ladd, D. R. (1990). Metrical representation of pitch register. In J. Kingston & M. Beckman (Eds.), Papers in Laboratory Phonology I: Between the grammar and physics of speech (pp. 35–37). Cambridge University Press. http://doi.org/10.1017/CBO9780511627736.003
Ladd, D. R. (1996). Intonational phonology. Cambridge University Press.
Ladd, D. R., & Arvaniti, A. (2023). Prosodic prominence across languages. Annual Review of Linguistics, 9, 171–193. http://doi.org/10.1146/annurev-linguistics-031120-101954
Ladd, D. R., & Morton, R. (1997). The perception of intonational emphasis: Continuous or categorical? Journal of Phonetics, 25(3), 313–342. http://doi.org/10.1006/jpho.1997.0046
Lorenzen, J. (2024). Individual variability in the encoding and decoding of prosodic prominence relations [Doctoral dissertation, University of Cologne.]
McQueen, J., & Dilley, L. (2021). Prosody and spoken-word recognition. In C. Gussenhoven & A. Chen (Eds.), The Oxford handbook of language prosody (pp. 509–521). Oxford University Press. http://doi.org/10.1093/oxfordhb/9780198832232.013.33
Mitterer, H., Kim, S., & Cho, T. (2024). Use of segmental detail as a cue to prosodic structure in reference to information structure in German. Journal of Phonetics, 103, 101297. http://doi.org/10.1016/j.wocn.2024.101297
Morrill, T., Dilley, L., & McAuley, J. D. (2014). Prosodic patterning in distal speech context: Effects of list intonation and f0 downtrend on perception of proximal prosodic structure. Journal of Phonetics, 46, 68–85. http://doi.org/10.1016/j.wocn.2014.06.001
Morrill, T., Dilley, L., McAuley, J. D., & Pitt, M. (2014). Distal rhythm influences whether or not listeners hear a word in continuous speech: Support for a perceptual grouping hypothesis. Cognition, 131(1), 69–74. http://doi.org/10.1016/j.cognition.2013.12.006
Mücke, D., & Grice, M. (2014). The effect of focus marking on supralaryngeal articulation–Is it mediated by accentuation? Journal of Phonetics, 44, 47–61. http://doi.org/10.1016/j.wocn.2014.02.003
Nieuwland, M., Ditman, T., & Kuperberg, G. (2010). On the incrementality of pragmatic processing: An ERP investigation of informativeness and pragmatic abilities. Journal of Memory and Language, 63(3), 324–346. http://doi.org/10.1016/j.jml.2010.06.005
Obama, B. (2013, November 28). Remarks of President Barack Obama: Weekly Address, November 28, 2013. The White House. https://obamawhitehouse.archives.gov/the-press-office/2013/11/28/weekly-address-wishing-american-people-happy-thanksgiving
Obama, B. (2014a, March 1). Remarks of President Barack Obama: Weekly Address, March 1, 2014. The White House. https://obamawhitehouse.archives.gov/the-press-office/2014/03/01/weekly-address-investing-technology-and-infrastructure-create-jobs
Obama, B. (2014b, March 8). Remarks of President Barack Obama: Weekly Address, March 8, 2014. The White House. https://obamawhitehouse.archives.gov/the-press-office/2014/03/08/weekly-address-time-congress-raise-minimum-wage-american-people
Obama, B. (2014c, April 5). Remarks of President Barack Obama: Weekly Address, April 5, 2014. The White House. https://obamawhitehouse.archives.gov/the-press-office/2014/04/05/weekly-address-president-s-budget-ensures-opportunity-all-hardworking-am
Orrico, R., & D’Imperio, M. (2020). Individual empathy levels affect gradual intonation-meaning mapping: The case of biased questions in Salerno Italian. Laboratory Phonology, 11(1), 1–39. http://doi.org/10.5334/labphon.238
Orrico, R., Gryllia, S., Kim, J., & Arvaniti, A. (2025). Individual variability and the H* ~ L + H* contrast in English. Language and Cognition, 17, e9. http://doi.org/10.1017/langcog.2024.62
Pagel, L., Roessig, S., & Mücke, D. (2024). The encoding of prominence relations in supra-laryngeal articulation across speaking styles. Laboratory Phonology, 15(1), 1–55. http://doi.org/10.16995/labphon.10900
Pierrehumbert, J. (1980). The phonology and phonetics of English intonation [Doctoral dissertation, Massachusetts Institute of Technology.] https://dspace.mit.edu/handle/1721.1/16065?show=full
R Core Team. (2024). R: A language and environment for statistical computing (Version 4.4.2) [Computer Software]. R Foundation for Statistical Computing. https://www.R-project.org
Riester, A., & Baumann, S. (2017). The RefLex Scheme–annotation guidelines. SinSpeC: Working Papers of the SFB 732. Incremental Specification in Context, 14, 1–31. University of Stuttgart. http://doi.org/10.18419/opus-9011
Roessig, S. (2024). The inverse relation of pre-nuclear and nuclear prominences in German. Laboratory Phonology, 15(1), 1–43. http://doi.org/10.16995/labphon.9993
Roessig, S., Mücke, D., & Grice, M. (2019). The dynamics of intonation: Categorical and continuous variation in an attractor-based model. PLOS One, 14(5), 1–36. http://doi.org/10.1371/journal.pone.0216859
Roessig, S., Winter, B., & Mücke, D. (2022). Tracing the phonetic space of prosodic focus marking. Frontiers in Artificial Intelligence, 5, 842546. http://doi.org/10.3389/frai.2022.842546
Selkirk, E. (1984). Phonology and syntax: The relation between sound and structure. MIT Press.
Selkirk, E. (1995). Sentence prosody: Intonation, stress, and phrasing. In J. Goldsmith (Ed.), The handbook of phonological theory (pp. 550–569). Blackwell. http://doi.org/10.1111/b.9780631201267.1996.00018.x
Sellers, K., Steeg Morris, D., Balakrishnan, N., & Davenport, D. (2018). multicmp: Flexible modeling of multivariate count data via the multivariate Conway-Maxwell-Poisson distribution (Version 1.1) [Computer software]. http://doi.org/10.32614/CRAN.package.multicmp
Thorson, J., & Burdin, R. S. (2024). Phonetic implementation and the interpretation of downstepping in Mainstream US English. Journal of Phonetics, 105, 101340. http://doi.org/10.1016/j.wocn.2024.101340
Turnbull, R. (2014). Assessing the listener-oriented account of predictability-based phonetic reduction [Doctoral dissertation, The Ohio State University.] https://rave.ohiolink.edu/etdc/view?acc_num=osu1429796768
Turnbull, R. (2019). Listener-oriented phonetic reduction and theory of mind. Language, Cognition and Neuroscience, 34(6), 747–768. http://doi.org/10.1080/23273798.2019.1579349
Turnbull, R., Royer, A., Ito, K., & Speer, S. (2017). Prominence perception is dependent on phonology, semantics, and awareness of discourse. Language, Cognition and Neuroscience, 32(8), 1017–1033. http://doi.org/10.1080/23273798.2017.1279341
Van den Brink, D., Van Berkum, J., Bastiaansen, M., Tesink, C., Kos, M., Buitelaar, J., & Hagoort, P. (2012). Empathy matters: ERP evidence for inter-individual differences in social language processing. Social Cognitive and Affective Neuroscience, 7(2), 173–183. http://doi.org/10.1093/scan/nsq094
Xiang, M., Grove, J., & Giannakidou, A. (2013). Dependency-dependent interference: NPI interference, agreement attraction, and global pragmatic inferences. Frontiers in Psychology, 4, 708. http://doi.org/10.3389/fpsyg.2013.00708







