Iconicity as a General Property of LanguageEvidence from Spoken and Signed Languages

Pamela Perniss, Robin L. Thompson, Gabriella ViglioccoView original
OverviewBalancedalloy voice
Let's start with a tension that sat at the heart of twentieth-century linguistics. For decades, the safe bet was arbitrariness: words are conventional symbols whose sounds have no natural tie to their meanings. That view fit beautifully with neat architectures of the mind, semantics over here, phonology over there, and it echoed a debate that goes back to Plato's Cratylus. But if you spend time with the research we're talking about today — the work synthesized by Gabriella Vigliocco, Pamela Perniss, and David Thompson — you come away with a different picture. Arbitrariness isn't the whole story. Iconicity, those motivated links between form and meaning, is woven through language in ways that matter for how we process words, how we learn them, and how languages are organized. Iconicity sounds abstract until you see it. It's the way sequence in language mirrors sequence in experience — veni, vidi, vici — or how related pieces of meaning cluster together in grammar because they belong together conceptually. Bill Croft put it bluntly: the structure of language reflects, in some way, the structure of experience. You find this in the minutiae of morphology and syntax — order, contiguity, repetition — and you find it in the lexicon. Spoken languages carry onomatopoeia and whole families of sound-symbolic words. Sign languages, with their visual-spatial canvas, let handshape, location, and movement map onto the world. The British Sign Language sign for cry traces the path of tears. Aeroplane draws wings and a flight path in space. And yet each system also preserves islands of arbitrariness. It's not either or. It's a balance. If your training is Indo-European, the scale of lexical iconicity elsewhere can be a surprise. Japanese alone catalogues more than one thousand seven hundred mimetic words. Across Siwu in West Africa, Japanese, and many languages in Australia and the Americas, you see patterns: reduplication for repeated or prolonged events; voicing and vowel quality nudging you toward large versus small, heavy versus light. That's not just folklore. Köhler started documenting these intuitions a century ago, and Ramachandran and Hubbard made it famous with the Bouba–Kiki effect: about ninety-five percent of people, whether they speak English or Tamil, match bouba to rounded shapes and kiki to jagged ones. The details differ by language, but the pull is real. Back vowels and voiced consonants feel big and soft. Front vowels and voiceless stops feel small and sharp. Sign languages push iconicity into the foreground, but not in a simple way. Signs are built from handshape, place, and motion, and many are transparently motivated by what they denote; others are completely conventional. Think of iconicity as a sliding scale, not a box to be checked. Classifier predicates place and move entities in sign space, mapping grammar to perceived world structure. Nonmanuals — facial expressions and mouth movements — interact with manual signs; echo phonology captures those mouth and hand correspondences. Ronnie Wilbur's Event Visibility Hypothesis goes further, tying aspectual meaning — how events unfold in time — to the phonological shape of predicate signs. The message isn't that everything is iconic. It's that iconic motivations thread through structure and usage, even as languages maintain the efficiency and open-endedness that arbitrariness affords. So how do we know iconicity isn't just a nice story we tell ourselves? Because it shows up in behavior, often when it shouldn't matter. Start with off-line judgments. Imai and colleagues asked Japanese and English speakers to judge whether novel verbs "fit" certain actions. Even without explicit knowledge of mimetics, people felt the pull of iconic matches. Nygaard and coworkers went after prosody — the music of speech. When the tune of a made-up word matched the meaning dimension it was supposed to signal, say big versus small or happy versus sad, people learned the mapping more easily. Misalign the tune and the meaning, and performance dropped. That suggests prosody doesn't just add emotion. It carries structured, domain-specific content. And in sign, Vigliocco and colleagues compared how native British Sign Language users and English speakers grouped items. Signers clustered tools together in ways that reflected the hand to tool relationship, showing that tool-use iconicity shapes how the lexicon is mentally organized. Move online, into processing, and the story tightens. In speech, there's a lively debate around phonesthemes — those bits like gl- in glimmer, glare, and glow. Bergen found faster access for words with phonesthemic patterns, and though Westbury pushed on whether it's true iconicity or just statistical familiarity, the effect is there: sound shapes are connected to pathways in meaning space. Shintel and colleagues tapped into something beautifully simple. Describe motion heading up, and speakers naturally raise pitch. Talk about fast movement, and they speak faster. Listeners then infer speed from nothing more than that acoustic shape. Form follows meaning into the sound stream, and comprehension reads it back out. These mappings show up early. Walker and colleagues tested three to four month old infants on pitch to space correspondences. High sounds go up, low sounds go down. When the pairings were congruent, babies looked longer; when they were flipped, they lost interest. That's a hint that some iconic alignments could be perceptual primitives, not cultural inventions. By two and a half years, Maurer and team reported that children already share adult-like Bouba–Kiki biases in matching sounds to shapes. Imai showed that three year olds in Japanese learn sound-symbolic verbs faster than arbitrary ones. Somewhere between infancy and preschool, iconic pulls become learning levers. In sign languages, the online evidence is especially clean. Thompson and colleagues asked American Sign Language users to recognize pictures after seeing a sign. When the part of the sign that made it iconic was also salient in the image — think of seeing a tear in a picture after the sign for cry — signers were faster. Non-signers, doing a parallel task with English words, got no boost. Vinson and collaborators replicated and broadened that result: more iconic signs sped recognition in general, regardless of whether the specific iconic feature was highlighted in the picture. That points to a general processing benefit for iconicity in the lexicon. And then comes a twist. In a follow-up phonological task, Thompson's group asked signers to make judgments about handshape that shouldn't require meaning. Iconicity still crept in, slowing people down and hurting accuracy for highly iconic signs. Meaning was accessed so automatically that it interfered with a purely form-based decision. Kovic and colleagues add a neural layer: when sound to shape mappings were congruent, they elicited an early change in the event-related potential, an N-200, hinting that the brain quickly flags iconic fit before higher reasoning can kick in. Development adds texture. Ormel and colleagues found that ten to twelve year old signers recognized highly iconic signs faster than less iconic ones in a picture to sign matching task. Meier's longitudinal work, on the other hand, didn't see iconicity reliably reduce children's sign errors, reminding us that not every learning path is dominated by iconic aid. Namy and coworkers offered a timeline across gesture: at eighteen months and again at four years, children learned arbitrary and iconic gestural labels about equally well. At twenty-six months, iconic labels pulled ahead. It's as if there's a window where analogy to the world gives learning a bump. Tolar's work helps explain why: across ages, mappings that were motoric — where the gesture mimicked an action — had an edge over purely perceptual iconicity, where the link was about static visual resemblance. Young learners, famously action-oriented, seem to leverage that motor bridge. Underneath all this is a pair of pressures that shape languages over development and across time. On one side, arbitrariness gives you a combinatorial explosion — the power to pack a giant vocabulary into distinct forms with minimal confusion. On the other, iconicity ties your code to embodied experience, letting perception and action scaffold your first steps into a lexicon. Vigliocco, Perniss, and Thompson argue we should stop treating these as adversaries and see them as interacting constraints. Think mechanistically: when you hear or produce a label, you co-activate sensory and motor traces. A simple Hebbian logic — neurons that fire together wire together — binds those traces to the form. If the form already resembles the experience, even a little, that bridge shortens. You learn it faster. You retrieve it more easily. And sometimes, as we saw in the phonological task with signers, you can't help but cross that bridge even when it gets in the way. Once you widen the lens to language as a multimodal system, the channels for iconic grounding multiply. Gesture rides alongside speech. Prosody shapes the analog side of the signal. Echo phonology links mouth and hands. Sign space makes events visible in the structure of predicates, just as Wilbur proposed. Croft's intuition lands here: structure reflects experience not because language is a pantomime, but because our communicative system exploits every available mapping that reduces the gap between form and world, while safeguarding the arbitrariness that keeps the system flexible. It's not all tidy. The review flags important caveats. There still aren't that many direct processing studies in some domains, and results vary with task and modality. Untangling motoric from perceptual and functional iconicity — when the link mirrors actions, when it mirrors shapes, and when it mirrors how things work — is hard, especially in development, where cognitive changes are fast and layered. And effects that look like iconicity can be confounded with distributional statistics, as Westbury cautioned in the phonestheme debate. The safest conclusion is a measured one: iconicity is not a marginal flourish. It's a general property of language that shows up in structure, speeds processing, and often helps learning, even as its strength and timing depend on the modality and on what kind of iconicity we're talking about. If you're a map person, here's the lay of the land. Off-line judgments show consistent biases — from Bouba–Kiki to action-verb intuitions — that cut across languages. On-line tasks reveal that form tracks meaning in production and perception, whether it's pitch following upward motion or iconic signs jumping the queue in recognition. Developmentally, infants are sensitive to congruent mappings within months; toddlers and preschoolers show selective learning advantages, especially when the link is through action. And at every step, sign languages remind us that iconicity can be structural, not just lexical, embedded in how events and participants are laid out in space. Where does that leave the old dual-system picture — semantics and phonology kept strictly apart? We still need it for parts of the story. Discrete categories and arbitrary conventions are why your language can grow without bound. But the day-to-day business of understanding and learning seems to rely on something messier and more grounded. Thompson's interference result is a nice emblem of that: even when you're trying to focus on handshape, meaning barges in. Iconic mappings aren't just present. They're active. Two brief thoughts to close. First, the methods that brought this field to life — from clever infant looking-time studies to sign-based reaction-time tasks and early neural markers — are still underused outside Indo-European speech and the handful of well-studied sign languages. Broader sampling will tell us which iconic pressures are truly universal and which are tightly tuned to community conventions. Second, the multimodal frame is a design principle, not just a research slogan. If prosody, gesture, and sign all carry iconic hooks into meaning, theories of language that ignore those hooks risk explaining away some of the very things that make language learnable and efficient. So next time someone tells you words are arbitrary, smile and say: often, and gloriously, yes. But not only. Sometimes the form leans toward the meaning, and our minds are built to feel that tug — fast, early, and across the ways humans make language.

Let's start with a tension that sat at the heart of twentieth-century linguistics. For decades, the safe bet was arbitrariness: words are conventional symbols whose sounds have no natural tie to their meanings. That view fit beautifully with neat architectures of the mind, semantics over here, phonology over there, and it echoed a debate that goes back to Plato's Cratylus.

But if you spend time with the research we're talking about today — the work synthesized by Gabriella Vigliocco, Pamela Perniss, and David Thompson — you come away with a different picture. Arbitrariness isn't the whole story. Iconicity, those motivated links between form and meaning, is woven through language in ways that matter for how we process words, how we learn them, and how languages are organized.

Iconicity sounds abstract until you see it. It's the way sequence in language mirrors sequence in experience — veni, vidi, vici — or how related pieces of meaning cluster together in grammar because they belong together conceptually. Bill Croft put it bluntly: the structure of language reflects, in some way, the structure of experience.

You find this in the minutiae of morphology and syntax — order, contiguity, repetition — and you find it in the lexicon. Spoken languages carry onomatopoeia and whole families of sound-symbolic words. Sign languages, with their visual-spatial canvas, let handshape, location, and movement map onto the world.

The British Sign Language sign for cry traces the path of tears. Aeroplane draws wings and a flight path in space. And yet each system also preserves islands of arbitrariness. It's not either or. It's a balance.

If your training is Indo-European, the scale of lexical iconicity elsewhere can be a surprise. Japanese alone catalogues more than one thousand seven hundred mimetic words. Across Siwu in West Africa, Japanese, and many languages in Australia and the Americas, you see patterns: reduplication for repeated or prolonged events; voicing and vowel quality nudging you toward large versus small, heavy versus light.

That's not just folklore. Köhler started documenting these intuitions a century ago, and Ramachandran and Hubbard made it famous with the Bouba–Kiki effect: about ninety-five percent of people, whether they speak English or Tamil, match bouba to rounded shapes and kiki to jagged ones. The details differ by language, but the pull is real.

Back vowels and voiced consonants feel big and soft. Front vowels and voiceless stops feel small and sharp.

Sign languages push iconicity into the foreground, but not in a simple way. Signs are built from handshape, place, and motion, and many are transparently motivated by what they denote; others are completely conventional. Think of iconicity as a sliding scale, not a box to be checked.

Classifier predicates place and move entities in sign space, mapping grammar to perceived world structure. Nonmanuals — facial expressions and mouth movements — interact with manual signs; echo phonology captures those mouth and hand correspondences. Ronnie Wilbur's Event Visibility Hypothesis goes further, tying aspectual meaning — how events unfold in time — to the phonological shape of predicate signs.

The message isn't that everything is iconic. It's that iconic motivations thread through structure and usage, even as languages maintain the efficiency and open-endedness that arbitrariness affords.

So how do we know iconicity isn't just a nice story we tell ourselves? Because it shows up in behavior, often when it shouldn't matter. Start with off-line judgments.

Imai and colleagues asked Japanese and English speakers to judge whether novel verbs "fit" certain actions. Even without explicit knowledge of mimetics, people felt the pull of iconic matches. Nygaard and coworkers went after prosody — the music of speech.

When the tune of a made-up word matched the meaning dimension it was supposed to signal, say big versus small or happy versus sad, people learned the mapping more easily. Misalign the tune and the meaning, and performance dropped. That suggests prosody doesn't just add emotion.

It carries structured, domain-specific content. And in sign, Vigliocco and colleagues compared how native British Sign Language users and English speakers grouped items. Signers clustered tools together in ways that reflected the hand to tool relationship, showing that tool-use iconicity shapes how the lexicon is mentally organized.

Move online, into processing, and the story tightens. In speech, there's a lively debate around phonesthemes — those bits like gl- in glimmer, glare, and glow. Bergen found faster access for words with phonesthemic patterns, and though Westbury pushed on whether it's true iconicity or just statistical familiarity, the effect is there: sound shapes are connected to pathways in meaning space.

Shintel and colleagues tapped into something beautifully simple. Describe motion heading up, and speakers naturally raise pitch. Talk about fast movement, and they speak faster.

Listeners then infer speed from nothing more than that acoustic shape. Form follows meaning into the sound stream, and comprehension reads it back out.

These mappings show up early. Walker and colleagues tested three to four month old infants on pitch to space correspondences. High sounds go up, low sounds go down.

When the pairings were congruent, babies looked longer; when they were flipped, they lost interest. That's a hint that some iconic alignments could be perceptual primitives, not cultural inventions. By two and a half years, Maurer and team reported that children already share adult-like Bouba–Kiki biases in matching sounds to shapes.

Imai showed that three year olds in Japanese learn sound-symbolic verbs faster than arbitrary ones. Somewhere between infancy and preschool, iconic pulls become learning levers.

In sign languages, the online evidence is especially clean. Thompson and colleagues asked American Sign Language users to recognize pictures after seeing a sign. When the part of the sign that made it iconic was also salient in the image — think of seeing a tear in a picture after the sign for cry — signers were faster.

Non-signers, doing a parallel task with English words, got no boost. Vinson and collaborators replicated and broadened that result: more iconic signs sped recognition in general, regardless of whether the specific iconic feature was highlighted in the picture. That points to a general processing benefit for iconicity in the lexicon.

And then comes a twist. In a follow-up phonological task, Thompson's group asked signers to make judgments about handshape that shouldn't require meaning. Iconicity still crept in, slowing people down and hurting accuracy for highly iconic signs.

Meaning was accessed so automatically that it interfered with a purely form-based decision. Kovic and colleagues add a neural layer: when sound to shape mappings were congruent, they elicited an early change in the event-related potential, an N-200, hinting that the brain quickly flags iconic fit before higher reasoning can kick in.

Development adds texture. Ormel and colleagues found that ten to twelve year old signers recognized highly iconic signs faster than less iconic ones in a picture to sign matching task. Meier's longitudinal work, on the other hand, didn't see iconicity reliably reduce children's sign errors, reminding us that not every learning path is dominated by iconic aid.

Namy and coworkers offered a timeline across gesture: at eighteen months and again at four years, children learned arbitrary and iconic gestural labels about equally well. At twenty-six months, iconic labels pulled ahead. It's as if there's a window where analogy to the world gives learning a bump.

Tolar's work helps explain why: across ages, mappings that were motoric — where the gesture mimicked an action — had an edge over purely perceptual iconicity, where the link was about static visual resemblance. Young learners, famously action-oriented, seem to leverage that motor bridge.

Underneath all this is a pair of pressures that shape languages over development and across time. On one side, arbitrariness gives you a combinatorial explosion — the power to pack a giant vocabulary into distinct forms with minimal confusion. On the other, iconicity ties your code to embodied experience, letting perception and action scaffold your first steps into a lexicon.

Vigliocco, Perniss, and Thompson argue we should stop treating these as adversaries and see them as interacting constraints. Think mechanistically: when you hear or produce a label, you co-activate sensory and motor traces. A simple Hebbian logic — neurons that fire together wire together — binds those traces to the form.

If the form already resembles the experience, even a little, that bridge shortens. You learn it faster. You retrieve it more easily.

And sometimes, as we saw in the phonological task with signers, you can't help but cross that bridge even when it gets in the way.

Once you widen the lens to language as a multimodal system, the channels for iconic grounding multiply. Gesture rides alongside speech. Prosody shapes the analog side of the signal.

Echo phonology links mouth and hands. Sign space makes events visible in the structure of predicates, just as Wilbur proposed. Croft's intuition lands here: structure reflects experience not because language is a pantomime, but because our communicative system exploits every available mapping that reduces the gap between form and world, while safeguarding the arbitrariness that keeps the system flexible.

It's not all tidy. The review flags important caveats. There still aren't that many direct processing studies in some domains, and results vary with task and modality.

Untangling motoric from perceptual and functional iconicity — when the link mirrors actions, when it mirrors shapes, and when it mirrors how things work — is hard, especially in development, where cognitive changes are fast and layered. And effects that look like iconicity can be confounded with distributional statistics, as Westbury cautioned in the phonestheme debate. The safest conclusion is a measured one: iconicity is not a marginal flourish.

It's a general property of language that shows up in structure, speeds processing, and often helps learning, even as its strength and timing depend on the modality and on what kind of iconicity we're talking about.

If you're a map person, here's the lay of the land. Off-line judgments show consistent biases — from Bouba–Kiki to action-verb intuitions — that cut across languages. On-line tasks reveal that form tracks meaning in production and perception, whether it's pitch following upward motion or iconic signs jumping the queue in recognition.

Developmentally, infants are sensitive to congruent mappings within months; toddlers and preschoolers show selective learning advantages, especially when the link is through action. And at every step, sign languages remind us that iconicity can be structural, not just lexical, embedded in how events and participants are laid out in space.

Where does that leave the old dual-system picture — semantics and phonology kept strictly apart? We still need it for parts of the story. Discrete categories and arbitrary conventions are why your language can grow without bound.

But the day-to-day business of understanding and learning seems to rely on something messier and more grounded. Thompson's interference result is a nice emblem of that: even when you're trying to focus on handshape, meaning barges in. Iconic mappings aren't just present. They're active.

Two brief thoughts to close. First, the methods that brought this field to life — from clever infant looking-time studies to sign-based reaction-time tasks and early neural markers — are still underused outside Indo-European speech and the handful of well-studied sign languages. Broader sampling will tell us which iconic pressures are truly universal and which are tightly tuned to community conventions.

Second, the multimodal frame is a design principle, not just a research slogan. If prosody, gesture, and sign all carry iconic hooks into meaning, theories of language that ignore those hooks risk explaining away some of the very things that make language learnable and efficient.

So next time someone tells you words are arbitrary, smile and say: often, and gloriously, yes. But not only. Sometimes the form leans toward the meaning, and our minds are built to feel that tug — fast, early, and across the ways humans make language.

More in Psychology