📝 This_Is_Your_Brain_on_Music_FULL_Summary.mdv4.5.1 · 2026-10-02

This Is Your Brain on Music -- full summary

Daniel J. Levitin, This Is Your Brain on Music: The Science of a Human Obsession (2006). A complete section-by-section summary of the whole book, made for Andrew from the text file he supplied. Introduction, chapters 1 to 9 and the appendices are all covered in the book's own order; the bibliographic notes are a reading list and are summarised by topic only.

Gaps from the scan: the instrument frequency-range figure and the interval table in chapter 1, the lyric and beat diagrams in chapter 2, the Ponzo and Shepard figures in chapter 3, and the Appendix A figure labels were damaged in the text file, so those pictures are described only as far as the surrounding prose allows.

Contents


Introduction: I Love Music and I Love Science, Why Would I Want to Mix the Two?

The chapter opens with an epigraph from Robert Sapolsky (Why Zebras Don't Get Ulcers). Sapolsky laments that many people fear science or believe that choosing it rules out compassion, art, or awe at nature. For him, science is meant to renew mystery rather than cure us of it.

How the author fell in love with sound

In the summer of 1969, aged eleven, Levitin spent a hundred dollars on a stereo. He had earned it weeding neighbours' gardens at seventy-five cents an hour. He spent long afternoons listening to Cream, the Rolling Stones, Chicago, Simon and Garfunkel, Bizet, Tchaikovsky, George Shearing and the saxophonist Boots Randolph. In college he later set his speakers on fire by playing them too loud.

His mother was a novelist who wrote daily in the den and played piano for an hour each night. His father was a businessman who worked eighty-hour weeks, forty of them in a home office. The father proposed a deal. He would buy headphones if the boy used them whenever he was home.

The headphones changed how Levitin listened. The new artists were experimenting with stereo mixing, and his cheap speakers had hidden the depth. He could now hear where instruments sat left to right and front to back in reverberant space. Records became about sound rather than songs, chords, lyrics or a singer's voice. His examples were the swampy ambience of Creedence's "Green River," the open pastoral feel of the Beatles' "Mother Nature's Son," and the faint oboes in Beethoven's Sixth under Karajan, soaked in the air of a large stone church. Headphones also made music feel personal, as if it came from inside his head. That connection, he says, drove him to become a recording engineer and producer.

Years later, Paul Simon told him that sound is what he too listens for in his own records. His first impression is of the overall sound, not the chords or lyrics.

Learning to hear, and years in the studio

Levitin dropped out of college after the speaker incident and joined a rock band. They became good enough to record at a twenty-four-track studio in California with engineer Mark Needham, who later recorded hits for Chris Isaak, Cake and Fleetwood Mac. Needham liked him, probably because he alone went into the control room to listen back while the others got high between takes. Needham treated him as a producer and taught him how microphone choice and placement change a sound. A mic close to a guitar amp sounds fuller, rounder and more even. A mic placed farther back picks up the room and sounds more spacious, but loses some midrange. Levitin could not hear some of these differences at first, but Needham taught him what to listen for.

The band grew moderately well known in San Francisco and got radio play. It broke up because of the guitarist's repeated suicide attempts and the singer's habits of inhaling nitrous oxide and cutting himself with razor blades. Levitin then produced other bands. He learned to distinguish microphones and even brands of tape:

Once he knew what to listen for, he could tell them apart as easily as an apple from a pear or an orange.

He worked with leading engineers: Leslie Ann Jones (Frank Sinatra, Bobby McFerrin), Fred Catero (Chicago, Janis Joplin) and Jeffrey Norman (John Fogerty, the Grateful Dead). Despite being the producer, he felt intimidated by them. He sat in on sessions with Heart, Journey, Santana, Whitney Houston and Aretha Franklin. He watched engineers discuss the articulation of a guitar part, the delivery of a vocal, and the syllables of a lyric, choosing among ten takes. He wondered how they trained their ears to hear what ordinary people cannot.

Contacts with studio managers and engineers led to better work. Once he spliced tape edits for Carlos Santana when an engineer failed to appear. Another time Sandy Pearlman left him in charge of finishing vocals on a Blue Oyster Cult session. He produced records for over a decade, working with famous musicians and with many talented unknowns who never made it.

Questions that sent him back to school

These experiences raised questions. Why do some musicians become household names while others stay obscure? Why does music come easily to some people and not others? Where does creativity come from? Why do some songs move us and others leave us cold? What role does perception play in the uncanny hearing of great musicians and engineers?

To look for answers, Levitin drove to Stanford twice a week with Sandy Pearlman to attend Karl Pribram's neuropsychology lectures. He decided psychology held answers about memory, perception, creativity and the brain that underlies them all. He came away with more questions, as often happens in science, and each one deepened his appreciation of the complexity of music and of human experience.

He cites philosopher Paul Churchland. In the last two hundred years, human curiosity has revealed the fabric of space-time, the makeup of matter, the forms of energy, the origin of the universe and the nature of life through DNA. The human genome was fully mapped just five years before the book. One mystery remains: how the brain gives rise to thoughts, feelings, hopes, desires, love and the experience of beauty, along with dance, visual art, literature and music. Levitin asks what music is, where it comes from, and why some sound sequences move us while barking dogs or screeching cars bother many people.

Artists, scientists, and "truth for now"

Some people feel that dissecting music resembles studying the chemistry of a Goya painting instead of seeing the art. Levitin answers with two sources. Oxford historian Martin Kemp observes that artists usually describe their work as experiments, a series of efforts to explore a shared concern or establish a viewpoint. His friend William Forde Thompson, a music cognition scientist and composer at the University of Toronto, notes that artists and scientists pass through similar stages. First comes a creative, exploratory brainstorming phase. Then come testing and refining phases that use set procedures but still call for creative problem-solving.

Studios and laboratories are also alike. Both run many projects at once in various stages of incompletion, and both need specialised tools. Their results are open to interpretation, unlike the plans for a suspension bridge or a bank's end-of-day tally. What the two groups share is a capacity to live with ongoing interpretation and reinterpretation.

Both pursue truth, but both know that truth is contextual, changeable and dependent on viewpoint. Today's truths become tomorrow's disproven hypotheses or forgotten art objects. In science, Piaget, Freud and Skinner held theories that were later overturned or heavily reassessed. In music, some groups were prematurely judged lasting. Cheap Trick were hailed as the new Beatles, and the Rolling Stone Encyclopedia of Rock once gave Adam and the Ants as much space as U2. People once could not imagine the names Paul Stookey, Christopher Cross or Mary Ford being forgotten.

An artist hopes to convey an aspect of universal truth that keeps moving people as societies change. A scientist aims at "truth for now," to be replaced someday, because that is how science advances.

Music's ubiquity and antiquity

Music is unusual among human activities for being both universal and ancient. No known culture, past or present, has lacked it. Some of the oldest artefacts from human and protohuman sites are instruments: bone flutes and animal skins stretched over tree stumps as drums. Music accompanies gatherings of every kind: weddings, funerals, graduations, soldiers going to war, sports events, nights out, prayer, romantic dinners, mothers rocking infants, and students studying. It is especially woven into daily life in nonindustrialised cultures.

Only about five hundred years ago did Western society split into performers and listeners. For most of history and across most of the world, making music was as natural as breathing or walking, and everyone took part. Concert halls are a recent development.

The Sotho and the idea that everyone sings

Levitin's high-school friend Jim Ferguson is now an anthropology professor, funny, brilliant and shy. He did his Harvard doctoral fieldwork in Lesotho, a small country surrounded by South Africa. After he had patiently earned villagers' trust, they invited him to join a song. He said softly that he did not sing. This was true, since despite being a fine oboist in the school band he could not carry a tune.

The villagers found this baffling. The Sotho treat singing as an everyday activity for all, young and old, men and women. They replied, "You talk!" Jim later said it was as strange to them as claiming he could not walk or dance despite having both legs. The Sesotho verb for singing, ho bina, also means to dance, as in many languages, because singing is assumed to involve movement.

Western culture and language, by contrast, separate expert performers such as Arthur Rubinstein, Ella Fitzgerald and Paul McCartney from everyone else, who pay to be entertained. Jim thought that singing or dancing in public implied he considered himself an expert.

Music in modern life and the expert listener

A couple of generations earlier, before television, families made music together. Today the emphasis is on technique and whether someone is good enough to play for others, so music-making has become a reserved activity and most people listen. The music industry is among the largest in the United States and employs hundreds of thousands. Album sales bring in $30 billion a year. That figure excludes concert tickets, the thousands of bands playing Friday nights in bars, and the thirty billion songs downloaded free through peer-to-peer sharing in 2005. Americans spend more on music than on sex or prescription drugs.

Given this, Levitin says most Americans count as expert listeners. We can detect wrong notes, find music we like, remember hundreds of melodies, and tap along in time. Tapping involves a process of meter extraction so complicated that most computers cannot perform it. Two concert tickets can cost a family of four's weekly food allowance, and one CD costs about as much as a work shirt, eight loaves of bread, or a month of basic phone service. Understanding why we like music is therefore a window on human nature.

Evolution and the mind

Asking about a basic, universal ability implicitly asks about evolution. Animals evolved physical forms in response to their environments, and traits that helped in mating passed on through genes. A subtler Darwinian point is coevolution. Organisms change in response to the world while the world, including other species, changes in response to them. If one species develops a defence against a predator, the predator is pressured to overcome it or find other food. Natural selection is thus an arms race of changing physical forms.

Evolutionary psychology, a fairly new field, applies this idea to the mind. Levitin's Stanford mentor, cognitive psychologist Roger Shepard, holds that our minds as well as our bodies are products of millions of years of evolution. This covers our thought patterns, our inclinations to solve problems in certain ways, and our senses, including colour vision and the particular colours we see. Shepard adds that minds coevolved with the physical world. Three of his students lead the field: Leda Cosmides and John Tooby (University of California, Santa Barbara) and Geoffrey Miller (University of New Mexico). They believe that considering the mind's evolution reveals much about behaviour.

So Levitin asks what function music served as humans evolved. Music of fifty or a hundred thousand years ago differed greatly from Beethoven, Van Halen or Eminem. Brains evolved and so did the music we make and want to hear. He also asks whether particular brain regions and pathways evolved specifically for making and hearing music.

Music across the brain

The old notion held that art and music live in the right hemisphere and language and mathematics in the left. Findings from Levitin's laboratory and his colleagues' show instead that music is distributed throughout the brain. Brain-damaged patients illustrate this. Some can no longer read a newspaper but can still read music. Others can play piano but lack the coordination to button a sweater. Listening, performing and composing engage nearly every brain area identified so far and nearly every neural subsystem. He wonders whether this explains claims that music exercises other mental abilities, such as the idea that twenty minutes of Mozart a day makes us smarter.

Music and emotion

Music's emotional power is used by advertisers, filmmakers, military commanders and mothers. Advertisers use it to make a soft drink, beer, running shoe or car seem hipper than rivals. Directors use it to tell us how to feel about ambiguous scenes or to heighten dramatic moments, as in a chase scene or a lone woman climbing the stairs of a dark mansion. We accept, and often enjoy, this manipulation. Mothers everywhere and always have used soft singing to soothe babies to sleep or distract them from distress.

Jargon, and knowing what you like

Many music lovers say they know nothing about it. Levitin finds that colleagues who study hard topics like neurochemistry or psychopharmacology feel unready for music research. Music theory has an arcane vocabulary as obscure as the most esoteric mathematics. Notation looks like set theory to a nonmusician, and keys, cadences, modulation and transposition can bewilder.

Yet every intimidated colleague can name the music they like. His friend Norman White is a world authority on the rat hippocampus and how rats remember places. White is a jazz fan who can talk expertly about favourite artists. He can instantly tell Duke Ellington from Count Basie, and early Louis Armstrong from late. He knows no music theory and cannot name chords, but he is an expert in knowing what he likes.

Levitin compares this to preferring one restaurant's chocolate cake to the coffee shop's. Only a chef could break the taste down into flour, shortening and chocolate type. Every field has specialised vocabulary, as a full blood-analysis report shows. But he thinks musicians, theorists and cognitive scientists should make their work more accessible, and this book tries to do so. The gap between performers and listeners has been paralleled by one between people who love music and talk about it, and those discovering how it works.

Fear that knowledge destroys pleasure

Levitin's students often confess they love life's mysteries and fear that too much education will rob them of simple pleasures. Sapolsky's students probably say the same. Levitin felt this anxiety in 1979 when he moved to Boston to attend the Berklee College of Music. He worried that scholarly analysis would strip music of mystery and that he would know so much he could no longer enjoy it.

That did not happen. He enjoys music as much as he did with that cheap hi-fi and those headphones. The more he learned about music and science, the more fascinating they became, and the more he valued people who excel at them. Like science, music has been an adventure never experienced the same way twice, and a source of continual surprise. Science and music turn out to be a good mix.

What the book will do

The book treats the science of music from cognitive neuroscience, the intersection of psychology and neurology. It will discuss recent studies by Levitin and others on music, musical meaning and musical pleasure. Two puzzles frame it. If we all hear music differently, why do pieces like Handel's Messiah or Don McLean's "Vincent (Starry Starry Night)" move so many? If we all hear it the same way, why do tastes differ so much, with one man's Mozart being another man's Madonna?

He credits recent advances for opening up the mind: brain imaging, drugs that manipulate neurotransmitters such as dopamine and serotonin, and ordinary scientific pursuit. Less known are advances in modelling how neurons network, thanks to the computer revolution. We understand the brain's computational systems as never before. Language appears substantially hardwired. Consciousness is no longer hopelessly mystical but seems to emerge from observable physical systems. Yet no one has gathered this work to illuminate what Levitin calls the most beautiful human obsession. Your brain on music is a way into the deepest mysteries of human nature, which is why he wrote the book.

Understanding what music is and where it comes from may clarify our motives, fears, desires, memories and communication generally. He poses further questions. Is listening like eating when hungry, satisfying an urge? Or is it like seeing a sunset or getting a backrub, triggering the brain's sensory pleasure systems? Why do people grow stuck in their tastes as they age and stop trying new music? The book tells the story of how brains and music coevolved: what music teaches about the brain, what the brain teaches about music, and what both teach about ourselves.


Chapter 1: What Is Music? From Pitch to Timbre

Opening: what counts as music

Levitin begins by showing that people disagree about what music is. For some it is only the great classical masters, for others it is hip-hop and electronica. One of his saxophone teachers at Berklee, like many traditional jazz fans, thought anything made before 1940 or after 1960 was not music. Friends in his 1960s childhood came to his house to hear the Monkees because their parents allowed only classical music, and others were allowed only hymns. When Bob Dylan played an electric guitar at the 1965 Newport Folk Festival, people walked out and some of those who stayed booed. The medieval Catholic Church banned polyphony, fearing it would make people doubt the unity of God. It also banned the augmented fourth, or tritone (C to F-sharp, the interval sung on "Maria" in West Side Story), as too dissonant, calling it the devil in music. Levitin sums up the two episodes: pitch caused the church's uproar, and timbre got Dylan booed.

He then points to avant-garde composers such as Francis Dhomont, Robert Normandeau and Pierre Schaeffer. They record found sounds (jackhammers, trains, waterfalls), edit them, alter their pitch and arrange them into collages with the same tension and release as traditional music. He compares them to painters who left realism behind, such as the cubists, Dadaists, Picasso, Kandinsky and Mondrian.

This raises the question of what Bach, Depeche Mode and John Cage share, and what separates a Busta Rhymes track or a Beethoven sonata from the noise of Times Square or a rainforest. He adopts Edgard Varèse's definition: music is organized sound. The book will take a neuropsychological view, but first he describes what music is made of.

The building blocks of music

The basic attributes of sound are loudness, pitch, contour, duration (rhythm), tempo, timbre, spatial location and reverberation. The brain organizes these into higher-level concepts, as a painter arranges lines into forms. Those concepts are meter, harmony and melody. Levitin defines each attribute:

These attributes can be varied independently, which is why they count as dimensions and can be studied one at a time. Music differs from random sound in how they combine and relate. The combinations give rise to higher-order concepts:

Visual art and dance work the same way. Visual perception has color (hue, saturation, lightness), brightness, location, texture and shape, but a painting becomes art through the relationships between lines and colors, which produce form, flow, perspective, foreground and emotion. Dance is likewise more than unrelated movements. Music also depends on which notes are left out. Miles Davis described his soloing the way Picasso described using a canvas: what matters most is the space between objects, the "air" between notes. Knowing exactly when to hit the next note and letting the listener anticipate it is a mark of his genius, especially on Kind of Blue.

Against jargon, and the definition of timbre

Levitin notes that technical terms (diatonic, cadence, even key and pitch) can intimidate nonmusicians, and that reviewers sometimes hide behind them. He gives mock reviews about an appoggiatura and a modulation to C-sharp minor. Readers really want to know whether the performance moved the audience and whether the singer inhabited the character. He compares this to a restaurant critic discussing the temperature at which lemon juice went into the hollandaise, or a film critic discussing lens aperture.

Even experts disagree on some terms. Timbre is the overall tonal color of an instrument, what separates a trumpet from a clarinet on the same written note, or your voice from Brad Pitt's saying the same words. Because no agreed definition exists, the Acoustical Society of America defines it negatively, as everything about a sound that is not loudness or pitch.

Pitch and frequency

Pitch has generated hundreds of articles and thousands of experiments. It relates to the frequency, or rate of vibration, of a string, air column or other source. A string moving back and forth sixty times a second has a frequency of sixty cycles per second, called Hertz (Hz) after Heinrich Hertz, the first to transmit radio waves. Asked what use radio waves might have, he reportedly said "none." Imitating a siren with your voice sweeps through pitches as vocal fold tension changes.

On a piano, keys at the left strike longer, thicker strings that vibrate slowly. Keys at the right strike shorter, thinner strings that vibrate faster. The string moves air molecules at the same frequency, and they move the eardrum in and out at that frequency. That motion is the only pitch information the brain receives, so the inner ear and brain must work out what vibrations in the world caused it.

Calling slow vibrations "low" and fast ones "high" is a convention. The Greeks used the opposite labels because their stringed instruments were vertical: short strings and organ tubes had their tops closer to the ground, so they were "low," while longer ones reached toward Zeus and Apollo. Levitin rejects the argument that the labels are intuitive because birds are high and bears are low. Thunder is low and comes from above, while crickets and crushed leaves are high and come from below.

As a first definition, pitch is the quality that mainly distinguishes one piano key from another. A hammer strikes a string, which stretches, springs back, overshoots, and oscillates with shrinking distance until it stops, which is why the note fades. The distance traveled is heard as loudness and the rate as pitch. These are independent. Distance depends on how hard the key is hit. Rate depends mainly on the string's size and tension.

Pitch is the mental representation of a sound's fundamental frequency, a wholly psychological phenomenon. Sound waves themselves do not have pitch, and it takes a human or animal brain to map them to it. Isaac Newton saw the same thing with color: he wrote that the waves themselves are not colored. Light has frequencies, and the retina sets off neurochemical events that produce the internal image of color. (Levitin adds that Newton, like Einstein, was a poor student who was eventually expelled.) What we perceive as color is not made of color, an apple's atoms are not red, and as Daniel Dennett notes, heat is not made of tiny hot things. Levitin offers further analogies: pudding has taste only in the mouth, and his kitchen walls are not "white" when nobody is looking. Sound waves hit the eardrums and pinnae and set off mechanical and neurochemical events that end in pitch. So the old question, first posed by George Berkeley, about a tree falling unheard is answered with no. Sound is a mental image, and a measuring device can register the frequency but it is not pitch until heard.

The range of hearing and of musical pitch

Vibrations can in theory run from just above 0 to 100,000 cycles per second or more, but each animal hears a subset, just as visible color is a small slice of the electromagnetic spectrum. Healthy humans hear roughly 20 Hz to 20,000 Hz. Low-end sounds are an indistinct rumble, like a passing truck (about 20 Hz) or a loud car subwoofer. Anything below 20 Hz is inaudible because of our ears' physiology.

Not everything we can hear sounds musical, since we cannot assign a clear pitch across the whole range, just as the infrared and ultraviolet ends of the spectrum look less well defined than the middle. A figure shows the range of musical instruments. Reference points include:

The glass breaks because every object has a natural vibration frequency. You can hear it by flicking the glass or rubbing a wet finger around a crystal rim. A singer hitting its resonant frequency makes it vibrate itself apart.

A standard piano has eighty-eight keys. Its lowest note is 27.5 Hz, roughly the rate at which a sequence of still images gives the illusion of motion. Film alternates stills with black frames at one forty-eighth of a second, too fast for the visual system to resolve, so we see smooth motion that is not really there. Cards in bicycle spokes show a related effect: at low speed you hear separate clicks, and above a threshold they fuse into a buzz you can hum, a pitch. Even so, the lowest note lacks a distinct pitch for most people, and the extreme top and bottom of the keyboard sound fuzzy, which composers use or avoid on purpose. Above about 6000 Hz, past the top piano note, sounds are mostly high whistling. Above 20,000 Hz most people hear nothing, and by sixty most adults lose hearing above about 15,000 Hz because the hair cells of the inner ear stiffen. The part of the keyboard with the strongest sense of pitch is about three quarters of the keys, roughly 55 Hz to 2000 Hz.

Pitch, emotion, melody and learning

Pitch is one of the main carriers of musical emotion, with mood, excitement, calm, romance and danger signaled by several factors. A single high note can convey excitement and a single low note sadness. Strung together, notes make stronger and subtler statements. A melody is the pattern of relations among successive pitches over time, so people easily recognize one played higher or lower than they have heard it. Many melodies, such as "Happy Birthday," have no correct starting pitch. A melody is an abstract prototype, an auditory object that keeps its identity through changes in key, tempo and instrumentation, as a chair stays a chair when moved, flipped or painted red. A song played louder is still the same song, and absolute pitches can shift so long as the relative distances hold.

Speech works similarly. A question's intonation rises at the end without matching any particular pitch. Linguists call this a prosodic cue, and it is a convention that must be learned, as is not the case in all languages. Western music has similar conventions: some pitch sequences evoke calm and others excitement, and the basis is mostly learning. Everyone can learn the linguistic and musical distinctions of their birth culture, and exposure shapes neural pathways so that people internalize the rules of their tradition.

Instruments occupy different parts of the pitch range, and this shapes how they convey emotion. The piano has the widest range. The high, shrill, birdlike piccolo tends to evoke flighty, happy moods whatever it plays, so composers use it for happy or rousing music like a Sousa march. A composer would give it sad sequences only for irony. In Prokofiev's Peter and the Wolf the flute stands for the bird and the French horn for the wolf, each with a leitmotiv, a melodic phrase tied to an idea, person or situation (a device especially associated with Wagner). Tuba and double bass evoke solemnity, gravity or weight.

How many pitches, and how the ear maps them

Because pitch comes from a continuum, there are in principle infinitely many, but not every frequency change is audible, just as a grain of sand does not change a backpack's weight. Most cultures use steps no smaller than a semitone, and most people cannot reliably hear changes under about a tenth of a semitone. Ability varies among people and animals, and training can help.

The basilar membrane in the inner ear holds hair cells that are frequency-selective, arranged from low to high like a keyboard laid across it. This is a tonotopic map. Cells fire according to frequency and send signals to the auditory cortex, which has its own tonotopic map. Pitch is so important that the brain represents it directly. Unlike almost any other musical attribute, electrodes in the brain could reveal which pitches were played. Paradoxically, music rests on relations between pitches, yet the brain attends to absolute values at its various stages.

Scales, tuning and note names

A scale is a subset of the infinite pitches, chosen by each culture from tradition or arbitrarily. Letter names such as A, B and C are arbitrary labels for frequencies. In Western music these are the only legal pitches, and most instruments are built to play only them. Trombone, cello and violin can slide between notes, so those players spend years learning to produce the exact frequencies. In-between sounds count as mistakes ("out of tune") unless used for expressive intonation, deliberately going briefly out of tune to add tension, or while passing between legal tones.

Tuning is the exact relation between a tone's frequency and a standard, or between tones sounded together. Orchestras tune up because wood, metal and strings drift with temperature and humidity, matching a standard or sometimes each other. Skilled players also bend pitch for expression, except on fixed-pitch instruments such as keyboards and xylophones, and adjust to ensemble mates who drift.

Note names run A to G, or in the alternate system do-re-mi-fa-sol-la-ti-do, which Rodgers and Hammerstein used in "Do-Re-Mi" from The Sound of Music. Names climb with frequency and restart after G. Notes with the same name have frequencies that are multiples of each other. If one A is 55 Hz, other As are two, three, four, five or a half times that.

The octave

Names repeat because doubling or halving a frequency yields a note that sounds remarkably like the original. This 2:1 ratio is the octave. Every known musical culture, whether Indian, Balinese, European, Middle Eastern or Chinese, uses the octave as its basis, whatever else differs. This gives circularity in pitch perception, like circularity in color, where red and violet are at opposite ends of the spectrum but look similar. Music thus has two dimensions: tones rising higher and higher, and the sense of coming home with each doubling.

Examples of octaves:

Intervals, semitones and the keyboard

An interval is the distance between two tones. Western music divides the octave into twelve logarithmically equal steps. A whole step (the author avoids the confusing term "tone") separates A and B, or do and re. The semitone, half a whole step, is one twelfth of an octave. Melody is relational, so it is defined by intervals, not the specific notes. Four semitones always make a major third, starting on A, G-sharp or anything. A table of intervals is given (its text is lost in the source). The perfect fourth and perfect fifth are so named because many people find them especially pleasing, and since the Greeks they have been at the heart of music, a backbone for at least five thousand years. There is no "imperfect fifth," it is only a name.

The brain areas responding to single pitches have been mapped, but the neural basis for encoding pitch relations has not. We know which cortex is involved in hearing C and E, or F and A, but not why both are heard as a major third. Those relations must come from poorly understood computations.

Why twelve notes but seven letters? Levitin jokes that it may be an invention by musicians, long relegated to servants' quarters, to make nonmusicians feel inadequate. The extra five have compound names such as E-flat and F-sharp. On the piano, white and black keys are unevenly arranged, but adjacent keys are always a semitone and keys two apart are a whole step. The same applies to guitar frets and to woodwind keys on clarinet and oboe.

White keys are A to G, and black keys carry compound names. The note between A and B is A-sharp or B-flat, interchangeable outside formal theory (and could also be called C double-flat). Sharp means high and flat means low. In do-re-mi, other syllables such as di and ra mark the in-between tones. These notes are not second-class. Some songs and scales use them exclusively, and the main accompaniment to Stevie Wonder's "Superstition" is on black keys only. The twelve tones and their octave repeats are the building blocks of every song in our culture, from "Deck the Halls" to "Hotel California," "Ba Ba Black Sheep" and the Sex and the City theme.

Musicians also use sharp and flat for out-of-tune playing: a tone slightly too high is sharp, too low is flat. Small errors go unnoticed, but when off by a quarter to half the distance to the next note, most listeners notice, particularly when it clashes with in-tune instruments.

A440 and equal spacing

Our standard is A440: the A in the middle of the piano is fixed at 440 Hz. This is arbitrary, and it could be 439, 444, 424 or 314.159. Mozart's time used other standards. Some claim the exact frequencies change the sound. Led Zeppelin often tuned away from A440 to get an uncommon sound, perhaps to link with the European children's folk songs behind some of their compositions. Baroque purists want period instruments partly because they play in the original tuning.

Frequencies can be fixed anywhere because music is defined by relations. The distance from note to note is not arbitrary: each note sounds equally spaced to our ears (not necessarily to other species'), though the change in Hz is not equal. Each note is about 6 percent higher in frequency than the one before. The auditory system is sensitive to relative, proportional change. Levitin illustrates with weight lifting: going from 5 to 50 pounds, adding 5 pounds each week is a doubling at first and a small increase later. Equal increases of effort mean adding a constant percentage, such as 50 percent: 5, 7.5, 11.25, 16.83. In the scale, twelve steps of 6 percent double the frequency; the exact ratio is the twelfth root of two, 1.059463.

Chromatic and major scales

The twelve notes form the chromatic scale. Any scale is a set of distinguishable pitches chosen as a basis for melodies. Western composers usually use seven (sometimes five) of the twelve. The most common seven-note scale is the major scale, or Ionian mode, of ancient Greek origin. It can start on any note, and it is defined by its pattern of steps: whole, whole, half, whole, whole, whole, half. Starting on C gives C-D-E-F-G-A-B-C, all white keys. Other major scales need black keys to keep the pattern. The starting note is the root.

The positions of the two half steps matter for the scale's identity and for musical expectation. Experiments show young children and adults learn melodies better from scales with unequal distances. The half steps orient experienced listeners within the scale. Everyone who has absorbed the style knows that a B in the key of C is the seventh degree and only a half step below the root, though most can't name the notes or know "root" or "scale degree." That knowledge is gained through passive exposure, not innate, like learning sunrise and sunset without cosmology.

The minor scale and how we infer key

Other step patterns yield other scales, the most common being the minor scale. A minor uses only white keys: A-B-C-D-E-F-G-A, so it is the relative minor of C major. Its pattern is whole, half, whole, whole, half, whole, whole. The half steps sit before the third and sixth degrees, versus before the root and fourth degree in major. There is still pull toward the root, but the chords producing it sound and feel different.

How do we know which scale we are in when the same white keys are played? Without awareness, the brain tracks how often notes sound, whether they fall on strong or weak beats, and how long they last, then infers the key. This works without training or declarative knowledge, the ability to talk about it. Listeners sense the intended tonal center, or key, and notice a return to the tonic or a failure to return. The simplest way to set a key is to play its root often, loud and long. If a composer who thinks he is in C major repeats a loud, long A, starts and ends on A and avoids C, listeners and theorists will hear A minor. Levitin compares it to speeding tickets: observed action counts, not intent.

Emotion, culture and scales

Culturally, major scales go with happy or triumphant feelings and minor with sad or defeated ones. Some studies suggest the link is innate, but since it is not universal, any innate tendency can be overridden by cultural exposure. Western theory recognizes three minor scales, each with a different flavor. Blues generally uses a five-note pentatonic scale that is a subset of the minor scale, and Chinese music uses a different pentatonic scale. Tchaikovsky chose scales typical of Arab or Chinese music in The Nutcracker, and within a few notes we are transported. Billie Holiday uses the blues scale to make a standard bluesy.

Composers use these associations on purpose, and brains know them from a lifetime of exposure. We link new patterns with whatever sights, sounds and sensory cues accompany them and form memory links to place, time or events. Anyone who has seen Hitchcock's Psycho thinks of the shower scene when hearing Bernard Herrmann's screeching violins. Anyone who has watched Warner Bros. "Merrie Melody" cartoons thinks of a sneaky character on stairs when hearing plucked violins on an ascending major scale. A few notes suffice, such as the first three of David Bowie's "China Girl" or Mussorgsky's "Great Gate of Kiev."

Nearly all this variety comes from dividing the octave, and in virtually every known case into no more than twelve tones. Claims that Indian and Arab-Persian music use microtuning, with steps much smaller than a semitone, fail on close analysis. Their scales also use twelve or fewer tones, and the rest are expressive variations: glissandos (continuous glides) and passing tones, like sliding into a note in American blues.

Tonal hierarchy, chords and early learning

Within any scale the tones differ in stability and finality, producing tension and resolution. In major, the most stable tone is the first degree, the tonic, and the others point toward it with varying force. The seventh degree (B in C major) points most strongly. The fifth (G in C) points least strongly because it feels stable, so ending a song on it does not leave us uneasy. Carol Krumhansl and colleagues showed that ordinary listeners have absorbed this hierarchy through passive exposure. She played a scale and asked people to rate how well other tones fit, and their subjective ratings recovered the theoretical hierarchy.

A chord is three or more notes played together, usually drawn from a scale to convey information about it. A typical chord uses the first, third and fifth notes. From C major that is C, E, G; from C minor, C, E-flat, G. The E versus E-flat third turns a major chord into a minor one. Even untrained listeners hear major as happy and minor as sad, reflective or exotic. The simplest rock and country songs use only major chords, such as "Johnny B. Goode," "Blowin' in the Wind," "Honky Tonk Women" and "Mammas Don't Let Your Babies Grow Up to Be Cowboys." Minor chords add complexity. In the Doors' "Light My Fire" the verses are minor and the chorus major. Dolly Parton mixes minor and major in "Jolene" for melancholy. Pink Floyd's "Sheep" uses only minor chords.

Chords also form a stability hierarchy by context. Every tradition has typical progressions, and by five most children have absorbed which are legal in their culture and can detect deviations as easily as a malformed sentence such as "The pizza was too hot to sleep." That requires neural networks forming abstract representations of musical structure and rules, automatically. Young brains are almost spongelike, soaking up sounds into their wiring. With age the circuits become less pliable, making new musical or linguistic systems harder to absorb deeply.

Overtones and the harmonic series

Pitch gets more complicated because of physics, and this gives instruments their richness. Natural objects vibrate in several modes at once: a piano string, a struck bell, a drum, air in a flute. Levitin offers analogies: the Earth spins daily, orbits the sun in 365.25 days, and travels with the solar system in the Milky Way; and a train, where you feel wind rocking the car about twice a second, then engine vibration, then wheel bumps over track joints. You know there are several vibrations but cannot easily count them or find their rates without instruments.

So a single note on any instrument, even a drum or cowbell, comprises many pitches at once, though most people are not consciously aware (some can train themselves to hear it). The slowest, lowest is the fundamental frequency, and the rest are overtones. They are often related by simple integer multiples. A string whose fundamental is 100 Hz also vibrates at 200 and 300 Hz. A flute at 310 Hz also sounds at 620, 930 and 1240. Such sounds are harmonic, and the pattern is the overtone series. Evidence suggests auditory cortex neurons respond to the components by firing in synchrony, a basis for the sound's coherence.

The brain is so tuned to this series that it fills in a missing fundamental. A sound with energy at 100, 200, 300, 400 and 500 Hz has a pitch of 100 Hz. Remove the 100 Hz and we still hear 100, not 200, because a normal 200 Hz sound would have 200, 400, 600, 800. Sequences that depart from the series, such as 100, 210, 302, 405, shift perceived pitch as a compromise between what is presented and what a harmonic series implies.

Levitin recounts that his graduate advisor, Mike Posner, told him about biology graduate student Petr Janata, a ponytailed, tie-dyed jazz and rock pianist who felt like a kindred spirit. Janata put electrodes in the inferior colliculus of the barn owl, part of the auditory system, and played a version of Strauss's "The Blue Danube Waltz" with the fundamentals removed. He predicted that if the missing fundamental is restored early in processing, the neurons would fire at the missing fundamental's rate, and they did. He routed the electrode signals through an amplifier to a loudspeaker, and the melody was clearly audible in the neurons' firing, which matched the missing fundamental's frequency. The overtone series therefore had a neural instantiation at early stages and in a different species.

Levitin argues that any advanced alien species would have some way to sense vibration. Wherever there is atmosphere, molecules vibrate, and knowing whether something is making noise or approaching, even unseen (in the dark, unattended, or while asleep), has survival value. Since most objects vibrate in several modes with simple integer relations, the overtone series should be found everywhere: North America, Fiji, Mars, planets around Antares. An organism evolving among vibrating objects should evolve brain machinery reflecting those regularities. Because pitch is a basic cue to identity, we should expect tonotopic maps and synchronous firing for octave and harmonic relations, helping a brain judge that the tones come from one object.

Terminology: the first overtone is the first frequency above the fundamental. A parallel harmonics system calls the fundamental the first harmonic, so the second harmonic equals the first overtone (Levitin says physicists seem to enjoy confusing undergraduates). Not all instruments vibrate so neatly. A piano, being percussive, has overtones close to but not exactly integer multiples, which adds to its sound. Percussion, chimes and other objects often have overtones clearly not integer multiples, called partials or inharmonic overtones. Such instruments lack a clear sense of pitch, possibly through lack of synchronous firing, yet still have some, which shows when notes are played in succession. A single woodblock or chime note cannot be hummed, but a melody on a set can be recognized because the brain follows changes in overtones. This is what happens when people play songs on their cheeks.

Timbre

A flute, violin, trumpet and piano can play the same written note with the same fundamental and the same heard pitch, yet sound very different. This is timbre (TAM-ber), the most important and ecologically relevant feature of auditory events. It separates a lion's growl from a cat's purr, thunder from ocean waves, a friend's voice from a bill collector's. Human timbral discrimination is so acute that most people recognize hundreds of voices and can tell if a close person is happy, sad, healthy or getting a cold.

Timbre results from overtones. Materials differ in density (metal sinks in a pond, wood floats), and density, size and shape affect the noise made when struck. A hammer tapped gently on a guitar gives a hollow wooden plunk, on a saxophone a tinny plink. Struck objects vibrate at several frequencies set by material, size and shape, and the intensities of the harmonics need not be equal and typically are not.

A saxophone playing a 220 Hz fundamental also sounds 440, 660, 880, 1200, 1420, 1640 and so on, with different intensities heard as different loudnesses. This loudness pattern gives the saxophone its color. A violin playing the same 220 Hz note has overtones at the same frequencies but a different loudness pattern. Each instrument has a unique pattern, like a fingerprint, and virtually all tonal variation, what gives a trumpet "trumpetiness" and a piano "pianoness," comes from how overtone loudness is distributed. Examples:

Individual instruments of one type differ too, though less than they differ from other types. Levitin jokes that to him all accordions sound alike and that the sweetest sound they could make would be burning in a bonfire. Master players can tell a Stradivarius from a Guarneri within a note or two. He can clearly distinguish his 1956 Martin 000-18, 1973 Martin D-18 and 1996 Collings D2H, which sound like different instruments.

Synthesis and the Stanford connection

Natural instruments produce energy at several frequencies because of their molecular structure. Levitin imagines a hypothetical "generator" producing only one frequency. A bank of them set to 110, 220, 330, 440, 550 and 660 Hz, with amplitudes matched to an instrument's overtone profile, would approximate a clarinet, flute or other instrument. This is additive synthesis. Pipe organs let you try it. Each key sends air through pipes of different sizes, each pitched by length like a mechanical flute, and one key sounds several pipes, the extras producing integer multiples or closely related tones. Organists pull drawbars to choose them, so knowing a clarinet is strong in odd harmonics, a clever organist could imitate it with a bit of 220 Hz, a dash of 330, a dollop of 440 and a heaping helping of 550.

From the late 1950s scientists built such synthesis into compact electronic devices, synthesizers. By the 1960s they appeared on Beatles records ("Here Comes the Sun," "Maxwell's Silver Hammer") and on Walter/Wendy Carlos's Switched-On Bach, followed by bands built around them, such as Pink Floyd and Emerson, Lake and Palmer. Many used additive synthesis, and later ones used more complex methods such as wave guide synthesis (Julius Smith, Stanford) and FM synthesis (John Chowning, Stanford). Copying only the overtone profile gives a pale imitation, so timbre involves more. Researchers still argue about what, but it is generally agreed that attack and flux are the other two attributes.

Levitin describes Stanford's setting near San Francisco: pastureland to the west, the Central Valley (raisins, cotton, oranges, almonds) to the east, Gilroy's garlic and Castroville, the "artichoke capitol," to the south. He once suggested the Castroville Chamber of Commerce change "capitol" to "heart," with little enthusiasm. Stanford is a home for computer scientists and engineers who love music. John Chowning, an avant-garde composer and music professor since the 1970s, pioneered using computers to create and store sound, and became founding director of CCRMA (pronounced CAR-ma, with a joke that the first c is silent). He is warm, and as an undergraduate Levitin found he would put a hand on his shoulder and ask about his work, treating a student as a chance to learn.

In the early 1970s, playing with sine waves, the building blocks of additive synthesis, Chowning noticed that changing their frequency while they played produced musical sounds. Controlling the parameters let him simulate many instruments. This was frequency modulation, or FM, synthesis, first built into Yamaha's DX9 and DX7 synthesizers, which revolutionized the industry on their 1983 release. Earlier synthesizers were expensive, clunky and hard to control, and new sounds needed time and know-how. FM let any musician get a convincing instrument sound at the touch of a button, so songwriters who could not hire a horn section or orchestra could try textures, and orchestrators could test arrangements first. The Cars, the Pretenders, Stevie Wonder, Hall and Oates and Phil Collins used it widely, and much of the "eighties sound" comes from it. The royalties let Chowning build CCRMA and attract students and faculty. Early celebrities there included John R. Pierce and Max Mathews.

John Pierce and the six songs

Pierce had been vice president of research at Bell Telephone Laboratories and supervised the team that built and patented the transistor, which he named (TRANSfer resISTOR). He is credited with the traveling wave vacuum tube and with launching the first telecommunications satellite, Telstar, and wrote science fiction as J. J. Coupling. He fostered an environment in which scientists felt empowered and creativity was valued. With AT&T's monopoly and cash reserves, Bell Labs was a playground, and Pierce let people be creative without worrying about profit or commercial use, since innovation needs freedom from self-censorship. Only a few ideas would be practical and fewer become products, but those would be innovative and profitable. Lasers, digital computers and the Unix operating system came from this environment.

Levitin met Pierce in 1990, when Pierce was eighty and lecturing on psychoacoustics at CCRMA. After his Ph.D. and return to Stanford they became friends, dining every Wednesday and discussing research. Pierce, knowing of Levitin's music-business past, asked him to explain rock and roll, which he had never understood, by playing six songs that captured it. The night before, Pierce called to say he had heard Elvis, so Elvis was dropped. Levitin brought:

  1. "Long Tall Sally," Little Richard
  2. "Roll Over Beethoven," the Beatles
  3. "All Along the Watchtower," Jimi Hendrix
  4. "Wonderful Tonight," Eric Clapton
  5. "Little Red Corvette," Prince
  6. "Anarchy in the U.K.," the Sex Pistols

Some choices paired great songwriters with different performers, and he would now make adjustments. Pierce kept asking who the artists were, what instruments he heard and how the sounds arose. He liked the timbres most, while the songs and rhythms interested him less. He found the sounds new and exciting, such as Clapton's fluid romantic solo with soft pillowy drums, and the Sex Pistols' dense wall of guitars, bass and drums. Distorted guitar was only part of it: the way bass, drums, electric and acoustic guitars and voice combined into a unified whole was new to him. For Pierce, timbre defined rock, and it was a revelation to both.

Timbre's growing role in Western music

Scales have changed little since the Greeks, except for the refinement into the equal-tempered scale in Bach's time. Rock and roll may be the last step in a millennium-long revolution that gave perfect fourths and fifths a prominence once reserved for the octave. For much of that time Western music was dominated by pitch, but for about two hundred years timbre has become increasingly important. Restating a melody on different instruments is standard across genres, from Beethoven's Fifth and Ravel's "Bolero" to the Beatles' "Michelle" and George Strait's "All My Ex's Live in Texas." New instruments were invented to widen the timbral palette, and we enjoy a melody repeated in a new timbre, as when a country or pop singer stops and another instrument takes the tune.

Attack, steady state and flux

In the 1950s Pierre Schaeffer did his famous "cut bell" experiments. He taped orchestral instruments and used a razor blade to cut off the beginnings of the sounds. This initial part is the attack, the sound of the first hit, strum, bow or blow. The body gesture that makes the sound strongly influences it, though most of that fades within seconds. Gestures are mostly impulsive, short bursts. With percussion the player usually stops touching the instrument afterward. With wind and bowed instruments contact continues, in smoother, less impulsive blowing or bowing.

The attack usually produces energy at many frequencies not related by integer multiples, so it has a noisy quality, more like a hammer on wood than on a bell or piano string, or wind rushing through a tube. Next comes a stable phase in which the material resonates and the tone takes on its orderly overtone pattern. This middle is the steady state, when the overtone profile is relatively stable.

With the attacks removed, most people could hardly identify the instruments. Pianos and bells sounded unlike themselves and much like each other. Splicing one instrument's attack onto another's steady state gives varied results. Sometimes you hear an ambiguous hybrid that sounds more like the instrument that gave the attack. Michelle Castellengo and others found that new instruments can be made this way, such as a violin bow sound spliced onto a flute tone sounding strongly like a hurdy-gurdy street organ. The experiments showed the importance of attack.

The third dimension, flux, is how the sound changes after it starts. A cymbal or gong has much flux, changing dramatically over time, while a trumpet has less and is more stable. Instruments also sound different across their ranges. Sting's straining, reedy high register in "Roxanne" conveys urgent pleading, unlike the deliberate, longing low sound at the start of "Every Breath You Take," which suggests a dull ache long endured but not yet at breaking point.

Timbre as a compositional tool

Timbre is more than the sounds instruments make. Composers choose instruments and combinations for emotion, atmosphere and mood. Examples are the almost comical bassoon opening Tchaikovsky's "Chinese Dance" in the Nutcracker Suite, and the sensuousness of Stan Getz's saxophone on "Here's That Rainy Day." Replace the electric guitars of the Rolling Stones' "Satisfaction" with piano and you have a different animal. Ravel used timbre as a device in Bolero, repeating the theme with different timbres, after brain damage impaired his ability to hear pitch. Jimi Hendrix is recalled mostly for the timbre of his guitars and voice.

Scriabin and Ravel called their works sound paintings in which notes and melodies are like shape and form and timbre is like color and shading. Stevie Wonder, Paul Simon and Lindsey Buckingham have described their songs that way, with timbre doing what color does in visual art, separating melodic shapes. Music differs from painting by being dynamic, changing across time, and what drives it forward is rhythm and meter. They are the engine of virtually all music and probably the first elements our ancestors used for protomusics, a tradition still heard in tribal drumming and the rituals of preindustrial cultures. Levitin ends by saying that he believes timbre is now central to our appreciation of music, while rhythm has held supreme power over listeners far longer.


Chapter 2: Foot Tapping — Discerning Rhythm, Loudness, and Harmony

Opening: Sonny Rollins and the primacy of rhythm

Levitin opens with a memory of seeing Sonny Rollins, one of the most melodic of saxophonists, in Berkeley in 1977. Almost thirty years on, he cannot recall a single pitch from the night, but he remembers rhythms. At one point Rollins improvised for three and a half minutes on one repeated note, varying only rhythm and subtle timing, and it was this, not melodic invention, that brought the crowd to its feet. Levitin uses this to make a general point: nearly every culture treats movement as part of making and hearing music, and rhythm is what people dance, sway and tap to. Even in jazz, the drum solo often excites audiences most. Playing an instrument means moving the body in coordinated, rhythmic ways and passing that energy into the instrument. Neurally, this recruits primitive "reptilian" structures (the cerebellum and brain stem), the motor cortex, and the planning regions of the frontal lobes, the most advanced part of the brain.

Rhythm, tempo and meter distinguished

Three related and often confused terms are defined. Rhythm is the lengths of notes and how they relate to one another. Tempo is the pace of the music, the rate at which you would tap your foot. Meter is the pattern of hard and light taps and how these group into larger units.

Rhythm and the 2:1 ratio

Rhythm is the relationship between the durations of notes, and it is a key ingredient in turning sound into music. Levitin gives the history of the "shave-and-a-haircut, two bits" rhythm, often used as a knock on a door: its first documented use was in an 1899 Charles Hale recording, "At a Darktown Cakewalk"; lyrics were attached in a 1914 song by Jimmie Monaco and Joe McCarthy; a 1939 song by Dan Shapiro, Lester Lee and Milton Berle used it with "shampoo," and how that became "two bits" is a mystery; Leonard Bernstein scored it in "Gee, Officer Krupke." It consists of long and short notes, the long ones twice the short ones.

The same two-length pattern appears in Rossini's William Tell overture (the Lone Ranger theme), in "Mary Had a Little Lamb" (six equal notes then one about twice as long), in the Mickey Mouse Club theme (three duration levels, each double the last), and in The Police's "Every Breath You Take" (durations 1, 1, 2, 2, 4 in arbitrary units). He suggests the 2:1 rhythmic ratio is a musical universal, like the octave in pitch.

Real music is rarely this simple. Just as a particular scale can evoke a particular culture, so can a particular arrangement of rhythms; most people could not reproduce a complex Latin rhythm but immediately recognise it as Latin rather than Chinese, Arabic, Indian or Russian. Organising rhythms into strings of varying lengths and emphases produces meter and establishes tempo.

Tempo, the beat, and tempo in real songs

Tempo is how quickly or slowly a piece goes by, like a song's gait or heartbeat. The beat (or tactus) is the basic unit, usually where you naturally tap, clap or snap. Some people tap at half or double the beat, owing to differences in neural processing, background and interpretation, and even trained musicians can disagree about tapping rate. They do agree on the underlying speed; the disagreement is over subdivisions or superdivisions.

Examples of tempos in beats per minute: Paula Abdul's "Straight Up" and AC/DC's "Back in Black" are both 96 (so you would step 96 or 48 times a minute, not 58 or 69); Aerosmith's "Walk This Way" 112; Michael Jackson's "Billie Jean" 116; the Eagles' "Hotel California" 75. In "Back in Black" the drummer's high-hat sounds the steady 96 at the start.

Equal tempos can feel very different. "Back in Black" has the cymbal playing eighth notes (two per beat) and a simple syncopated bass locked with the guitar. "Straight Up" is dense: complex, irregular drumming with sixteenth-note bursts and gaps ("air") typical of funk and hip-hop, a similarly syncopated bass line, and in the right channel a Latin shaker, the afuche or cabasa, which alone plays on every beat. Putting the key rhythm on a light, high instrument inverts convention, and synthesizers, guitar and percussion effects stress various beats unpredictably, which gives the song staying power over many listenings.

Tempo, emotion and memory for tempo

Fast tempos tend to be heard as happy and slow as sad; this is an oversimplification but holds across many cultures and across the lifespan. Levitin and Perry Cook published a 1996 experiment in which people sang favourite rock and pop songs from memory, to see how close they came to the recorded tempo. As a baseline, the average listener can detect a tempo change of about 4 percent (for a 100 bpm song, a shift to 96 goes unnoticed by most people and even some professional musicians, though drummers, who must keep time without a conductor, are more sensitive). Most of the nonmusician participants sang within 4 percent of the original tempo.

The likely neural basis is the cerebellum, thought to hold timekeepers for daily life and to synchronise with heard music; it seems to store the settings used for synchronising and recall them when we sing from memory. The basal ganglia, which Gerald Edelman called "the organs of succession," are almost certainly involved in generating and shaping rhythm, tempo and meter too.

Meter: strong and weak beats

Meter is the grouping of beats. Some beats feel stronger, as if played louder and heavier, and weaker ones follow until the next strong one. Every known musical system has such patterns. The commonest Western pattern is a strong beat every four, with the third beat somewhat stronger than the second and fourth, giving a hierarchy: first, third, then second and fourth. Less often there is one strong beat in three, the waltz beat. We count accordingly (ONE-two-three-four; ONE-two-three).

Music would be dull with only straight beats, so beats may be omitted for tension. "Twinkle, Twinkle Little Star," which Mozart wrote at age six, leaves a rest on the fourth beat of alternate bars. "Ba Ba Black Sheep," set to the same tune, subdivides the beat: "have-you-any" goes twice as fast as "ba ba black," because quarter notes are halved.

In Elvis Presley's "Jailhouse Rock" (by Leiber and Stoller) the strong beat falls on the first sung note and every fourth note after. In music with words, syllables do not always match downbeats: "began" begins before a strong beat and ends on it, which suits the natural accent on its second syllable and adds momentum. Simple nursery rhymes and folk songs such as "Frère Jacques" avoid this.

Note durations, time signatures and syncopation

Western notation names durations in a relative way, like intervals (a perfect fifth is seven semitones from any starting note). A whole note lasts four beats regardless of tempo (at 60 bpm, as in the Funeral March, four seconds). A half note is half that, a quarter note half again. In most popular and folk music the quarter note is the basic pulse, and 4/4 means four quarter-note beats to each measure or bar. This only describes the counting; a bar may hold notes of any length or rests. "Ba Ba Black Sheep" has four quarter notes in its first measure and eighth notes plus a quarter rest in the second.

Buddy Holly's "That'll Be the Day" starts with a pickup note ("Well") before the first strong beat, then the strong beat falls on every fourth note, and Holly, like Presley, splits a word ("day") across lines. Most listeners tap four times between downbeats. A tap sometimes lands mid-word: the foot is in the air when "say" starts and comes down partway through it. A note that comes slightly earlier than the strict beat is syncopation, which surprises us and ties into expectation and emotional impact. Some people feel the song in half time (two taps), a valid reading. Holly also violates expectation by delaying words: in lines two and four of the verse the downbeat arrives and he is silent, whereas nursery rhymes put a word on every downbeat. Withholding what we expect builds excitement.

Backbeat, "in two," and waltz time

When people clap or snap spontaneously they often do so on beats two and four, not the downbeat: the backbeat Chuck Berry mentions in "Rock and Roll Music." John Lennon said rock writing meant simple English, a rhyme, and a backbeat. In most rock the snare drum plays only on two and four, opposing the strong beat on one and the secondary strong beat on three, as in Lennon's "Instant Karma." Queen's "We Will Rock You" has two stamps then a clap, repeated; the clap is the backbeat.

Sousa's "The Stars and Stripes Forever" is "in two": you tap once per two quarter notes. Rodgers and Hammerstein's "My Favorite Things" is in 3/4 waltz time, a strong beat followed by two weak.

Quantisation and complex ratios

Small-integer duration ratios are the commonest and are probably easier to process neurally, yet Eric Clarke notes they almost never occur exactly in real performances. Levitin infers a quantisation process in perception: the brain treats similar durations as equal, rounding to simple ratios such as 2:1, 3:1 and 4:1. Some music uses more complex ratios; Chopin and Beethoven have passages with seven or five notes in one hand against four in the other. In theory any ratio is possible, but perception, memory, style and convention limit what is used.

Common and unusual meters

The most common Western meters are 4/4, 2/4 and 3/4; others include 5/4, 7/4 and 9/4. 6/8 (six eighth-note beats) resembles 3/4 but is felt in groups of six with the eighth note as pulse; it can be counted as two groups of three or as six with a secondary accent on four. Listeners find this a subtlety for performers, but brains may differ: there are neural circuits for tracking meter, and the cerebellum sets an internal timer. Nobody has tested whether 6/8 and 3/4 have different neural representations, but since musicians treat them as different, Levitin thinks it likely. He states the principle that behavioural differences must have a neural basis.

Even meters (4/4, 2/4) suit walking, dancing and marching because the same foot always lands on the strong beat; 3/4 is awkward for marching. Five-beat meter is rare: Lalo Schifrin's "Mission: Impossible" theme and Dave Brubeck's "Take Five" (with a secondary accent on four, so many musicians hear alternating 3 and 2; Mission: Impossible has no clear split). Tchaikovsky used 5/4 in the second movement of his Sixth Symphony. Pink Floyd's "Money" and Peter Gabriel's "Solsbury Hill" use 7/4, requiring a count of seven between strong beats.

Loudness

Levitin treats loudness briefly. Like pitch, it is purely psychological: stereo volume changes the amplitude of molecular vibration, but only a brain experiences loudness. Oddities: loudnesses do not add as amplitudes do (loudness is logarithmic); a pure tone's pitch can shift with its amplitude; and sounds can seem louder after processing such as dynamic range compression, common in heavy metal.

Loudness is measured in decibels (after Alexander Graham Bell), a dimensionless ratio like percent and like a musical interval rather than a note name. Doubling a source's intensity adds 3 dB. A log scale suits the ear's sensitivity: the loudest tolerable sound versus the softest detectable is a million to one in sound pressure, or 120 dB. That span is the dynamic range; a recording with 90 dB between softest and loudest is considered high fidelity and beyond most home systems.

The ear compresses very loud sounds to protect the middle and inner ear: inner hair cells have a 50 dB range though we hear over 120 dB, with each 4 dB increase in the world passing as 1 dB to the hair cells. We can often hear this compression as a different quality.

Acousticians use a reference of 20 micropascals, roughly the hearing threshold (a mosquito ten feet away), and write dB (SPL). Landmarks:

Foam earplugs block about 25 dB (not evenly across frequencies); at a Who concert they could bring levels to 100–110 dB. Firing-range and airport ear protectors are often supplemented with inner plugs.

Many people love very loud music and describe a special thrilling state above about 115 dB. The reason is unknown; perhaps loud sound saturates the auditory system so many neurons fire maximally, giving an emergent, qualitatively different brain state. Others simply dislike it.

Loudness is one of seven major elements (pitch, rhythm, melody, harmony, tempo, meter). Tiny changes carry emotional weight: a pianist playing five notes with one slightly louder changes that note's role. Loudness also cues rhythm and meter, since it determines how notes group.

Key, harmony and expectation

Returning to pitch: rhythm is a game of expectation (tapping predicts what comes next), and pitch has its own expectation game with rules of key and harmony. A key is the tonal context of a piece. Some music has none (African drumming, Schoenberg's twelve-tone music), but nearly all Western music, from jingles to Bruckner, Mahalia Jackson to the Sex Pistols, has a tonal centre it returns to. The key can change (modulation) but generally holds for minutes. A melody built on the C major scale is "in C," pulling back to C even if it does not end there; outside notes are departures, like a film flashback after which we know the main plot will resume. (Appendix 2 covers theory.)

A note's sound depends on context: preceding melody and accompanying chords. Levitin compares this to flavour: oregano with tomato sauce versus banana pudding, cream on strawberries versus in coffee. In the Beatles' "For No One" one note is sung for two bars while chords change its mood; Jobim's "One Note Samba" features one note under shifting chords, sounding bright in some and pensive in others. Nonmusicians also recognise chord progressions without melody: the Eagles need play only three chords of the "Hotel California" sequence (he lists it: B minor, F-sharp major, A major, E major, G major, D major, E minor, F-sharp major) before fans know the song, even in a different instrumentation or a Muzak version at the dentist.

Consonance and dissonance

Some sounds are unpleasant without clear reason; fingernails on a chalkboard bother humans, but in the one experiment done, monkeys liked it as much as rock. Some people hate distorted guitars; others like little else. At the harmonic level, musicians call pleasing intervals and chords consonant and unpleasing ones dissonant. There is no agreement why. We know the brain stem and dorsal cochlear nucleus, primitive structures all vertebrates have, can tell consonance from dissonance before the cortex gets involved.

Agreed consonant intervals: the unison (1:1) and octave (2:1), whose waveform peaks half line up and half fall exactly between. Halving the octave gives the tritone, widely judged the most disagreeable interval, and its ratio (given as 43:32) is not simple. A 3:1 ratio is two octaves; 3:2 is the perfect fifth (C to G); G to the next C is a perfect fourth, 4:3.

The major scale's notes trace to the Greeks' consonance ideas. Adding perfect fifths repeatedly from C yields frequencies near the major scale: C, G, D, A, E, B, F-sharp, C-sharp, G-sharp, D-sharp, A-sharp, E-sharp (F), then back to C: the circle of fifths. The overtone series also yields frequencies near the major scale.

A lone note cannot be dissonant but can clash with a chord, especially one implying a key the note is not in. Two notes may clash together or in sequence if they break learned idiom, and chords drawn from outside the established key can clash. The composer must balance all this; when the balance is slightly off, betrayed expectations drive us to change the station, remove earphones or leave.

Gestalt psychology and melody as a whole

Having reviewed pitch, timbre, key, harmony, loudness, rhythm, meter and tempo, Levitin notes that neuroscientists separate them to find where each is processed, but real music succeeds through their relationships; changing a rhythm may require changing pitch, loudness or chords. One approach to relationships goes back to the Gestalt psychologists of the late 1800s.

In 1890 Christian von Ehrenfels puzzled over transposition, singing a song in different pitches. A person starting "Happy Birthday" may pick any note, even one between C and C-sharp, and few notice; three renditions in a week can use entirely different pitches, each a transposition of the others.

The Gestaltists (von Ehrenfels, Max Wertheimer, Wolfgang Köhler, Kurt Koffka) studied how parts form wholes that differ from, and cannot be understood from, their parts. A suspension bridge is a Gestalt: cables and girders do not explain it, and the same parts could make a crane. The Mona Lisa's features scattered would not be the painting. They asked how a melody keeps its identity when every pitch changes: so long as the relations among pitches hold, it is the same melody, across instruments, half or double speed, or all transformations together. They never answered it but formed the school that produced the Gestalt Principles of Grouping for vision.

Albert Bregman (McGill) has spent thirty years on grouping principles for sound. Fred Lerdahl (Columbia) and Ray Jackendoff (Brandeis, now Tufts) described grammar-like rules for musical composition, including grouping. The neural basis is incomplete, but behavioural experiments have clarified the phenomena.

Auditory grouping

In vision, grouping is how elements combine or stay separate, the problem of "what goes with what." It is partly automatic and unconscious. Helmholtz described it as unconscious inference about which things belong together from their attributes. From a mountaintop, trees form a forest group because of similar shape, size and colour; in a mixed forest, smooth white alders pop out from dark craggy pines; focusing on one tree reveals bark, insects and moss; a lawn is not seen as blades unless attended to. Grouping is hierarchical. Some factors are intrinsic (shape, colour, symmetry, contrast, continuity of lines), others psychological (attention, memory, expectation).

Sounds group too. Most people cannot pick out one violin or trumpet; sections form groups, and an entire orchestra can be one stream (Bregman's term). At an outdoor concert with several ensembles, the one in front forms a stream separate from the others, and by attention you can pick out its violins, as you follow one conversation in a crowded room.

One case is the fusion of an instrument's harmonics into one percept: we hear an oboe, not harmonics. With an oboe and trumpet together, the brain analyses dozens of frequencies and builds an image of each and of their combination, the basis for appreciating timbre combinations, which is what Pierce marvelled at in rock's electric bass and guitar: two distinct instruments creating a new sound that can be heard, discussed and remembered.

The system exploits the harmonic series because our brains coevolved with sounds sharing such properties. By Helmholtz's unconscious inference and a likelihood principle, the brain assumes one object produces the harmonic components rather than many sources each making one. Even people who cannot name an oboe can tell two instruments are playing, as people who cannot name notes can tell different notes. Overtones group like blades of grass into "lawn." Different fundamentals yield different overtone sets, so a trumpet and an oboe on different notes are separated by a computer-like process. For the same note, where overtone frequencies nearly coincide (amplitudes differ), the system relies on simultaneous onsets: sounds starting together group together. Since Wilhelm Wundt's first psychology lab in the 1870s, hearing is known to be sensitive to onset differences of a few milliseconds. So a trumpet and oboe on one note are told apart because one spectrum begins perhaps thousandths of a second before the other: grouping both integrates and segregates.

Simultaneous onset is part of a wider temporal positioning principle (sounds now versus tomorrow night). Other cues:

Closing: neural basis and what comes next

The neural subsystems for these attributes separate early, at low brain levels, suggesting grouping uses general mechanisms operating somewhat independently. Yet attributes interact, and experience and attention influence grouping, so part is under conscious control. How conscious and unconscious processes interact is still debated, but progress in the past ten years lets scientists locate areas involved in particular aspects of music and even the part that governs attention. He closes with questions for later chapters: how thoughts form, whether memories are stored in a specific place, why songs get stuck in the head, and whether the brain enjoys tormenting us with commercial jingles.


Chapter 3: Behind the Curtain: Music and the Mind Machine

Mind versus brain

Levitin opens by separating two terms. For cognitive scientists the mind is the part of a person that holds thoughts, hopes, desires, memories, beliefs and experiences. The brain is a bodily organ in the skull, made of cells, water, chemicals and blood vessels, and its activity gives rise to the contents of the mind. A common analogy treats the brain as a computer's hardware and the mind as the software running on it. Different programs can run on much the same hardware, so different minds can arise from very similar brains. He jokes that it would be nice if one could buy a memory upgrade.

Western culture has inherited dualism from Descartes, who held that mind and brain are entirely separate. Dualists say the mind existed before birth and that the brain is only an instrument for carrying out the mind's will, moving muscles and keeping the body stable. Our experience of being "me" makes this feel right, and it is hard to see how that feeling reduces to axons, dendrites and ion channels. Levitin suggests the feeling may be an illusion, much as the earth feels still although it spins at about a thousand miles an hour. Most scientists and contemporary philosophers regard brain and mind as two aspects of one thing, and some think the distinction itself is flawed. The dominant view is that all our thoughts, beliefs and experiences are patterns of electrochemical firing. If the brain stops, the mind is gone, though the brain could still sit in a jar in a laboratory.

Evidence from localised damage

The evidence comes from neuropsychology. Stroke, tumours, head injury and other trauma can damage a specific region and remove a specific function. When dozens or hundreds of cases link the same lost function to the same region, we infer that the region is involved in, or responsible for, that function. More than a century of this work has produced maps of functional areas. The prevailing picture is of the brain as a computational system in which networks of interconnected neurons compute on information and combine the results into thoughts, decisions, perceptions and finally consciousness. Examples given:

Since the case of Phineas Gage in 1848 we have known the frontal lobes relate to self and personality. Yet our knowledge of personality and neural structure remains vague. No "patience," "jealousy" or "generosity" region has been found, and Levitin thinks none will be. Complex personality traits are spread widely through the brain.

The brain has four lobes (frontal, temporal, parietal, occipital) plus the cerebellum. Levitin warns that these are rough generalisations, because behaviour is not reducible to simple mappings.

A lobotomy surgically separates part of the frontal lobe (the prefrontal cortex) from the thalamus. This leads to a joke about the Ramones' song "Teenage Lobotomy," whose line about having no cerebellum is anatomically wrong, though Levitin forgives it for the rhyme.

Music across the brain

Music uses nearly every brain region and subsystem we know of. The brain employs functional segregation, with feature detectors analysing particular aspects of the signal such as pitch, tempo and timbre. Some of this overlaps with other sound analysis. Understanding speech, for example, means segmenting sound into words, sentences and phrases and grasping things beyond the words, such as sarcasm. The several dimensions of a musical sound are handled by quasi-independent processes and then combined into one coherent representation.

Listening begins in subcortical structures (the cochlear nuclei, brain stem, cerebellum) and then reaches the auditory cortices on both sides. Following music you know, or a familiar style such as baroque or blues, adds the hippocampus (memory) and the inferior frontal cortex, in the lowest part of the frontal lobe. Tapping along, aloud or mentally, uses the cerebellum's timing circuits. Performing, whether playing, singing or conducting, uses the frontal lobes for planning, the motor cortex for movement and the sensory cortex for touch feedback that the right key was pressed or the baton moved as intended. Reading music uses the visual cortex in the occipital lobe. Lyrics engage language centres, including Broca's and Wernicke's areas and others in the temporal and frontal lobes. The emotions music evokes involve deeper, primitive structures: the cerebellar vermis and the amygdala.

Alongside regional specificity there is distribution of function. The brain is a massively parallel device. There is no single language centre and no single music centre. Some regions perform component operations and others coordinate them. The brain is also far more capable of reorganising itself than once believed. This neuroplasticity means that regional specificity may be temporary, since processing for important functions can move elsewhere after trauma or damage.

The scale of the brain and of its connections

The brain's numbers are hard to grasp. An average brain has about a hundred billion neurons. If each were a dollar handed out at one per second, around the clock, starting on the day Jesus was born, you would by now have given away only about two thirds. Handing out hundred-dollar bills each second would take thirty-two years. The real power lies in the connections. Each neuron links to roughly a thousand to ten thousand others. The number of ways neurons can be connected grows exponentially. Four neurons give 64 possibilities, and the author lists the series: 2 neurons give 2, 3 give 8, 4 give 64, 5 give 1,024 and 6 give 32,768. The count of possible brain states exceeds the number of known particles in the universe, so we are unlikely ever to understand all the connections.

Music has a similar property. All songs, past and future, can be built from twelve notes (ignoring octaves). Each note can move to any other note, repeat or go to a rest, and each choice leads to twelve more. Adding the many possible note lengths makes the possibilities multiply very fast.

Parallel processing

Much of the brain's power also comes from parallel rather than serial processing. A serial processor is like an assembly line that handles one item at a time. A computer asked to download a song, report the weather in Boise and save a file does these one by one, only seeming simultaneous because it is fast. Brains work on many things at once. The auditory system need not know a sound's pitch before working out where it comes from, since the two circuits work concurrently. A circuit that finishes early passes its result on to connected regions. If late information changes the interpretation, the brain can "change its mind," and it revises its opinions hundreds of times a second without our knowing.

The telephone-friends analogy

To explain how neurons connect, Levitin asks you to imagine a neutral Sunday morning at home with a network of one-dimensional friends you can phone. Hannah makes you happy. Sam makes you sad, because a mutual friend died and Sam reminds you. Carla makes you calm, recalling times meditating with her in a sunny forest clearing. Edward energises you, and Tammy makes you tense. These are your connections, and using them changes your state. Talking to Hannah and Sam together would net out to neutral. A weight on each connection, meaning how close you feel to the person at that moment, sets how much influence they have. If you are twice as close to Hannah as to Sam, you end up happy, though less so than with Hannah alone.

The friends can also talk to each other and change one another's states. Cheerful Hannah can be dampened by Sad Sam. Edward, if he has just spoken to Tense Tammy, who has just spoken to Jealous Justine, might leave you with a new mix, a tense jealousy with energy to act on. Any friend might call you at any time, and you leave your mark on them in turn. With thousands of friends and telephones ringing all day, the range of possible emotional states is huge.

Thoughts and memories are generally thought to arise from such connections. Not all neurons are active at once, because that would be a cacophony, which is in fact what happens in epilepsy. Particular groups, which can be called networks, become active during particular activities and can in turn activate others.

Three examples: a stubbed toe, a car horn, Rachmaninoff

Stubbing a toe sends signals from receptors to the sensory cortex, setting off a chain that produces pain, withdrawal of the foot and perhaps an involuntary shout.

Hearing a car horn sends electrical signals to the auditory cortex. Some neurons process pitch, so that I can tell the horn from a truck's air horn or the airhorn-in-a-can at a football game. Another group works out where the sound came from. This triggers a visual orienting response, and if needed a jump back, driven by the motor cortex together with the amygdala signalling danger.

Hearing Rachmaninoff's Piano Concerto no. 3 involves the following:

Some overlap persists. Deep double bass vibrations may fire touch-sensitive neurons that fired when I stubbed my toe. If the horn is A440, neurons tuned to that frequency fire for the horn and again for an A440 in Rachmaninoff. But my inner experience differs because the context and recruited networks differ. The concerto's oboes and violins may calm me rather than startle me, using neurons active when I feel safe.

Why do I associate horns with danger? Some sounds are intrinsically soothing or frightening. Despite individual variation, we are born predisposed to interpret sounds in certain ways. Abrupt, short, loud sounds are read by many animals as alerts, as in the alarm calls of birds, rodents and apes, while slow-onset, long, quiet sounds read as calming or neutral. He contrasts a sharp dog bark with a cat purring in a lap. Composers know this and use many shadings of timbre and note length to convey emotion.

Haydn's "Surprise Symphony"

In Haydn's Symphony no. 94 in G major, second movement (andante), soft violins carry the main theme, which is soothing. The short pizzicato accompaniment sends a gently contradictory hint of danger, producing mild suspense. The theme spans just over half an octave, a perfect fifth. Its contour goes up, down, then up again, creating a parallelism that prepares us for another "down." Haydn instead rises slightly with the rhythm unchanged and rests on the fifth, a relatively stable tone. Since this is the highest note so far, we expect the next to be lower, starting the return to the tonic and closing the gap. Instead a loud note an octave higher bursts in, played by brash horns and timpani. He thus violates expectations for direction, contour, timbre and loudness at once. Even a listener with no musical knowledge is surprised by the timbral shift from soft violins to the alert call of horns and drums, and a trained listener is also surprised because stylistic conventions are broken. How the brain does surprise, expectation and analysis in neurons is still largely a mystery, though there are clues.

The author's bias for mind over brain

Levitin admits a preference for studying the mind rather than the brain. Part of it is personal. As a child he would not collect butterflies with his class because all life seemed sacred to him. Brain research has typically meant poking into the brains of live animals, often monkeys and apes, then "sacrificing" them. He spent one miserable semester in a monkey lab dissecting dead monkeys' brains for microscopy, walking past live ones' cages each day, and had nightmares.

Intellectually, he is more fascinated by thoughts than by the neurons behind them. Functionalism, held by many prominent researchers, says similar minds can arise from quite different brains, which are just the wiring and modules that carry out thought. Whether or not that is true, it suggests studying brains alone limits what we can learn about thought. A neurosurgeon once told Daniel Dennett, a prominent functionalist, that he had seen hundreds of live thinking brains but never a thought.

Meeting Michael Posner

Choosing a graduate school and mentor, Levitin admired Michael Posner, who pioneered mental chronometry (measuring how long thoughts take to learn about mental organisation), ways to study the structure of categories and the Posner Cueing Paradigm for attention. Rumour said Posner was abandoning the mind for the brain, which Levitin did not want. As an undergraduate (somewhat older than usual) finishing his B.A. at Stanford, he attended the American Psychological Association meeting in San Francisco, forty miles away. Posner's talk was full of slides of brains at work. Afterwards Levitin chased him through the conference centre, out of breath and nervous. He had read Posner's textbook in his first psychology class at MIT, where he began his degree before transferring, and his first psychology professor there, Susan Carey, had spoken reverently of Posner as one of the smartest and most creative people she knew. Stammering, he asked whether it was true that Posner had moved entirely to the brain, adding that he had applied to the University of Oregon to work with him.

Posner replied that he was a little interested in the brain, but saw cognitive neuroscience as a way to constrain theories in cognitive psychology, helping decide whether a model has a plausible anatomical basis.

What cognitive neuroscience is for

Many neuroscientists come from biology or chemistry and focus on how cells communicate. For the cognitive neuroscientist, brain anatomy and physiology is an intellectual puzzle, like a complicated crossword, but not the goal. The goal is understanding thoughts, memories, emotions and experiences, with the brain simply the box they happen in. Returning to the telephone analogy, mapping all the phone lines would only help a little in predicting your mood tomorrow. It matters more to know each friend's tendencies, who will call and what effect they have. Ignoring connectivity would also be a mistake: a broken line, a missing connection, or a person who can only reach you through another all constrain predictions.

This shapes how Levitin studies the neuroscience of music. He rejects fishing expeditions that try every stimulus to see where it lands in the brain, and he and Posner have often discussed the current rush to produce atheoretical brain cartography. His aim is to understand how regions coordinate, and how neuron firing and neurotransmitters yield thoughts, laughter, joy, sadness and lasting art. Where something happens in the brain matters only if it tells us how and why, and cognitive neuroscience assumes it can.

What makes a good experiment

Of the infinite experiments possible, the worthwhile ones advance understanding of how and why. A good experiment is theoretically motivated and makes clear predictions about which of two or more competing hypotheses will be supported. One likely to support both sides of a dispute is not worth doing, since science advances by eliminating false hypotheses.

A good experiment also generalises to other people, other music and other situations. Much behavioural research uses few subjects and artificial stimuli. His laboratory uses musicians and nonmusicians where possible, and nearly always real recordings of real musicians playing real songs, so as to learn about the music most people hear, not music that exists only in labs. This makes rigorous control harder but not impossible, and requires more planning, but it has paid off. It means they study the brain doing what it normally does, not responding to pitchless rhythms or rhythmless melodies. Splitting music into components, if done badly, risks sound sequences that are unmusical.

Preferring mind to brain does not mean ignoring the brain. But similar thoughts can arise from different architectures. He compares watching the same programme on an RCA, a Zenith, a Mitsubishi or a computer screen, devices whose architectures differ enough that the patent office issued separate patents. His dog Shadow has a very different brain organisation, anatomy and neurochemistry, so the firing patterns when Shadow is hungry or hurts a paw probably differ greatly from his own when hungry or stubbing a toe, yet he believes Shadow has substantially similar mind states.

Illusions and misconceptions about perception

First, many people, including scientists in other fields, intuit that the brain holds an isomorphic representation of the world (from Greek roots for "same" and "form"). The Gestalt psychologists, right about much else, articulated this: looking at a square activates a square-shaped pattern of neurons, and looking at a tree might activate tree-shaped neurons with roots at one end and leaves at the other. Likewise, a song in the head feels as if it plays over neural loudspeakers.

Dennett and V. S. Ramachandran argue against this. If a mental picture is itself a picture, some part of the mind must be looking at it. Dennett describes the intuition of a screen or theatre in the mind, which would require an audience member, and that member's mental image would need another viewer, and so on in an infinite regress. The same applies to hearing. Because we can zoom and rotate mental images, or speed up and slow down a song in our heads, we feel there is a home theatre in the mind, but logic rules it out.

Second, we feel we simply open our eyes and see, or a bird chirps and we instantly hear. Perception builds mental representations so fast and seamlessly it seems effortless. This is an illusion, as perceptions are the end of a long chain of neural events. Our strongest intuitions often mislead, like the flat earth, and so does the belief that our senses give an undistorted view.

Since Aristotle it has been known that senses distort. Levitin's teacher Roger Shepard, a perception psychologist at Stanford, used to say that a properly working perceptual system is supposed to distort the world. John Locke noted that all we know comes through the senses, and we assume the world is as we perceive it, but experiments show otherwise. Visual illusions are the most compelling evidence. The Ponzo illusion makes two equal lines look unequal. Shepard's "Turning the Tables," related to the Ponzo, shows two tabletops that look different but are identical in size and shape, which can be checked by tracing one on paper or cellophane and laying it over the other. It exploits depth perception, and even knowing it is an illusion does not turn the mechanism off, so it keeps surprising us because the brain is supplying misinformation. In the Kanizsa illusion (spelled Kaniza in the text) a white triangle seems to lie on a black-outlined one, but no triangles are drawn, because the perceptual system fills in absent information.

Why perception fills in: Warren's experiment

The best guess is that this was evolutionarily adaptive. Sights and sounds often arrive partially obscured, such as a tiger half hidden by trees, or a lion's roar partly masked by nearby rustling leaves. A system that restores missing information speeds decisions in danger. It is better to run than to work out whether two broken pieces of sound were one roar.

The auditory system has its own completion, demonstrated by cognitive psychologist Richard Warren. He recorded the sentence "The bill was passed by both houses of the legislature," cut a piece out of the tape and replaced it with static of equal length. Nearly everyone heard both a sentence and static, but many could not say where the static was, because the system had filled in the missing speech so that the sentence seemed uninterrupted. Static and speech formed separate perceptual streams because their timbres differ, which Bregman calls streaming by timbre. This is a distortion, but it has adaptive value in a life-or-death situation.

Perception as inference

According to Helmholtz, Richard Gregory, Irvin Rock and Roger Shepard, perception is inference involving probabilities. The brain must determine the most likely arrangement of objects in the world given the pattern reaching the receptors (retina for vision, eardrum for hearing). That input is usually incomplete or ambiguous, with voices mixed with other voices, machines, wind and footsteps. Listening around you now, outside a sensory isolation tank you could likely identify half a dozen sounds, which is remarkable given what the receptors pass up. Grouping principles by timbre, spatial location, loudness and so on help separate sources, but much is unknown and nobody has built a computer that can do sound source separation.

The eardrum problem

The eardrum is a membrane stretched across tissue and bone, and nearly all our impressions of the auditory world come from how it moves in response to air molecules. The pinnae and skull bones contribute somewhat. Levitin pictures a woman reading in her living room with six identifiable sounds: the heating blower, the refrigerator hum, street traffic (itself possibly dozens of sounds), leaves rustling, a purring cat, and a recording of Debussy preludes. Each is an auditory object or sound source with its own sound.

Sound is carried by molecules vibrating at certain frequencies. They push the eardrum in and out depending on force (amplitude, or volume) and speed (pitch). But nothing in the molecules says where they came from or which object they belong to. Cat-purr molecules carry no "cat" tag and may hit the same part of the eardrum at the same time as those from the refrigerator, heater and Debussy.

He offers the pillowcase analogy: stretch a pillowcase over a bucket while people throw Ping-Pong balls at it from various distances, any number at any rate. From the pillowcase's movement alone you must work out how many people there are, who they are, and whether they approach, recede or stand still. The auditory system faces this task, and the question is how the brain does it for music.

Feature extraction, feature integration, bottom-up and top-down

The brain first does feature extraction, then feature integration. Specialised neural networks break the signal into pitch, timbre, spatial location, loudness, reverberant environment, tone durations and note onset times, including those of tone components. These computations run in parallel and fairly independently, so the pitch circuit does not wait for the duration circuit. Processing driven only by information in the stimulus is bottom-up. In both the world and the brain these attributes are separable, as one can change a visual object's shape without changing its colour.

Low-level processing deals with elemental attributes and happens in peripheral and phylogenetically older brain parts. High-level processing happens in more sophisticated regions that receive projections from receptors and many low-level units, and combines elements into an integrated representation, where understanding of form and content arises. Low-level vision sees ink blobs and perhaps recognises the letter A, while high-level processing joins three letters into ART and brings its meaning to mind.

While extraction goes on in the cochlea, auditory cortex, brain stem and cerebellum, higher centres, mostly frontal cortex, receive a constantly updated stream about what has been extracted, which typically overwrites older information. They work to predict what comes next in the music, based on four things:

These frontal-lobe calculations are top-down processing and can influence the low-level modules during their bottom-up work. Top-down expectation can cause misperception by resetting circuitry in the bottom-up processors, which is partly the neural basis of perceptual completion and other illusions.

The two directions inform each other continuously. Higher, more phylogenetically advanced regions integrate features into a perceptual whole, building a representation of reality as a child builds a fort from Lego. In doing so the brain draws inferences from incomplete or ambiguous data, and when they are wrong we get visual and auditory illusions, which show the system guessed wrongly about what is out there.

Three difficulties and the unconscious nature of inference

In identifying auditory objects the brain faces three problems. The incoming information is undifferentiated, it is ambiguous (different objects can produce similar or identical eardrum patterns), and it is seldom complete because sounds may be masked or lost. So the brain makes a calculated guess, quickly and mostly subconsciously. Illusions and these operations lie outside awareness. Knowing that perceptual completion explains the Kanizsa triangles does not switch it off, and you stay surprised.

Helmholtz called this "unconscious inference" and Rock called it "the logic of perception." George Miller, Ulrich Neisser, Herbert Simon and Roger Shepard describe perception as a "constructive process." All mean that what we see and hear is the end of a long chain of mental events producing an impression of the physical world. Many brain functions, including colour, taste, smell and hearing, arose under evolutionary pressures, some now gone. Steven Pinker and others suggest music perception was an evolutionary accident, with survival and sexual-selection pressures creating a language and communication system that we learned to exploit for music. This is contentious, and the archaeological record rarely gives a decisive "smoking gun."

Illusions in music

Filling in is not just a lab curiosity, and composers exploit it, knowing a melodic line seems to continue even when instruments partly hide it. When we hear the lowest piano or double bass notes, we are not really hearing 27.5 or 35 Hz, since those instruments produce little energy that low, and our ears fill in the tone. Further examples:

Production illusions and hyperreality

Most modern recordings contain another sort of illusion. Artificial reverberation makes vocalists and lead guitars seem to come from the back of a concert hall even through headphones an inch from our ears. Microphone techniques can make a guitar seem ten feet wide with our ears at the soundhole, impossible in reality since the strings cross the soundhole and the guitarist would be strumming our nose. The brain uses cues about the spectrum and echoes to understand the auditory world, as a mouse uses whiskers to understand its physical one, and engineers mimic these cues to give studio recordings a lifelike quality.

This also partly explains the appeal of recorded music, especially now with personal players and headphones. Engineers and musicians create effects that exploit circuits evolved to detect important features of our environment. Like 3-D art, films or visual illusions, which are too recent for dedicated brain mechanisms, they borrow systems built for other purposes. Using those circuits in novel ways makes them interesting, and modern recordings work the same way.

Our brains can judge an enclosed space's size from reverberation and echo. Without knowing the equations, anyone can tell a small tiled bathroom from a mid-sized concert hall or a high-ceilinged church, and can tell from a recording what size room a singer or speaker was in. Engineers create what Levitin calls "hyperrealities," the recorded counterpart of a cinematographer mounting a camera on a speeding car's bumper, giving impressions we never get in real life.

Our brains are also extremely sensitive to timing, localising sounds from arrival differences of a few milliseconds between the ears. Many recorded effects exploit this. The guitar sounds of Pat Metheny and David Gilmour of Pink Floyd use multiple delays to create an otherworldly, haunting effect, simulating a cave with many echoes that never occurs in nature, like the infinitely repeating reflections of barbershop mirrors.

The illusion of structure and musical grammar

Perhaps the ultimate musical illusion is of structure and form. Nothing in a note sequence, a scale, a chord or a chord sequence intrinsically produces emotional associations or an expectation of resolution. Making sense of music depends on experience and on neural structures that learn and modify themselves with each new song and each replay of an old one. Our brains learn a musical grammar specific to our culture, as we learn our culture's language.

Noam Chomsky proposed that we are born with an innate capacity to understand any human language, and that experience with a particular language shapes, builds and finally prunes a complex, interconnected network of circuits. The brain does not know beforehand which language it will meet, but brains and languages coevolved so that all languages share basic principles, and the brain can absorb any of them almost effortlessly through exposure in a critical developmental stage.

Similarly we seem to have an innate capacity to learn any of the world's musics, though they differ substantively. After birth the brain develops rapidly for the first years, forming new connections faster than at any other time. In mid-childhood it prunes them, keeping the most important and most used. This becomes the basis of our understanding of music and of what we like, what moves us, and how. Adults can still learn to appreciate new music, but basic structural elements get built into the brain's wiring by early listening.

Music as perceptual illusion, and the lead-in to expectation

Music can therefore be seen as a perceptual illusion in which the brain imposes structure and order on a sequence of sounds. Why this produces emotion remains part of the mystery, since we do not weep over other kinds of order, such as a balanced chequebook or neatly arranged first-aid products in a drugstore. What is it about musical order that moves us? Scale and chord structure play a part, as does brain structure. Feature detectors extract information from the incoming sound, and the brain's computational system assembles it into a whole, partly from what it thinks it should be hearing and partly from expectations. Where expectations come from is key to understanding how music moves us, when it does, and why some music just makes us reach for the off button. Musical expectation is, Levitin says, perhaps the area of music cognitive neuroscience where music theory and neural theory, and musicians and scientists, come together most harmoniously. To understand it fully we need to study how particular musical patterns produce particular patterns of neural activation, which leads into the next chapter.


Chapter 4: Anticipation (What We Expect from Liszt and Ludacris)

Opening: Music, Emotion and Expectation

Levitin begins with a personal note: at weddings it is the music, not the sight of the couple, that makes him cry, and in films the music tips him over when lovers are reunited. He returns to his earlier definition of music as organized sound and adds that the organization must contain some surprise, or the result is flat and robotic. Appreciating music depends on learning its underlying structure, which is comparable to grammar in spoken or signed language, and on predicting what comes next. Composers and performers produce thrills, chills and tears by knowing what listeners expect and deliberately controlling when those expectations are met and when they are not.

Ways Composers Violate Expectations

The best-documented trick in Western classical music is the deceptive cadence. A cadence is a chord sequence that sets up an expectation and normally closes with a satisfying resolution. In the deceptive version the composer repeats the sequence until listeners are convinced the resolution is coming, then supplies a chord that stays within the key but does not fully resolve. Haydn used it almost obsessively. Perry Cook compares this to a magician's trick, in which expectations are set up and defied without the audience knowing how or when. The Beatles' "For No One" ends on the V chord, leaving the resolution hanging, and the next song on Revolver opens with the chord we were waiting for.

Levitin then lists many other examples:

The Brain Constructs Musical Reality

These features are not directly represented in the brain, at least in early processing. The brain builds its own version of reality, based partly on what is present and partly on the role each tone plays in a learned musical system. Language works the same way: nothing about the word "cat" is catlike, and we have simply learned the association. Likewise we learn which tones go together, and our brains perform a statistical analysis of how often pitches, rhythms and timbres have co-occurred, producing expectations. Levitin rejects the appealing idea that the brain stores an accurate, one-to-one copy of the world; it stores distortions and illusions and extracts relationships. As evidence, light varies in one dimension, wavelength, yet we perceive colour in two dimensions (the colour circle from earlier). Similarly, pitch arises from one continuum of vibration speeds, but the brain builds a pitch space of three, four or even five dimensions in some models. Adding so many dimensions may help explain our deep reactions to well-constructed sound.

Schemas

For cognitive scientists, a violated expectation is an event at odds with what could reasonably have been predicted. We know a lot about standard situations that differ only in insignificant details. Reading is the example: feature extractors in the brain learn the unvarying essentials of letters and ignore details such as font unless we attend to them. Seeing every word in a different font is jarring, but the point stands that the detectors extract "the letter a" rather than the typeface.

The brain extracts what several situations share and builds a framework called a schema. A schema for the letter a would include its shape and memory traces of every a we have seen, with their variability. Schemas shape daily life: the birthday-party schema differs by culture and age and carries expectations, some flexible and some not. Levitin lists typical elements: a person celebrating the anniversary of their birth, other people celebrating with them, a cake with candles, presents, festive food, and hats, noisemakers and decorations. A missing item would not shock us, but the more are missing the less typical the party. For an eight-year-old we might expect pin-the-tail-on-the-donkey, but not single-malt scotch.

Musical schemas begin forming in the womb and are refined with every listening experience. Our Western schema includes implicit knowledge of the scales normally used, which is why Indian or Pakistani music sounds strange to us at first but not to people from those countries or to infants. It sounds strange only because it is inconsistent with what we have learned to call music. By age five, children recognize chord progressions of their culture.

We also form schemas for genres and styles, and style is just another word for repetition. A Lawrence Welk concert schema has accordions and no distorted guitars, and a Metallica schema is the reverse. A Dixieland schema is foot-tapping and up-tempo, with no overlap with a funeral procession's repertoire unless the band is being ironic. Schemas extend memory: we recognize that we have heard something before, and whether in the same piece or another. According to Eugene Narmour, listening requires holding in memory the notes just heard along with knowledge of all the other music we know that resembles the current style. The latter is less vivid but provides necessary context.

The main schemas cover genres, styles, eras, rhythms, chord progressions, phrase structure, song length and which notes tend to follow which. The four- or eight-measure phrase is part of our schema for late twentieth-century popular songs, absorbed from thousands of songs without being able to state the rule. So "Yesterday" still surprises after thousands of hearings, because the violated schema is more entrenched than our memory of the particular song. Songs that last for years play with expectations enough to stay slightly surprising, which is why some people never tire of Steely Dan, the Beatles, Rachmaninoff or Miles Davis.

Melody, Gap Fill and Beethoven

Melody is a main tool for controlling expectations. Theorists identify gap fill: after a large leap up or down, the next note should change direction. Typical melodies move mostly by step, and after a leap the melody tends to "want" to return toward its starting point or harmonic home.

Neural Codes

Because the brain does not hold an isomorphic copy of the world, Levitin asks what its neurons do hold: mental or neural codes. Neuroscientists decipher these at the level of neurons, and cognitive psychologists at the level of general principles.

He compares this to a picture on a computer. A grandmother's photograph is a physical object, but a computer image is a file of 0s and 1s, the same code used for everything. A corrupt file or failed attachment shows gibberish, a sort of intermediate hexadecimal code. In a simple black-and-white image a 1 might mean a black dot and a 0 a white one, but the digits form a long line, not a triangle, and the computer has instructions for where each belongs. Someone good at reading such files could guess the nature of the image; people who work with image files can tell about redness, grayness and sharpness of edges, though not whether it shows a human or a horse. Audio files are likewise binary, indicating sound in parts of the frequency spectrum, so a sequence indicates a bass drum or a piccolo.

Computers decompose objects into small components (pixels, or sine waves with frequency and amplitude) and translate them to code, while software hides this: we double-click and see or hear the original. This effortless illusion mirrors the neural code of millions of nerves firing at different rates, all invisible to us; we cannot feel them or speed them up, slow them down, start them in the morning or stop them at night.

Levitin and Perry Cook once read about a man who could identify the music on a record from its grooves with the label hidden. They examined old records and saw regularities: low notes make wide grooves, high notes narrow ones, and the needle moves thousands of times a second to trace the wall. Someone who knew many pieces could characterize them by how many low notes there are (lots in rap, few in baroque concertos) and whether they are steady or percussive (walking bass in jazz swing versus slapping bass in funk). The skill is extraordinary but explicable.

We meet auditory code readers daily: the mechanic who diagnoses clogged fuel injectors or a slipped timing chain from engine sound, the doctor who hears an arrhythmia, the detective who hears lying in vocal stress, and the musician who tells a viola from a violin or a B-flat from an E-flat clarinet. In all, timbre helps unlock the code.

Neurons and Neurotransmitters

To study neural codes, some neuroscientists examine what makes neurons fire, how fast, and their refractory period (recovery time), plus how neurons communicate and the role of neurotransmitters. Little is known about the neurochemistry of music, though Levitin promises new results from his laboratory in Chapter 5.

Neurons are the brain's primary cells, also found in the spinal cord and peripheral nervous system. Outside activity can fire them, as when a tone excites the basilar membrane, which signals frequency-selective neurons in the auditory cortex. Contrary to a century-old view, neurons do not touch; the gap is the synapse. Firing sends an electrical signal that releases a neurotransmitter, a chemical that crosses the synapse and binds to receptors on a neighbouring neuron, as keys fit locks, and only certain neurotransmitters fit certain receptors. Neurotransmitters generally make the receiving neuron fire or prevent it, and are then absorbed by reuptake; without it they would keep stimulating or inhibiting.

Some neurotransmitters work throughout the nervous system and others only in certain regions or neuron types. Serotonin, produced in the brain stem, is linked to mood and sleep regulation. Antidepressants such as Prozac and Zoloft are selective serotonin reuptake inhibitors (SSRIs), letting existing serotonin act longer; how this eases depression, obsessive-compulsive disorder and mood and sleep disorders is unknown. Dopamine, released by the nucleus accumbens, helps regulate mood and movement and is famous for the pleasure and reward system: it is released when addicts get their drug, gamblers win, or chocoholics get cocoa. Its role in music, and that of the nucleus accumbens, was unknown until 2005.

Hemispheric Specialization

Cognitive neuroscience has advanced greatly in the last decade. A popular macro-level notion is that the left and right hemispheres perform different functions. That is true but more nuanced. The research was done on right-handed people. Left-handers (about 5 to 10 percent) and ambidextrous people sometimes share right-handers' organization but more often differ, either as a simple mirror image or in poorly documented ways. So generalizations apply only to right-handers.

Writers, businessmen and engineers call themselves left-brain dominant, and artists, dancers and musicians right-brain. The idea that left is analytical and right artistic has some merit but is too simple: both sides analyse and think abstractly, and these activities need both hemispheres, though some functions are lateralized.

Speech processing is mainly left-hemisphere, but global aspects like intonation, emphasis and pitch pattern, collectively prosody, are more often disrupted by right-hemisphere damage; telling a question from a statement or sarcasm from sincerity relies on these cues. One might expect music to be the mirror image, and there are cases of left-hemisphere damage that remove speech but spare music, and vice versa, suggesting shared circuits but not complete overlap.

Local features of speech, such as telling speech sounds apart, are left-lateralized. For music, overall melodic contour (shape ignoring intervals) and fine discrimination of close pitches are processed on the right. The left is involved in naming songs, performers, instruments and intervals. Musicians using the right hand or reading with the right eye engage the left brain, which controls the right body. New evidence says following the development of a musical theme, thinking about key, scale and whether the music makes sense, is lateralized to the left frontal lobes. Musical training shifts some processing from the right (imagistic) hemisphere to the left (logical), as musicians use linguistic terms. Development also increases specialization: children show less lateralization of musical operations than adults, musicians or not.

Studying Musical Expectation with EEG

Expectation is best examined by how we track chord sequences, since music unfolds over time and its tones lead the brain to predict what follows. Neural firing makes a small electric current, measured by the electroencephalogram (EEG), with electrodes painlessly placed on the scalp. EEG is very sensitive to timing, down to a millisecond. Limits: it cannot tell whether activity releases excitatory, inhibitory or modulatory neurotransmitters such as serotonin and dopamine, and since a single neuron's signal is weak, it only detects synchronous firing of large groups.

Its spatial resolution is also poor, because of the inverse Poisson problem. Levitin uses an analogy: someone inside a stadium shines a flashlight at a spot on a semitransparent dome, while an observer outside must guess where the person stands; any spot on the field gives the same view, and mirrors would make it worse. Brain signals can come from many sources, on the surface or deep in the sulci, and can bounce off them before reaching the scalp. Still, EEG has the best temporal resolution of common tools, which suits time-based music.

Stefan Koelsch, Angela Friederici and colleagues played chord sequences that either resolved in the standard way or ended on unexpected chords. Brain activity linked to musical structure appeared 150 to 400 ms after the chord began, and activity linked to musical meaning about 100 to 150 ms later. Structural processing (musical syntax) was localized to the frontal lobes of both hemispheres, adjacent to and overlapping with speech syntax areas such as Broca's area, and appeared whether or not listeners were trained. Musical semantics, linking tone sequences to meaning, seems to lie in the rear temporal lobes on both sides, near Wernicke's area.

Music and Language: Shared and Separate

Many case studies show patients losing one faculty but not the other after injury, indicating functional independence. The most famous is Clive Wearing, a musician and conductor with herpes encephalitis damage who, as Oliver Sacks reported, lost all memory except musical memories and memory of his wife. Others lost music but kept language and memory. When parts of his left cortex deteriorated, Ravel lost his sense of pitch but kept timbre, a loss said to have inspired Bolero, which stresses timbre variation. The simplest explanation is that music and language share some neural resources but also have independent pathways. Their closeness and partial overlap in frontal and temporal lobes suggest the circuits start undifferentiated and experience and development separate them.

Babies are thought to be synesthetic, unable to separate the senses, so that five might be red, cheddar might taste like D-flat, and roses might smell like triangles. Maturation creates distinctions by pruning connections, so a cluster that responded to every sense becomes a specialized network. Music and speech may likewise share origins, regions and networks, then develop dedicated pathways with experience, perhaps still sharing resources, as Ani Patel proposes in his shared syntactic integration resource hypothesis (SSIRH).

Levitin and Menon's fMRI Study

Levitin and his friend Vinod Menon, a Stanford systems neuroscientist, wanted to pin down the Koelsch and Friederici findings and support Patel's hypothesis. EEG's spatial resolution was too coarse, so they used another method. Haemoglobin is slightly magnetic, so blood flow changes can be tracked by magnetic resonance imaging (MRI), a giant electromagnet. (MRI development was done by the British company EMI, largely funded by Beatles profits, so "I Want to Hold Your Hand" might have been "I Want to Scan Your Brain.") Active neurons need more oxygen, so regions with most blood flow are the ones engaged; this use is functional MRI (fMRI).

fMRI shows a living brain thinking. Imagining a tennis serve shifts blood to the arm region of the motor cortex, and doing a math problem sends it to frontal regions for arithmetic. Will brain imaging allow mind reading? Probably not, and certainly not soon, since thoughts involve too many regions. fMRI can tell music listening from watching a silent film, but not hip-hop from Gregorian chant, let alone a song or thought.

fMRI locates activity within a couple of millimetres, but its temporal resolution is poor because of hemodynamic lag. Others had studied when structure is processed; they wanted where, and whether it involved speech areas. Attending to musical structure activated the left pars orbitalis (part of Brodmann Area 47) in the frontal cortex, overlapping partly with language-structure studies but with unique activations too, plus an analogous right-hemisphere area. So musical structure needs both hemispheres, whereas language structure needs only the left.

Most astonishingly, the left-hemisphere regions tracking musical structure are those active when deaf people use sign language. So the region does not just judge whether chords or sentences make sense; it responds to the visual organization of American Sign Language words. It appears to process structure in general conveyed over time, with inputs from different populations and outputs through distinct networks, and it kept appearing in tasks involving organizing information over time.

The Emerging Picture

All sound starts at the eardrum and is quickly segregated by pitch. Soon after, speech and music probably diverge. Speech circuits break the signal into phonemes. Music circuits separately analyze pitch, timbre, contour and rhythm, and their outputs go to frontal regions that assemble everything and look for structure in the temporal pattern. The frontal lobes then consult the hippocampus and interior temporal regions: have I heard this before, when, what does it mean, is it part of a larger unfolding sequence? With the neurobiology of structure and expectation settled, Levitin says the next step is the brain mechanisms of emotion and memory.


Chapter 5: You Know My Name, Look Up the Number: How We Categorize Music

Opening memory and the questions of the chapter

Levitin begins with one of his earliest memories: at about three he lay on a shaggy green wool carpet under the family grand piano while his mother, who was German, played. He could see only her legs working the pedals, but the sound surrounded him and vibrated through the floor and his body, low notes on his right and high notes on his left. He recalls the dense chords of Beethoven, the acrobatic notes of Chopin and the almost militaristic rhythms of Schumann. The music held him in a trance and time seemed to stop.

From this he asks how memories of music differ from other memories, why music can bring back memories that seemed lost, how expectation produces emotion in music, and how we recognize songs we have heard before.

Tune recognition and the lookup table

Recognizing a tune means the brain must ignore features that change between hearings and keep the ones that stay the same. Volume, instrumentation, tempo and pitch can all vary without changing the song's identity, so they must be set aside while the essentials are abstracted. Otherwise a song played at a different volume would sound like a different song. Separating invariant from momentary properties is a huge computational problem.

Levitin worked in the late 1990s for an Internet company whose software identified MP3 files, many of which were misnamed or unnamed (a misspelled "Elton John," or Elvis Costello's "Alison" filed under its chorus line "My Aim Is True"). Naming them was fairly easy because every recording has a digital fingerprint, and the task was to search a database of about half a million songs efficiently. Computer scientists call this a lookup table. It is like finding a Social Security number from a name and birth date: one performance maps to one sequence of digital values. The program could not, however, find other versions of the same song. He might have eight versions of "Mr. Sandman," but given Chet Atkins's it could not locate Jim Campilongo's or the Chordettes'. The stream of numbers does not translate readily into melody, rhythm or loudness. A program would need to detect constant melodic and rhythmic intervals while ignoring performance details. The brain does this easily, and no computer can yet begin to.

Constructivist versus record-keeping memory

This human/computer difference ties to a century-old debate over whether memory is relational or absolute. The relational, or constructivist, school holds that memory stores relations among objects and ideas rather than sensory details, so we rebuild reality from those relations and fill in details on the spot; memory's job is to discard irrelevant details and keep the gist. The rival record-keeping theory likens memory to a tape recorder or digital camera that preserves experience with near-perfect fidelity.

Music bears on the debate because, as the Gestalt psychologists noted over a hundred years ago, melodies are defined by pitch relations (constructivist) yet are made of precise pitches (record-keeping, if those pitches are stored).

Evidence for the constructivists. People who hear or read something and then report it recall the general content but not the exact wording. Elizabeth Loftus of the University of Washington, interested in eyewitness testimony, showed subjects videotapes of cars scraping each other and asked how fast the cars were going when they "scraped" or "smashed" into each other. The one-word change produced very different speed estimates. When subjects returned up to a week later and were asked how much broken glass they saw (there was none), those who had heard "smashed" were more likely to report remembering glass. Their memory had been rebuilt from the earlier question. Levitin concludes that memory is built from disparate, possibly inaccurate pieces, and that retrieval resembles perceptual filling-in. He compares this to retelling a dream at breakfast: the memory comes in fragments, and we fill gaps while telling it (his example has a ladder, a Sibelius concert and Pez candy raining from the sky).

Such story-making is the job of the left brain, probably the orbitofrontal cortex behind the left temple, which makes up stories from limited information and will go far to sound coherent. Michael Gazzaniga found this in commissurotomized patients, whose hemispheres were surgically separated to treat severe epilepsy. Because inputs and outputs are largely contralateral, a chicken talon could be shown only to the left brain and a snow-covered house only to the right. The patient pointed to a chicken with his right hand (left brain) and a shovel with his left hand (right brain). Asked why he picked the shovel, his left hemisphere, which had seen both images, invented that you need a shovel to clean the chicken shed, unaware of the house or of inventing the story.

In the early 1960s at MIT, Benjamin White, following the Gestalt psychologists, altered well-known songs such as "Deck the Halls" and "Michael, Row Your Boat Ashore." He transposed pitches, kept contour while shrinking or stretching the intervals, played tunes backward and forward, and changed rhythms. Almost always the altered tune was recognized more often than chance. Listeners recognized transposition almost at once without error. The constructivist reading is that memory extracts and stores generalized invariant information; under record-keeping, each transposition would require fresh comparison against a single stored performance.

Evidence for record-keeping. The Gestalt idea was that each experience leaves a trace that is reactivated at retrieval. Roger Shepard showed people hundreds of photographs for a few seconds each, then a week later showed pairs of old and new ones, where new ones differed subtly (the angle of a sailboat's sail, the size of a background tree), and subjects were astonishingly accurate. Douglas Hintzman showed letters differing in font and capitalization, and subjects remembered the specific font. People also recognize hundreds or thousands of voices: your mother's within one word, your spouse's, and whether he or she has a cold or is angry, all by timbre. Famous voices (Woody Allen, Nixon, Drew Barrymore, W. C. Fields, Groucho Marx, Katharine Hepburn, Clint Eastwood, Steve Martin) are remembered with their catchphrases, which favors storing specifics.

Yet impressionists are funniest saying things the celebrity never said, which implies a stored trace of timbre independent of words. That might undercut record-keeping, but Levitin notes that timbre may be a separable attribute stored as specific values, which would also explain recognizing a clarinet playing an unfamiliar tune.

The neuropsychological case of S., a Russian patient of A. R. Luria, shows the extreme. He had hypermnesia (remembering everything) and could not see that different views and expressions belonged to one person; a smile was one face, a frown another. "Everyone has so many faces!" he complained. Only his record-keeping system worked; he could not abstract. Yet understanding speech requires setting aside differences in how people and contexts pronounce words.

Categories: from Aristotle to Rosch

Scientists dislike two theories that make different predictions, but Levitin's answer to which is right is that neither is. The resolution came from work on categories and concepts, which was developing at the same time. Categorization is basic to living things: every object is unique, but we treat objects as members of classes.

Aristotle held that categories come from lists of defining features. The category "triangle" holds images of every triangle we have seen plus imagined ones, with its boundary set by a definition (three-sided; or, more mathematically, closed, interior angles summing to 180 degrees), with subcategories such as isosceles, equilateral and right triangles. New items are assigned by comparing their properties to the definition. From Aristotle through Locke to modern times, membership was thought to be a matter of logic: in or out.

After 2,300 years Ludwig Wittgenstein asked what a game is, launching empirical work on categories. Eleanor Rosch wrote her Reed College philosophy thesis on Wittgenstein, felt a year with him "cured her" of philosophy, and wanted to study philosophical questions empirically; she took her Ph.D. in cognitive psychology at Harvard and is a professor at Berkeley, where Levitin taught. Many cognitive psychologists now describe the field as "empirical philosophy."

Wittgenstein argued no definition covers all games. Proposed features are: done for fun, a leisure activity, found among children, has rules, competitive, involves two or more people. Each fails: Olympic athletes may not be having fun, pro football is not leisure, poker and jai alai are not childish, a child throwing a ball at a wall has no rules, ring-around-the-rosy is not competitive, solitaire is solitary. His alternative was family resemblance: something is a game if it resembles things already called games, as at a family reunion where cousins share Aunt Tessie's eyes, the family chin, Grandpa's forehead or Grandma's red hair but no single required feature. The feature list can be dynamic (red hair dies out, then returns). Levitin says this foreshadows multiple-trace memory models, worked on by Hintzman and recently by Stephen Goldinger of Arizona.

Genres as family resemblance

Levitin applies this to music. Defining heavy metal as having distorted guitars, heavy loud drums, three or power chords, shirtless sweaty singers and umlauts in band names is easy to refute. Michael Jackson's "Beat It" has distorted guitar (a solo by Eddie Van Halen), and so does a Carpenters song. Led Zeppelin has songs with none ("Bron-y-aur," "Down by the Seaside," "Goin' to California," "The Battle of Evermore"), and "Stairway to Heaven" lacks heavy drums and distortion for 90 percent of its length and has more than three chords. Raffi's songs use three or power chords. Metallica's singer is not called sexy, and umlauts appear in Motley Crue, Blue Oyster Cult, Motorhead, Spinal Tap and Queensryche but not Led Zeppelin, Metallica, Black Sabbath, Def Leppard, Ozzy Osbourne or Triumph. Something is heavy metal if it resembles heavy metal.

Rosch's three insights

First, using Wittgenstein, Rosch said membership comes in degrees. A robin is clearly a bird; chickens and penguins qualify after a pause but are poor examples, shown by hedges such as "technically a bird." Boundaries are fuzzy and open to disagreement: is white a color, is hip-hop music, is Queen without Freddie Mercury still Queen (worth $150 a ticket?), is a cucumber a fruit or vegetable, is someone my friend. A person can even disagree with himself over time.

Second, earlier category experiments used artificial concepts and stimuli with little relation to the real world, and were unintentionally biased toward the experimenters' theories. This reflects a general tension in science between experimental control and real-world validity. Levitin cites Alan Watts: to study a river you do not stare at a bucket of its water, because you lose its motion and flow. He says much music neuroscience of the past decade likewise uses artificial melodies and sounds so removed from music that it is unclear what is learned.

Third, some stimuli hold a privileged place in perception or concept and become prototypes around which categories form. Red and blue arise from retinal physiology, with certain reds more vivid because a particular wavelength maximally fires red receptors. Rosch tested the Dani of New Guinea, whose language has just two color words, mili and mola, roughly light and dark. With no word for red, they had no training in what counts as good red. Shown chips of many reds, they overwhelmingly chose the same best red as Americans and remembered it better, and did the same for unnamed greens and blues.

Her conclusions were: (a) categories form around prototypes; (b) prototypes can have a biological foundation; (c) membership is a matter of degree, with better and worse exemplars; (d) new items are judged against prototypes, giving gradients of membership; and (e) members need not share any common attribute, and boundaries need not be definite.

Levitin's lab informally found the same for musical genres: people agree on prototypical songs for "country," "skate punk" and "baroque," and judge some as lesser examples (the Carpenters are not really rock, Sinatra is less jazz than Coltrane). Prototypes appear even within one artist. Choosing "Revolution 9" as a Beatles song would draw the complaint "that's not what I meant." Neil Young's doo-wop album Everybody's Rockin' and Joni Mitchell's jazz work with Charles Mingus are atypical, and their labels threatened each with contract cancellation for off-brand music.

Shepard's three appearance-reality problems

Understanding begins with particular cases that the brain handles as category members. Roger Shepard frames this evolutionarily: to find food, water and shelter, escape predators and mate, animals must solve three problems.

  1. Objects that present similarly may be different (the apple on the tree versus the one in hand; several violins playing one note are several instruments).
  2. Objects that present differently may be identical (an apple seen from above or the side; your voice heard in person with both ears versus on the phone in one ear), requiring integration into a single representation.
  3. Objects that present differently may be of the same natural kind. This is categorization, the most powerful principle, which higher mammals, many lower mammals, birds and even fish manage: red and green apples are both apples, and mother and father both caregivers.

Adaptive behavior requires separating invariant properties from momentary circumstances. Leonard Meyer notes that classification lets composers, performers and listeners internalize stylistic norms and notice deviations. Levitin quotes Shakespeare on giving "airy nothing a local habitation and a name."

Posner and Keele's dot patterns

Rosch's work prompted challenges. Posner and Keele made prototype patterns of dots placed in a square, like dice faces with random dots, then shifted dots a millimeter or so randomly to make distortions, some so large they were hard to tie to a prototype. Levitin compares this to jazz artists varying a standard: Sinatra's "A Foggy Day" versus Ella Fitzgerald and Louis Armstrong's, baroque and Enlightenment musicians such as Bach and Haydn performing variations, and Aretha Franklin's "Respect" versus Otis Redding's, all considered the same song. He asks whether such versions share a family resemblance or vary an ideal prototype.

Subjects saw many distorted patterns, never the prototypes and without being told they existed. A week later they judged which patterns were old, doing well. Prototypes had been slipped in, and subjects often said they had seen the two unseen prototypes before. That suggests prototypes are stored, so memory must process stimuli beyond preserving them, which seemed a blow to record-keeping.

Work by White and later Jay Dowling (University of Texas) shows music is robust to transformation: transposition, tempo, instrumentation, intervals, scales, even major to minor, and arrangement. Levitin owns the Austin Lounge Lizards playing Pink Floyd's "Dark Side of the Moon" on banjos and mandolins and the London Symphony Orchestra playing the Rolling Stones and Yes. Memory seems to extract a formula allowing recognition, so the constructivist view fits music and, via Posner and Keele, vision.

Absolute pitch and Levitin's Stanford experiments

In 1990 Levitin took a Stanford course on psychoacoustics and cognitive psychology for musicians, team-taught by John Chowning, Max Mathews, John Pierce, Roger Shepard and Perry Cook. Cook suggested a project on memory for pitches and whether people can attach arbitrary labels to them, uniting memory and categorization. Theory predicted no reason to retain absolute pitch, given easy recognition in transposition, and only about one in ten thousand people can name notes: those with absolute pitch (AP).

AP possessors name notes as effortlessly as most of us name colors, including a C-sharp played on piano and pitches of car horns, fluorescent hum or knives on plates. Color and pitch are both psychophysical fictions, structure imposed by the brain on a single continuum (light or sound frequency). Most of us do identify sounds effortlessly, but by timbre: car horn, trumpet, grandmother Sadie with a cold. Why some have AP remains unsolved; the late Dixon Ward quipped the real question is why we don't all.

Levitin found roughly a hundred AP papers between 1860 and 1990 and another hundred in the following fifteen years. All tests required note names, so nonmusicians could not be tested. He and Cook gave nonmusicians tuning forks, to bang on their knees and listen to several times a day for a week, calling the pitch Fred or Ethel (after the neighbors on I Love Lucy, surname Mertz, a rhyme with Hertz noticed only years later). Half the forks were middle C and half G. After a week without the forks, half of subjects sang their pitch and half picked it from three keyboard notes, and they overwhelmingly succeeded, suggesting ordinary people can remember notes tied to arbitrary names.

Shepard then asked whether nonmusicians remember song pitches without names. Levitin cited Andrea Halpern, who had nonmusicians sing "Happy Birthday" or "Frere Jacques" on two occasions: they differed from each other in key but were consistent individually, implying long-term encoding of pitch. Skeptics proposed muscle memory of the vocal cords (Levitin counts that as memory anyway). But Ward and Ed Burns had AP singers sight-sing (only AP singers can sing the right key from a score alone), then used headphones with loud noise so that they relied on muscle memory; they were off by about a third of an octave on average.

Halpern's songs have no correct key, no canonical version or standard recording. Rock and pop songs (Rolling Stones, Police, Eagles, Billy Joel) have a single canonical version, always heard in one key, as with M. C. Hammer's "U Can't Touch This" or U2's "New Year's Day." Levitin recruited forty nonmusicians, paid five dollars for ten minutes (brain imaging pays about fifty), for a vaguely described "memory experiment," and had them sing a favorite song from memory, excluding songs with multiple recordings; examples were Basia's "Time and Tide," Paula Abdul's "Opposites Attract," Madonna's "Like a Virgin" and Billy Joel's "New York State of Mind." Many protested they could not sing. They sang at or very near the original absolute pitch, and again with a second song. Memory held details of a performance, not just an abstraction, including vocal quirks: Michael Jackson's "ee-ee" in "Billie Jean," Madonna's "Hey!", Karen Carpenter's syncopation in "Top of the World," and Springsteen's rasp on the first word of "Born in the U.S.A." Levitin put subjects and records on two stereo channels and it sounded like singing along with the record, though they sang to an internal representation. Most also sang at the right tempo, across a wide range of tempos, and described "singing along with an image" or "recording" in their heads.

Neural basis and ear worms

Mike Posner pointed him to Petr Janata's EEG work: brain waves while listening to and imagining music were nearly impossible to tell apart, suggesting remembering and perceiving use the same regions. Perceiving is a pattern of neurons firing in a configuration (smelling a rose versus rotten eggs use different circuits in the olfactory system, and even the same neurons can have different settings). Remembering may re-recruit those neurons, "re-membering" them into the original club.

This helps explain ear worms (from German Ohrwurm), or stuck song syndrome. Little research exists. Musicians and people with obsessive-compulsive disorder report them more, and OCD medication sometimes helps. Levitin's best explanation is that circuits get stuck in playback mode. Usually a piece under about 15 to 30 seconds, the span of echoic memory, loops rather than a whole song, and simple songs and jingles stick more; this simplicity bias reappears in preference formation, Chapter 8.

Timbre and soundscape

Other labs replicated the singing result. Glenn Schellenberg (Toronto; a founding member of Martha and the Muffins) played Top 40 snippets of about a tenth of a second, a finger snap, shorter than one or two notes, so only timbre was available; subjects matched them to a list of titles at a significant rate, even played backward. Paul Simon, Levitin recalls from the introduction, listens first for timbre.

Songs have a sonic color, like the look of Kansas plains versus northern California coastal forests versus Colorado mountains: we grasp the overall scene before details. This is why early Beatles recordings are identifiable even when the song is unfamiliar, and why Eric Idle and colleagues could create the Rutles as Beatles satire by using the timbral elements. Soundscapes mark eras: 1930s and early 1940s classical records, 1980s rock, heavy metal, 1940s dance hall, late 1950s rock and roll; producers recreate them through microphones and mixing. Echo is a clue: slap-back echo on Gene Vincent's "Be-Bop-A-Lula," Ricky Nelson, Elvis's "Heartbreak Hotel" and Lennon's "Instant Karma," versus the warm tiled-room echo on the Everly Brothers' "Cathy's Clown" and "Wake Up Little Susie."

These findings strongly support encoding of absolute features, and there is no reason musical memory differs from visual, olfactory, tactile or gustatory memory. But constructivist evidence must still be explained, along with our ability to scan songs mentally and imagine transformations.

Scanning a song in the mind

In a demonstration based on Halpern, readers are asked whether "at" appears in "The Star-Spangled Banner" ("what so proudly we hailed at the twilight's last gleaming"). Three things happen: you sing faster than ever heard, which a tape recording could not do; you vary tempo independent of pitch, unlike a tape that rises in pitch when sped up; and on reaching the target you continue to the rest of the phrase. This implies hierarchical encoding, with entry and exit points at phrases.

Musicians confirm this: they learn in units (note groups into phrases, then verses, choruses or movements) and usually cannot start from a few notes before or after a phrase boundary, even from a score. They recall notes faster and more accurately at phrase starts or downbeats than mid-phrase or on weak beats. Notes fall into categories of important and unimportant. Amateur singers store the important tones and the contour and fill in the rest, lowering memory load.

Exemplar theory and the limits of prototypes

Memory theory has converged with category research, so which memory theory is right affects category theory. Hearing a new version of a favorite song, we place it in a category of all its versions. A fan can even displace a prototype: for "Twist and Shout," your prototype may be the Beatles' or the Mamas and Papas' version, but learning the Isley Brothers had a hit two years before the Beatles may reorganize the category. Such top-down reorganization shows categories involve more than prototype theory states. Prototype theory parallels constructivist memory (details discarded, gist kept).

Record-keeping's parallel is exemplar theory. In the 1980s Edward Smith, Douglas Medin and Brian Ross found weaknesses in prototypes:

A theory must therefore handle categories with no prototype, context, and on-the-spot categories, which suggests original details are retained. Gist-only storage could not build "songs with the word love in them but not in the title" (examples: "Here, There and Everywhere," "Don't Fear the Reaper," "Something Stupid," "Cheek to Cheek," "Hello Trouble," "Can't You Hear Me Callin'").

Exemplar theory (Smith and Medin) says every experience, word, kiss, object or song is stored as a trace, a descendant of the Gestalt residue theory. Something belongs to a category if it resembles that category's members more than a competing category's. This indirectly accounts for the Posner and Keele result: a never-seen prototype, being the average, resembles all stored examples most and is categorized quickly. Levitin says this matters for how we can like new music instantly, the subject of Chapter 6.

Multiple-trace models and neural plausibility

The convergence is the family of multiple-trace memory models: each experience is stored with high fidelity, and distortion arises from interference from similar competing traces or degradation of details through normal neurobiology.

The test is whether they predict data on prototypes, constructive memory and abstraction such as transposed recognition. Leslie Ungerleider (NIH) and colleagues found with fMRI that categories (faces, animals, vehicles, foods) occupy specific cortical regions, and lesion patients lose some categories and not others. To test whether detail storage can behave like abstraction, cognitive science uses neural network or parallel distributed processing (PDP) models, computer brain simulations pioneered by David Rumelhart (Stanford) and Jay McClelland (Carnegie Mellon): parallel, layered, with flexibly connected simulated neurons that can be pruned or added. If a model behaves like humans, the theory is plausible.

Douglas Hintzman's influential MINERVA (named for the Roman goddess of knowledge, 1986) stored individual examples yet behaved like a prototype system by comparing new instances to stored ones, as Smith and Medin describe. Goldinger found further support using words spoken in specific voices. There is now an emerging consensus that neither view is right but a hybrid, multiple-trace memory, is, consistent with the music memory experiments and with the exemplar view of categorization.

To account for extracting invariants while listening, we must compute melodic intervals and tempo-free rhythm alongside absolute pitch, rhythm, tempo and timbre. Robert Zatorre and colleagues at McGill found melodic calculation centers in the dorsal temporal lobes, above the ears, tracking interval size and distances between pitches, building a pitch-free template for recognizing transposition. Levitin's own imaging shows familiar music activates those regions plus the hippocampus, crucial for memory encoding and retrieval, suggesting both abstract and specific information are stored, perhaps for all senses.

Cues, repertoire and why music retrieves memories

Because multiple-trace models preserve context, they explain retrieving nearly forgotten memories: a long-unsmelled odor, or an old song, bringing back a long-ago event. We keep a repertoire of memories, like a photo album, recalled to remind us who we are, like a musician's repertoire.

Every experience is potentially encoded, not in one place (the brain is no warehouse) but in groups of neurons configured in particular ways. The obstacle to recall is finding the right cue, not absence of storage; the more a memory is accessed, the more active its retrieval circuits and the easier the cues. With the right cues any past experience could be accessed.

Levitin's example is your third-grade teacher: generic cues (desks, hallways, playmates) bring little, but a class photo would release names, subjects and lunchtime games. A song is a specific, vivid set of cues. Because context is encoded with traces, music is cross-coded with the events of its time. A memory maxim says unique cues work best; the more contexts a cue is tied to, the weaker it is. So songs that kept playing (classic rock or classical stations with limited repertoires) are poor cues, but one not heard since a particular time opens the floodgates. Because memory and categorization are linked, a song can also trigger categorical memories: hearing "YMCA" may bring "I Love the Nightlife" and "The Hustle."

Memory, repetition and emotion

Levitin says without memory there would be no music, which is based on repetition (as John Hartford's song "Tryin' to Do Something to Get Your Attention" notes). Music works because we remember tones just heard and relate them to those now playing; phrases return in variation or transposition, tickling memory while engaging emotional centers. In the last decade neuroscience has shown how closely memory and emotion are linked: the amygdala, long considered the seat of emotion, sits next to the hippocampus and is highly active for experiences or memories with strong emotional content. Every neuroimaging study in Levitin's lab has shown amygdala activation to music, not to random sounds or tones. Skillful repetition by a master composer is emotionally satisfying to our brains and makes listening pleasurable.


Chapter 6: After Dessert, Crick Was Still Four Seats Away from Me — Music, Emotion, and the Reptilian Brain

Pulse, meter and how composers play with it

Levitin opens by noting that most music is foot-tapping music: it has a regular, evenly spaced pulse that we can tap along to, physically or in our heads. That pulse makes us expect events at particular moments, and like the clickety-clack of a train it tells us we are moving forward and all is well.

Composers sometimes suspend the pulse. The opening of Beethoven's Fifth stops after its short phrase, so we do not know when sound will return; after the repeated phrase on different pitches, a steady beat takes over. Other composers state the pulse but present it softly before arriving at a heavy statement of it. "Honky Tonk Women" starts with cowbell, adds drums, then electric guitar, with the meter unchanged while the intensity of the strong beats grows (and on headphones the cowbell sits in one ear). "Back in Black" opens with high-hat and muted guitar chords for eight beats before the full guitar arrives, and Hendrix's "Purple Haze" opens with eight quarter notes on guitar and bass that set the meter before Mitch Mitchell's drums enter. Some composers tease by setting up a meter and then changing it when the other instruments join, as in Stevie Wonder's "Golden Lady" and Fleetwood Mac's "Hypnotized"; Frank Zappa was a master of this.

Groove

Some music is more rhythmically driven than other music. "Eine Kleine Nachtmusik" and "Stayin' Alive" both have a clear meter, but the second makes people dance. A readily predictable beat helps us be moved physically and emotionally. Composers get this by subdividing the beat in different ways and accenting some notes more than others, and performance matters a great deal too.

Groove, Levitin explains, is the quality that creates strong forward momentum, like a book you cannot put down. A song with good groove pulls us into a sonic world we do not want to leave; we sense the pulse, yet outside time seems to stop. Groove belongs to a particular performer and performance rather than the written score, and it can come and go from day to day even with the same band. Listeners disagree, but he suggests common ground: "Shout" by the Isley Brothers, "Super Freak" by Rick James, "Sledgehammer" by Peter Gabriel, "I'm On Fire" by Springsteen, "Superstition" by Stevie Wonder and "Ohio" by the Pretenders all have great groove, though they differ widely. There is no formula, as R&B musicians trying to copy the Temptations or Ray Charles know, and the small number of songs that have it shows how hard it is.

He analyses "Superstition." Part of the secret is Stevie Wonder's drumming. In the opening seconds the high-hat plays alone; drummers treat the high-hat as their timekeeper and reference point even when it is inaudible in loud passages. Wonder never plays the pattern the same way twice, adding extra taps, hits and rests, and every cymbal note has a slightly different volume, which adds tension. The snare enters and the high-hat pattern follows, which the author writes out as syllables. He keeps listeners oriented by holding part of the pattern constant, repeating the start of each line while varying the second half in a call-and-response way, and in one place he strikes the cymbal differently so it seems to change its vowel sound.

Musicians generally agree that groove works best when it is not strictly metronomic. Drum machines have produced danceable songs, such as Michael Jackson's "Billie Jean" and Paula Abdul's "Straight Up," but the gold standard is a drummer who shifts tempo slightly for aesthetic and emotional reasons, so that the rhythm track "breathes." Steely Dan spent months editing, shifting, pushing and pulling the drum-machine parts on Two Against Nature to make them sound human. Such local tempo changes do not alter meter, which is the structure of the pulse and whether beats group in twos, threes or fours, nor the global pace.

Classical music is not usually discussed in terms of groove, but operas, symphonies, sonatas, concertos and quartets also have meter and pulse, shown by the conductor, who may stretch or compress beats for emotional effect. Real human exchanges such as pleas for forgiveness, anger, courtship, storytelling and parenting do not run at a machine's rate, so music reflecting emotional life must swell, contract, speed up, slow down and pause.

Expectation and the brain's model of the pulse

We can only notice such timing variations if a brain system has extracted when beats should occur. The brain builds a model, a schema, of a constant pulse so it can detect departures, just as it needs a mental representation of a melody to appreciate a musician's liberties with it. Metrical extraction is therefore central to musical emotion. Music communicates emotion through systematic violations of expectation, in pitch, timbre, contour, rhythm, tempo or other domains, and some violation must occur. Organized sound with no unexpected element is flat and robotic; scales are organized, but parents tire of hearing children practise them after five minutes.

Neural basis of rhythm, meter and melody

Lesion studies show rhythm and metrical extraction are neurally unrelated to each other. Left-hemisphere damage can remove the ability to perceive and produce rhythm while meter extraction survives, and right-hemisphere damage shows the opposite pattern. Both are separate from melody processing: Robert Zatorre found right temporal lobe lesions harm melody perception more than left ones, and Isabelle Peretz found a right-hemisphere contour processor that outlines a melody for later recognition, dissociable from rhythm and meter circuits.

The Desain and Honing foot-tapping computer

Computer models help us understand the brain. Peter Desain and Henkjan Honing in the Netherlands built a beat-extraction model that relied chiefly on amplitude, since meter is defined by loud and soft beats alternating at regular intervals. For showmanship they connected it to a small motor in a shoe, and Levitin saw it demonstrated at CCRMA in the mid-nineties. A man's size-nine black wingtip hung from a metal rod, wired to the computer, and after a few seconds of "listening" to any CD handed over it tapped on a piece of plywood. Perry Cook afterward asked whether it came in brown. The system shared human weaknesses, sometimes tapping in half time or double time relative to professional musicians, as amateurs often do; a model that errs like a human is better evidence that it replicates the computations behind thought.

The cerebellum, the "reptilian brain"

The cerebellum, Latin for "little brain," hangs beneath the cerebrum at the back of the neck, has two sides divided into subregions, and is among the oldest brain parts in evolutionary terms, which is why it is sometimes called the reptilian brain. It is about 10 percent of the brain's weight but holds 50 to 80 percent of its neurons. Its function, timing, is crucial to music. Traditionally it guides movement, which in most animals is repetitive and oscillatory, such as a steady gait in walking or running or steady fin and wing rates in fish and birds; the cerebellum helps maintain that rate. Difficulty walking is a hallmark of Parkinson's disease, which is accompanied by cerebellar degeneration.

In Levitin's lab, the cerebellum was strongly active when people listened to music but not noise, apparently tracking the beat. It also appeared when people heard music they liked versus disliked, or familiar versus unfamiliar music. He and others wondered whether those activations were errors. In summer 2003 Vinod Menon told him about Harvard's Jeremy Schmahmann, who has argued against traditionalists that the cerebellum serves emotion as well as timing and movement, using autopsies, neuroimaging, case studies and other species. This would explain activation to liked music. Schmahmann notes massive connections from the cerebellum to the amygdala (remembering emotional events) and the frontal lobe (planning and impulse control). Why emotion and movement share a region found even in snakes and lizards is unknown, but informed speculation came from James Watson and Francis Crick.

The Cold Spring Harbor workshop

Cold Spring Harbor Laboratory on Long Island, directed by Nobel laureate Watson, works on neuroscience, neurobiology, cancer and genetics, and offers degrees through SUNY Stony Brook. Levitin's colleague Amandine Penel, who took her Ph.D. in music cognition in Paris while he did his at Oregon, was a postdoc there. The lab runs multi-day residential workshops that gather world experts, often holding opposing views, hoping agreement on parts of a problem will speed science.

Levitin was surprised to receive, among mundane McGill e-mails, an invitation to a four-day workshop titled "Neural Representation and Processing of Temporal Patterns." Its text asked how time is represented and how complex temporal patterns are perceived or produced, and aimed to unite psychologists, neuroscientists and theorists, and to extend work on single time intervals to patterns of multiple intervals.

He thought his name was a mistake, since the list held giants of the field. Paula Tallal had found with Mike Merzenich that dyslexia relates to a timing deficit in children's auditory systems, and had published influential fMRI work on speech in the brain. Rich Ivry, Levitin's "intellectual cousin," a student of Steve Keele at Oregon, had done pioneering work on the cerebellum and motor control and is low-key and incisive. Randy Gallistel was a mathematical psychologist modelling learning and memory in humans and mice. Bruno Repp had been Penel's postdoctoral advisor and a reviewer of Levitin's first two papers (on people singing pop songs near the right pitch and tempo). Mari Reiss Jones had done key work on attention in music cognition and a model of how accents, meter, rhythm and expectation yield musical structure knowledge. John Hopfield, inventor of Hopfield nets, a major PDP network class, was also there. Levitin felt like a girl backstage at a 1957 Elvis concert.

The meeting was intense. Participants disagreed on basic matters, such as how to distinguish an oscillator from a timekeeper and whether estimating a silent interval differs neurally from estimating one filled with regular pulses. They realized that progress was impeded because people used different terms for the same things, and one word like "timing" for very different things, with differing assumptions. Even "planum temporale" was defined anatomically by one person and functionally by another. They argued over gray versus white matter and over whether synchronous events must coincide exactly or only appear simultaneous perceptually. Evenings brought catered dinners, beer and red wine; Levitin's student Bradley Vines attended as observer and played saxophone, Levitin played guitar with other musicians, and Amandine sang.

Most attendees, focused on timing, had paid little attention to Schmahmann, but Ivry knew and liked his work. He showed Levitin similarities between music perception and motor action planning that Levitin had not seen in his own experiment, and agreed the mystery of music must involve the cerebellum. Watson told Levitin he too found a connection among cerebellum, timing, music and emotion plausible. But what was it, and what was its evolutionary basis?

Visiting the Salk Institute

A few months later Levitin visited his collaborator Ursula Bellugi at the Salk Institute in La Jolla, which sits on land overlooking the Pacific. A student of Roger Brown at Harvard, Bellugi runs the Cognitive Neuroscience Laboratory. She was first to show that sign language is truly a language with syntax, so Chomsky's linguistic module is not only for speech, and has done major work on spatial cognition, gesture, neurodevelopmental disorders and neuroplasticity. She and Levitin had worked ten years on the genetic basis of musicality, and he visits yearly so they can sit at one screen and discuss chromosome diagrams and brain activations.

Salk held a weekly professors' lunch around a square table with its director Francis Crick, rarely open to visitors, where scientists speculated freely. Levitin had dreamed of attending. He found Crick's The Astonishing Hypothesis, arguing consciousness arises from neurons, glial cells and their molecules, interesting but disliked mapping the mind for its own sake. What drew him was Crick's What Mad Pursuit and a passage about the end of the war. Crick took stock of a mediocre degree, Admiralty work, narrow knowledge of magnetism and hydrodynamics he did not love, and no publications, then realized lack of qualification could be an advantage, since most scientists by thirty are trapped in their expertise, whereas he knew nothing and so was free to choose. This encouraged Levitin, who also started science late, to treat inexperience as licence to think differently.

One morning Levitin arrived at seven; Bellugi had been in since six. She asked whether he would like to meet Francis that day, and he panicked, recalling meeting Watson only months earlier. The panic recalled his days as a young record producer. Michelle Zarin, manager of San Francisco's Automatt studio, held Friday wine-and-cheese gatherings for an inner circle, and while Levitin worked with unknown bands like the Afflicted and the Dimes he watched figures like Carlos Santana, Huey Lewis and producers Jim Gaines and Bob Johnston go in. One Friday Ron Nevison, engineer of his favourite Led Zeppelin records, was to attend. Zarin placed him in a semicircle, but Nevison ignored him; twenty minutes passed while Boz Scaggs played ("Lowdown," "Lido"), and when "We're All Alone" came on its lyrics got to him, so he introduced himself. Nevison shook his hand and went back to his conversation. Zarin scolded him: waiting for her introduction would have reminded Nevison he was the thoughtful young apprentice she had mentioned. He never saw Nevison again.

The lunch with Crick

At lunch they climbed three flights to the professors' room, Ursula seating him about four places from Crick, then in his late eighties and frail. The talk was a cacophony: a cancer gene, squid visual-system genetics, a drug to slow Alzheimer's memory loss. Crick mostly listened and spoke too softly to hear. After dessert he was still four seats away, talking with someone facing away. Levitin wanted to discuss the book and the relation among cognition, emotion and motor control, and a genetic basis for music. Ursula promised an introduction on the way out; he expected a quick hello. She took him by the elbow (she is four foot ten) to Crick, who was discussing leptons and muons, and introduced him as working on Williams syndrome and music. She started pulling him toward the door, but Crick's eyes lit up at "music," he sent his lepton colleague away, and Ursula noted they had time now.

Crick asked about music neuroimaging, and Levitin described their cerebellum findings. The cerebellum's role in helping performers and conductors keep time was known, and many assumed it did the same for listeners, but where did emotion fit?

Emotion, evolution and the cerebellum

Scientists cannot agree what emotions are. Levitin distinguishes emotions (temporary states, usually caused by an external event present, remembered or anticipated), moods (longer-lasting, with or without external cause) and traits (tendencies, as in a generally happy person). Some use "affect" for positive or negative valence, with emotion for specific states: positive ones include happiness and satiety, negative ones fear and anger.

Evolutionarily, he and Crick agreed, emotions were tied to motivation: neurochemical states prompting action for survival. Seeing a lion produces fear, from a particular mix of neurotransmitters and firing rates, and we run without thinking. Bad food produces disgust, with a scrunched nose, protruding tongue and constricted throat. Finding water after wandering produces elation and satiety, helping us remember the spot. Not all emotions lead to movement, but many important ones do, especially running, which is faster and more efficient with a regular gait, a cerebellar job. Ancestors whose emotional systems linked directly to motor systems reacted faster and survived to pass on genes.

Crick cared more for data than origins. He knew Schmahmann was reviving forgotten ideas, such as a 1934 paper linking the cerebellum to arousal, attention and sleep modulation. In the 1970s lesions to certain cerebellar regions were found to change arousal dramatically: monkeys with one lesion showed "sham rage" with no environmental cause (though the surgery gave reason), while lesions elsewhere caused calm and were used clinically to soothe schizophrenics. Electrical stimulation of the vermis, a thin central strip, can cause aggression in humans, and another region reduces anxiety and depression.

Crick pushed away his dessert plate and gripped a glass of ice water; Levitin could see the veins in his hands and almost his pulse. The room was still and waves crashed outside.

Ear-to-cerebellum pathways, redundancy and startle

They discussed 1970s work showing the inner ear does not send all connections to auditory cortex. In cats and rats there are direct projections from the ear to the cerebellum, coordinating orientation toward sounds, along with location-sensitive cerebellar neurons. These areas project to frontal regions, the inferior frontal and orbitofrontal cortex, that Levitin's work with Menon and Bellugi found active in both language and music. Why would ear connections bypass auditory cortex for a motor (and perhaps emotional) center?

Redundancy and distribution of function are core neuroanatomical principles: organisms must survive to reproduce, injury is likely, and so important systems evolved backup pathways. Perception is tuned to detect change, which may signal danger. Vision sees millions of colors and works at one photon in a million, but is most sensitive to sudden change, and area MT detects motion. An insect on the neck is slapped because touch notices subtle pressure change; a change in smell, such as pie on a windowsill, alerts us. Sound causes the greatest startle, jumping, ducking or ear-covering. Auditory startle is the fastest and arguably most important, because a moving large object disturbs air, which we hear. Redundancy means the nervous system must react to sound even if partly damaged. Deeper looking reveals latent circuits; people with cut visual pathways can orient to and sometimes identify objects despite claiming blindness. An auditory backup involving the cerebellum preserves quick emotional and motor reactions to dangerous sounds.

Habituation

Linked to startle and sensitivity to change is habituation: a fridge hum fades from notice. A rat in its hole hears a loud noise overhead that could be a predator's step or a branch tapping in the wind. After a dozen or two taps he should ignore it, but if intensity or frequency changes, conditions have changed and he should attend: the wind may have strengthened and threaten his burrow, or died down so he can go out. Habituation separates threatening from nonthreatening stimuli. The cerebellum acts as a timekeeper, so when it is damaged, tracking regularity fails and habituation is lost.

Williams syndrome

Bellugi told Crick of Albert Galaburda's finding at Harvard that people with Williams syndrome have cerebellar formation defects. Williams arises when about twenty genes are missing on chromosome 7, in one of twenty thousand births, about a quarter as common as Down syndrome, from an early fetal transcription error. Losing those twenty of roughly twenty-five thousand genes is devastating: profound intellectual impairment, with few able to count, tell time or read. Yet language is largely intact, they are very musical, unusually outgoing, pleasant, emotional and gregarious, loving music and meeting people. Schmahmann found cerebellar lesions can produce Williams-like over-friendliness toward strangers.

Levitin describes visiting Kenny, a cheerful, music-loving fourteen-year-old with WS whose IQ under fifty meant a seven-year-old's capacity. He had very poor eye-hand coordination: his mother helped with sweater buttons, he wore Velcro shoes, and he struggled with stairs and getting food to his mouth. Yet he played clarinet, executing complex finger movements for several learned pieces without naming notes or saying what he was doing, as though his fingers had their own mind; the coordination problem vanished while playing, then returned when he needed help opening the case.

Allan Reiss at Stanford found the neocerebellum, the newest cerebellar part, is larger than normal in WS. Movement entrained to music seemed to differ from other movement, suggesting the cerebellum is the part with a "mind of its own" and might reveal how it normally affects music processing. The cerebellum thus relates to emotion (startle, fear, rage, calm, gregariousness) and to auditory processing.

The binding problem and "look at the connections"

Crick raised the binding problem. Objects have features processed by separate subsystems (for vision: color, shape, motion, contrast, size), and the brain must bind them into a whole. Patients with Balint's syndrome recognize only one or two features without holding them together; some can say where an object is but not its color, or vice versa; others hear timbre and rhythm but not melody. Peretz found a patient with absolute pitch who is tone deaf: he names notes perfectly but cannot sing.

Crick proposed synchronous firing of neurons across the cortex as a solution, and his "astonishing hypothesis" held that consciousness emerges from 40 Hz synchronous firing. Neuroscientists considered cerebellar operations preconscious, controlling running, walking, grasping and reaching, but Crick saw no reason cerebellar neurons could not fire at 40 Hz and contribute, though we do not attribute humanlike consciousness to cerebellum-only animals such as reptiles. "Look at the connections," he said. He had taught himself neuroanatomy at Salk and had little patience for cognitive neuroscientists who ignored their founding principle of constraining hypotheses with the brain; he thought progress needed rigorous study of structural and functional detail. His colleague returned about an appointment, and as they stood Crick again said "Look at the connections." Levitin never saw him again; Crick died months later.

Connecting the cerebellum, frontal lobe and reward system

The cerebellum-music link was not hard to see. At Cold Spring Harbor people noted that the frontal lobe, the seat of advanced cognition, connects directly and in both directions to the cerebellum, the most primitive part. Frontal regions Tallal studied for fine speech-sound distinctions also connect to it, and Ivry's work showed links among frontal lobes, occipital cortex, motor strip and cerebellum. Another player lay deep in the brain.

In 1999 Anne Blood, a postdoc with Robert Zatorre at the Montreal Neurological Institute, showed that intense musical emotion, which subjects called "thrills and chills," was tied to reward, motivation and arousal regions: ventral striatum, amygdala, midbrain and frontal cortex. Levitin focused on the ventral striatum, which includes the nucleus accumbens (NAc), the brain's reward center, important in pleasure and addiction, active when gamblers win or drug users take their drug, and tied to opioid transmission through dopamine release. Avram Goldstein showed in 1980 that music's pleasure could be blocked by naloxone (the text spells it "nalaxone"), believed to interfere with dopamine in the NAc. But PET lacks the resolution to detect the small NAc, whereas Levitin and Menon had higher-resolution fMRI data.

To prove the story, they needed to show the NAc acted at the right time in a sequence, after frontal structures processing musical structure and meaning; that its activity coincided with other dopamine-producing and transmitting structures, so it was not coincidence; and that the cerebellum, which has dopamine receptors, also appeared. Menon had read Karl Friston's papers on functional and effective connectivity analysis, which reveals how regions interact during cognition, constrained by anatomical connections, allowing moment-by-moment examination of music-induced networks, something Crick would have wanted. It was hard: scans yield millions of data points, a session can fill an ordinary hard drive, ordinary analysis takes months, and no ready software existed. Menon spent two months working out the equations, then they reanalyzed their data from people listening to classical music.

The findings: a cascade through the brain

The results matched their hopes. Music activated regions in order: auditory cortex first, for initial sound processing; then frontal regions such as BA44 and BA47, previously linked to musical structure and expectation; then the mesolimbic system, involved in arousal, pleasure, opioid transmission and dopamine production, culminating in the nucleus accumbens. Cerebellum and basal ganglia were active throughout, presumably supporting rhythm and meter. Music's rewarding, reinforcing effects thus seem mediated by rising NAc dopamine and the cerebellum's emotional regulation via frontal and limbic connections. Current theories link positive mood and affect to higher dopamine, which is why many newer antidepressants act on the dopamine system. Music clearly improves mood, and now Levitin thinks he knows why.

Closing synthesis

Music mimics some features of language and conveys some of the same emotions as vocal communication, but nonreferentially and nonspecifically. It uses some of the same regions as language but, far more, taps primitive structures of motivation, reward and emotion. At the first cowbell hits of "Honky Tonk Women" or notes of "Sheherazade," brain systems synchronize neural oscillators to the pulse and predict the next strong beat, keep updating estimates, are satisfied when mental and real beats match, and delight when a musician skilfully violates the expectation, a joke we are all in on. Music breathes and varies like the real world, and the cerebellum enjoys adjusting to stay synchronized.

Groove is subtle timing violation. Just as the rat responds emotionally to a break in the branch's rhythm, we respond to timing violations in groove, but the rat, lacking context, feels fear, while we know from culture and experience that music is not threatening and take pleasure and amusement. This response runs by the ear-cerebellum-nucleus accumbens-limbic circuit, not the ear-auditory cortex circuit, so it is largely pre- or unconscious, going through the cerebellum rather than the frontal lobes. Remarkably, all these pathways combine into one experience of a song.

The story of the brain on music is an orchestration of the oldest and newest brain parts, from the cerebellum at the back of the head to the frontal lobes behind the eyes, and a choreography of neurochemical release and uptake between logical prediction systems and emotional reward systems. Loved music reminds us of other music and activates memory traces of emotional times. As Crick repeated while leaving the lunchroom, it is all about connections.


Chapter 7: What Makes a Musician? Expertise Dissected

Opening: Sinatra and the question of expertise

Levitin opens with Frank Sinatra's album Songs for Swinging Lovers. He says he is not a Sinatra fan: he owns only about six of the more than two hundred albums, dislikes the films, finds much of the repertoire sappy and finds the post-1980 Sinatra too cocky. Billboard once hired him to review Sinatra's last album, the duets record with Bono, Gloria Estefan and others, and he panned it, saying Sinatra sang like a man who had just had somebody killed. On Swinging Lovers, though, every note is placed perfectly in time and pitch. The timing departs from the notated score, yet it fits emotions that words cannot describe. The phrasing is so detailed and idiosyncratic that nobody Levitin has met can match it when singing along.

This leads to the chapter's questions. How do people become expert musicians? Why do so few of the millions of children who take lessons keep playing as adults? Many people tell him their lessons "didn't take," and he thinks they are too hard on themselves. The gap between experts and everyday players discourages people, and it does so more for music than for other skills. Most of us cannot play like Shaquille O'Neal or cook like Julia Child, but we still enjoy a backyard basketball game or a holiday meal. He sees the gap as cultural and specific to modern Western society. Laboratory work shows that even a little childhood training builds more efficient neural circuits for music. Lessons teach us to listen better and to perceive structure and form faster, which helps us work out what we like.

He then asks about true experts such as Alfred Brendel, Sarah Chang, Wynton Marsalis and Tori Amos. Do they differ from the rest of us in kind or only in degree? Do composers and songwriters have a different set of skills from performers?

Talent versus practice

Expertise has been a major topic in cognitive science for about thirty years, and musical expertise is usually studied as part of it, defined as technical mastery of an instrument or of composition. Michael Howe, Jane Davidson and John Sloboda started an international debate by asking whether the everyday idea of "talent" can be defended scientifically. They posed a dichotomy: high achievement comes either from innate brain structures (talent) or from training and practice. They defined talent by four features:

Early identification means studying children, and talent may show up differently in different children.

Children clearly acquire skills at different speeds, as the ages for walking, talking and toilet training show, even within one family. Genes may contribute, but motivation, personality and family dynamics are hard to separate from them, and they can mask genetic effects on musical ability. Brain studies have not helped much because cause and effect are hard to separate. Gottfried Schlaug at Harvard found that the planum temporale, a region of auditory cortex, is larger in people with absolute pitch than in others. It is unclear whether it starts larger or grows because absolute pitch is acquired. The picture is clearer for motor areas. Thomas Elbert's studies of violinists show that the brain region controlling the left hand, which needs the most precision, grows with practice. Whether some people are predisposed to such growth is still unknown.

The best evidence for talent is that some people simply learn faster. The evidence for practice comes from how much the high achievers train. As in mathematics, chess and sport, musical experts need long periods of instruction and practice. In several studies the best conservatory students had practiced the most, sometimes twice as much as students rated less good. In another study, students were secretly split into two groups according to teachers' judgments of their talent. Years later, the top performers were those who had practiced most, whichever group they had been put in. This suggests practice causes achievement rather than merely accompanying it. It also suggests "talent" is used circularly, applied only after someone has already achieved something.

Ericsson and the ten-thousand-hour figure

Anders Ericsson of Florida State and his colleagues treat musical expertise as one case of how people become expert at anything, so they study writers, chess players, athletes, artists and mathematicians alongside musicians. First, an expert is someone who has reached a high level relative to others, so expertise is a social judgment, and it concerns a field we care about. Sloboda notes that one could become expert at folding one's arms or saying one's own name, but that is not equivalent to expertise at chess, Porsche repair or stealing the British crown jewels without being caught.

The emerging finding is that about ten thousand hours of practice are needed for world-class mastery in anything. The figure turns up for composers, basketball players, fiction writers, ice skaters, concert pianists, chess players and master criminals. It equals roughly three hours a day, or twenty a week, for ten years. It does not explain why some people get nowhere or why some gain more from practice, but nobody has found world-class expertise reached faster. It seems the brain needs this long to absorb what mastery requires.

Levitin says this fits what we know about learning. Learning means consolidating information in neural tissue, and each further experience strengthens the memory trace. This holds under multiple-trace theory and its variants: memory strength depends on how often the stimulus was experienced. Strength also depends on how much we care. Neurochemical tags mark emotional experiences, positive or negative, as important. He tells students that to do well on a test they must care about the material. Caring may partly explain early differences in how fast people learn. A piece I love will be practiced more, and I will tag its sounds, finger movements and, for wind players, breathing as important. Liking an instrument's sound also makes one attend to subtle differences in tone and how to shape it. Caring leads to attention, and together they produce measurable neurochemical change: dopamine, linked to emotional regulation, alertness and mood, is released and helps encode the memory. Some people are less motivated, and their practice is less effective because of motivation and attention.

The theory persuades because the number recurs across domains, and scientists favor a number or formula that recurs. Like many theories, though, it has holes and must face rebuttals.

The Mozart rebuttal

The classic objection is Mozart, said to have composed symphonies at four, which cannot be ten thousand hours even at forty hours a week from birth. Levitin first corrects the facts: Mozart began composing at six and wrote his first symphony at eight. That is still unusual, but precocity is not expertise, and many children write music, some even large works by eight. Mozart also had intensive training from his father, regarded as Europe's greatest music teacher and a stern taskmaster. If Mozart began at two and worked thirty-two hours a week, he would reach ten thousand hours by eight. The theory also does not say that ten thousand hours are needed to write a symphony. The real question is whether that first symphony was the work of an expert or whether his expertise came later.

John Hayes of Carnegie Mellon asked exactly this: would Symphony no. 1 look like genius if Mozart had written nothing else? It may be of historical rather than aesthetic interest. Hayes examined the programs of leading orchestras and the catalog of commercial recordings, assuming better works are performed and recorded more. Mozart's early works were rarely performed or recorded, and musicologists treat them as curiosities that did not predict the later masterpieces. The works regarded as truly great came after his ten thousand hours.

As in the earlier debates about memory and categorization, Levitin concludes that the truth lies between the extremes, so he turns to what geneticists say.

Genetics, twins and the limits of the evidence

Geneticists look for gene clusters tied to observable traits and expect a genetic contribution to music to run in families, since siblings share half their genes. Separating genes from environment is hard. Environment includes the womb: the mother's diet, smoking and drinking, and the nutrients and oxygen the fetus gets. Even identical twins can have different womb environments depending on space, movement and position. Music runs in families, but musical parents also encourage their children, and the siblings get similar support. By analogy, French speaking runs in families, yet nobody calls it genetic.

One method is studying identical twins, especially those raised apart. The Minnesota twins registry, kept by David Lykken, Thomas Bouchard and colleagues, follows identical and fraternal twins raised apart and together. Fraternal twins share 50 percent of genetic material and identical twins 100 percent, so a genetic trait should appear more often in identical pairs and should persist when they grow up in separate environments. Behavioral geneticists look for such patterns and estimate heritability.

The newest approach examines gene linkages, finding genes linked to a heritable trait. Levitin avoids saying "responsible for" because gene interactions are complicated and a single gene cannot be said to cause a trait. Having a gene also does not mean it is active, since not all genes are expressed at all times. Gene chip expression profiling shows which genes are active at a given moment. About twenty-five thousand genes direct protein synthesis for all biological functions, such as hair growth and color, digestive fluids, saliva and whether you end up six feet or five feet tall. Something starts the growth spurt at puberty and stops it years later. A sample of your RNA could show whether your growth gene is currently expressed. Analyzing gene expression in the brain is not practical, because it needs a piece of brain tissue, and most people find that unpleasant.

Twins separated at birth, sometimes unaware of each other, and raised in places that differed in geography (Maine versus Texas, Nebraska versus New York), wealth and values, show striking similarities when found twenty or more years later. One woman backed into the sea at the beach, and so did her twin. One man sold life insurance, sang in a church choir and wore Lone Star beer belt buckles, and so did his twin. Such studies suggested musicality, religiosity and criminality have strong genetic components.

Levitin offers two alternative explanations. The first is statistical. With enough comparisons you will find odd coincidences between any two strangers, who share distant ancestors anyway. His example is a person who washes hair on Tuesdays and Fridays with particular shampoos, then reads The New Yorker and listens to Puccini. We all differ in thousands of ways, and occasional overlaps are no more surprising than guessing a number from one to a hundred, which succeeds one time in a hundred if you play long enough.

The second is social psychological. Appearance, assumed genetic, affects how others treat us, which shapes us. Literature from Cyrano de Bergerac to Shrek dramatizes people shunned for their looks. In the other direction, attractive people tend to earn more, get better jobs and report greater happiness. Features that signal trustworthiness, such as large eyes and raised eyebrows, draw trust, and tall people may get more respect. Our encounters are shaped by how others see us, so identical twins may develop similar personalities. Someone with downturned eyebrows looks angry and is treated so, a defenseless-looking person gets exploited, and someone who looks like a bully is repeatedly invited to fight and becomes aggressive. Actors such as Hugh Grant, Judge Reinhold, Tom Hanks and Adrien Brody have innocent faces. Here genes influence personality only indirectly.

Levitin applies this to musicians, especially singers. Doc Watson's voice sounds sincere and innocent, and perhaps he succeeded because of how people react to his voice. He means expressiveness, not having a great instrument like Ella Fitzgerald's or Placido Domingo's. In Aimee Mann's singing he hears a little girl's vulnerable innocence, as if confessing to a close friend. She may have a vocal quality that makes listeners attribute those feelings to her whether or not she has them. The essence of performance is conveying emotion, and it may not matter whether the artist feels it or was born sounding as if she does.

He adds that none of this means these people did not work. Every successful musician he knows worked hard, and the "overnight sensations" he has known spent five or ten years getting there. Genetics is a starting point that can influence personality, career and the choices made in it. Tom Hanks will not get Arnold Schwarzenegger's roles. Schwarzenegger worked at bodybuilding but had a predisposition for it, and being six foot ten predisposes someone to basketball rather than horse racing, though he must still learn the game and practice for years. Body type, largely though not wholly genetic, creates predispositions in basketball, acting, dancing and music. Musicians use their bodies, so genes may strongly influence which instrument one plays well and whether one becomes a musician, less so for composing and arranging.

Levitin's own story: small hands and the guitar

At six, Levitin saw the Beatles on The Ed Sullivan Show and decided to play guitar. His old-school parents did not consider the guitar a serious instrument and told him to play the family piano. He cut out pictures of classical guitarists like Andrés Segovia and left them around the house. He also had a prominent lisp, which he kept until age ten, when a school speech therapist took him out of fourth grade and spent two grueling years, three hours a week, changing how he said s. He argued that the Beatles must be serious artists since they shared a stage with Beverly Sills, Rodgers and Hammerstein and John Gielgud, and he rendered this with his lisp ("therious," "artithts").

By 1965, at eight, the guitar was everywhere, and with San Francisco fifteen miles away he felt a cultural revolution. His parents were still reluctant, perhaps because of the guitar's link to hippies and drugs, or because he had failed to practice piano the previous year. He pointed out that the Beatles had now been on Ed Sullivan four times, and they half-relented and asked a friend. Jack King, an old college friend, visited, a big man with large hands and a short black crew cut, holding a classical guitar like a baby. He played, then, without talking to Levitin or looking at him, pressed his palm against the boy's and told his mother the hands were too small for the guitar.

Levitin later learned about three-quarter and half-size guitars (he owns one) and about Django Reinhardt, one of the greatest guitarists ever, who had only two fingers on his left hand. But an eight-year-old finds adults' words unbreachable. By 1966 he had grown a little and was playing clarinet, glad to make music while the Beatles' "Help" played. At sixteen he bought his first guitar and learned reasonably well, since the rock and jazz he plays do not need classical guitar's long reach. His first song was Led Zeppelin's "Stairway to Heaven." Some parts will always be hard for him, but that is true for every instrument. Putting his hands into Jimmy Page's handprints in the cement on Hollywood Boulevard, he was surprised that Page's hands were no bigger than his.

He also once shook hands with the jazz pianist Oscar Peterson, whose hands were the largest he has shaken, at least twice the size of his own. Peterson began with stride piano, a style from the 1920s in which the left hand plays octave bass and the right plays melody. Stride players must reach far-apart keys with minimal movement, and Peterson could stretch an octave and a half. His style is tied to chords small hands cannot play. Forced to play violin as a child, he would have struggled to play a semitone on the small neck with such wide fingers.

Some people thus have a biological predisposition for particular instruments or for singing. A cluster of genes might also work together to produce component skills a musician needs: eye-hand coordination, muscle and motor control, tenacity, patience, memory for structures and patterns, and a sense of rhythm and timing. Determination, self-confidence and patience help in becoming great at anything. Successful people also have, on average, more failures than unsuccessful ones, because failure happens, sometimes randomly, and what matters is what you do afterward. They do not quit. Levitin lists the president of FedEx, the novelist Jerzy Kosinski, van Gogh, Bill Clinton and Fleetwood Mac as people with many failures who learned and carried on. This trait may be partly innate, but environment must play a role.

Genes plus environment: population-level prediction

The scientists' best current guess is that genes and environment each account for about half of complex cognitive behavior. Genes may transmit a tendency toward patience, coordination or passion, but life events, broadly construed to include the food you ate and your mother ate while pregnant, affect whether the tendency is realized. Early trauma such as losing a parent or suffering abuse is an obvious case where a predisposition is heightened or suppressed. Because of this interaction, predictions can only be made for populations, not individuals. Knowing someone has a predisposition to criminal behavior says nothing about whether he will be jailed in five years, but among a hundred such people some percentage probably will be, and some will never get in trouble.

Any musical genes found would work the same way: a group with them is more likely to yield expert musicians, but we cannot say which individuals. That assumes we can find the genetic correlates and agree what musical expertise is. It must be more than strict technique. Listening, enjoyment, musical memory and engagement with music are all part of a musical mind, so musicality should be defined inclusively, to avoid excluding people who are musical broadly but not technically. Irving Berlin, among the twentieth century's most successful composers, was a poor instrumentalist who could barely play piano.

Emotion and expressivity

Even elite classical players are more than technicians. Arthur Rubinstein and Vladimir Horowitz, regarded as two of the century's greatest pianists, made small technical mistakes surprisingly often: wrong, rushed or badly fingered notes. A critic wrote that he would take Rubinstein's passionate interpretations, mistakes included, over a twenty-two-year-old technical wizard who plays the notes but misses the meaning.

Most of us turn to music for emotional experience. We do not audit performances for wrong notes, and unless they break our reverie we do not notice them. Levitin argues that research on expertise has looked in the wrong place, at finger facility rather than emotional expression. He asked the dean of a top North American music school when emotion and expressivity are taught. She said they are not: the approved curriculum (repertoire, ensemble and solo training, sight singing, sight reading, theory) leaves no time. Expressive musicians usually arrive knowing how to move listeners or work it out themselves. Occasionally, for an exceptional student who already solos with the school orchestra, there is time in the final weeks of the last semester to coach emotion, which she said almost in a whisper. So at one of the best schools, music's central purpose is taught to a few, in the last weeks of a four- or five-year program.

Even uptight, analytic people expect to be moved by Shakespeare and Bach. We admire the craft, but facility must serve a different kind of communication. Jazz fans are especially demanding of post-big-band heroes from the Miles Davis, John Coltrane and Bill Evans era, and call lesser, emotionally detached players "shucking and jiving," pleasing the audience through obsequiousness rather than soul.

Why are some musicians better emotionally? Nobody knows. Musicians have not yet performed with feeling inside brain scanners, because scanners require total stillness, though that may change within about five years. Interviews and diaries from musicians such as Beethoven, Tchaikovsky, Rubinstein, Bernstein, B. B. King and Stevie Wonder suggest that conveying emotion involves partly technical, mechanical factors and partly something mysterious.

Brendel says he does not think about notes onstage but about creating an experience. Stevie Wonder told Levitin in 1996 that when performing he tries to return to the frame of mind and "frame of heart" he had when writing the song, which helps his delivery. How this changes his singing or playing is unknown, but neuroscientifically it makes sense. Remembering music means returning the neurons active during the original perception to their earlier state, with the same connectivity pattern and firing rates as closely as possible. This recruits neurons in the hippocampus, amygdala and temporal lobes, orchestrated by attention and planning centers in the frontal lobe.

The neuroanatomist Andrew Arthur Abbie speculated in 1934 about a link between movement, brain and music, now being confirmed. He wrote that pathways from the brain stem and cerebellum to the frontal lobes can weave sensory experience and coordinated muscle movement into a homogeneous fabric, yielding man's highest powers as expressed in art, a pathway devoted to movements that carry creative purpose. Studies by Marcelo Wanderley of McGill and Levitin's former doctoral student Bradley Vines (now at Harvard) show that nonmusician listeners are very sensitive to musicians' gestures. Watching a performance with the sound off and attending to arm, shoulder and torso movement, ordinary listeners detect much of the musician's expressive intention. Adding sound produces an emergent understanding beyond what sound or image gives alone.

If music conveys feeling through gesture and sound, the musician's brain state needs to match the emotion expressed. Levitin bets that B.B. King's neural signatures while playing the blues resemble those while feeling the blues, though scientists would have to subtract the motor-command and listening processes versus sitting with head in hands. Listeners probably match some of the performers' brain states too. As a recurring theme, even those without explicit training have musical brains and are expert listeners.

Charisma, celebrity and luck

Musical expertise takes many forms, technical and emotional. Drawing us into a performance until we forget everything else is a special ability too. Many performers have a personal magnetism independent of other abilities. We cannot take our ears off Sting, Miles Davis or Eric Clapton, and it is not about the notes, which many good musicians could produce, perhaps with better technique. Record executives call it star quality. A model who has it is photogenic, and Levitin calls the musical counterpart on records phonogenic.

Celebrity and expertise must be distinguished, and their causes may be unrelated. Neil Young told Levitin he does not consider himself especially talented but one of the lucky ones who became commercially successful. Few get a major-label deal and fewer sustain decades-long careers. Young, Stevie Wonder and Eric Clapton all credit much of their success to a lucky break. Paul Simon said he has been lucky to work with amazing musicians, most of whom nobody has heard of.

Untrained musicians and Joni Mitchell

Francis Crick turned his lack of training into an advantage, free of scientific dogma to open his mind. An artist who brings that freedom to music can produce astonishing results. Many great musicians lacked formal training: Sinatra, Louis Armstrong, John Coltrane, Eric Clapton, Eddie Van Halen, Stevie Wonder and Joni Mitchell; in classical music George Gershwin, Mussorgsky and David Helfgott; and Beethoven judged his own training poor in his diaries.

Joni Mitchell sang in school choirs but never took guitar or other lessons. Her music is variously called avant-garde, ethereal, and a bridge between classical, folk, jazz and rock. She uses many alternate tunings, setting the strings to pitches of her choosing. She does not play notes others cannot, since the chromatic scale has twelve notes, but she can fingerwise reach combinations other guitarists cannot, whatever their hand size. A more important effect concerns sound production. Each string has a tuned pitch, and pressing it against the neck shortens it so it vibrates faster and sounds higher. A fretted string sounds slightly deadened by the finger, while open strings ring clearer and longer. When several open strings ring together a distinctive timbre emerges. By retuning, Joni changes which notes sound when strings are open, so we hear ringing notes and combinations not usually heard, as in "Chelsea Morning" and "Refuge of the Roads."

There is more, since guitarists such as David Crosby, Ry Cooder, Leo Kottke and Jimmy Page also use own tunings. Over dinner in Los Angeles Joni talked about bass players she has worked with: Jaco Pastorius, Max Bennett, Larry Klein, and Charles Mingus, with whom she wrote a whole album. She speaks compellingly for hours about tunings, comparing them to van Gogh's colors. She told how Jaco argued with her and caused mayhem backstage. When Roland hand-delivered the first Roland Jazz Chorus amplifier for her to use, Jaco moved it to his corner of the stage, growled that it was his, and gave her a fierce look. Levitin, a fan of Jaco from Weather Report, interrupted to ask what playing with him was like musically. She said he was the only bass player up to then who really understood what she was trying to do, which is why she tolerated his aggression.

She also recounted how, when she started, the record company wanted to assign a hit-making producer, but Crosby warned that a producer would ruin her and proposed putting his own name on the record as producer so the company would trust him and stay out of her way. Then the session musicians all had ideas, the worst being bass players, who always asked for the root of the chord. In music theory, the root is the note a chord is named for and built around: C is the root of C major, and E-flat of E-flat minor. Her chords, because of her composing and playing, are not typical and cannot easily be labeled. She told bassists to play whatever sounded good, but they said they had to play the root or it would not sound right.

Since Joni had no theory and could not read music, she could not name roots. She had to tell the bassists each guitar note, and they worked out chords painstakingly one at a time. Here psychoacoustics and theory collide. Standard chords like C major, E-flat minor and D7 are unambiguous, with one possible root. Joni's genius is making ambiguous chords that could have two or more roots. With no bass, as in "Chelsea Morning" or "Sweet Bird," the listener is left with open aesthetic possibilities, because expectations about what follows are less certain than with traditional chords. Strings of such chords multiply harmonic complexity, since each sequence can be interpreted dozens of ways depending on how each member is heard. Because listeners hold what they just heard in immediate memory and integrate it with new input, even nonmusicians can revise many interpretations as the piece unfolds, and each new listening brings new contexts, expectations and interpretations. Levitin calls it as close to impressionist visual art as anything he has heard.

A bass note fixes one interpretation and ruins the delicate ambiguity. Every bassist before Jaco insisted on roots or what they took for roots. Jaco, Joni said, instinctively wandered the possibility space, giving the different interpretations equal emphasis and holding the ambiguity in suspended balance, so her songs could have bass without losing their expansiveness. This, they worked out at dinner, is one reason her music sounds unlike anyone else's: the harmonic complexity comes from her insistence that it not be anchored to one interpretation. Add her phonogenic voice and the listener enters a unique soundscape.

Musical memory

Memory is another aspect of expertise. Some people remember details others cannot, like a friend who remembers every joke while others cannot retell one heard that day. Richard Parncutt, musicologist and music cognition professor at the University of Graz, once played piano in a tavern to fund graduate school. When he visits Levitin in Montreal he plays at the living room piano while Levitin sings, and can play any song named from memory, including different versions: asked for "Anything Goes," he asks whether Levitin wants Sinatra's, Ella Fitzgerald's or Count Basie's. Levitin can play or sing about a hundred songs from memory, which is typical for people who have played in bands or orchestras, but Parncutt seems to know thousands, chords and lyrics. How does he do it, and can ordinary people?

At Berklee College of Music in Boston, Levitin met Carla, who could recognize and name a piece within three or four seconds. He does not know how well she could sing from memory, since they spent their time trying to stump her with melodies, which was hard. She later worked at ASCAP, the composers' rights organization, which monitors radio playlists to collect royalties for members. Staff sit in a Manhattan room listening to excerpts from stations nationwide and must name song and performer within three to five seconds, then log it and move on; that is also a hiring requirement.

A third example is Kenny, the clarinet-playing boy with Williams syndrome from earlier. Playing Scott Joplin's "The Entertainer" he struggled with a passage and asked, with the eagerness to please typical of Williams syndrome, to try again, but went back to the very beginning rather than a few notes. Levitin has seen master musicians from Carlos Santana to the Clash in studios restart from the beginning of a phrase if not the whole piece, as if executing a memorized sequence of muscle movements that must start from its start.

What do these have in common, and how do they differ from ordinary musical memory? Expertise in any domain brings superior memory only within that domain. Parncutt still loses his keys. Chess grandmasters have memorized thousands of positions, but their memory covers only legal positions; given random arrangements they do no better than novices, so their knowledge is schematized around legal moves. Likewise, expert musicians excel at remembering chord sequences that make sense within the harmonic systems they know, but do no better than others on random chords.

Musicians therefore memorize by relying on a structure into which details fit. This is efficient: instead of every note, they build a framework that can hold many pieces. Learning Beethoven's "Pathétique" Sonata, a pianist learns the first eight measures and then needs only to know that the next eight repeat the theme an octave higher. Any rock musician can play the Beatles' "One After 909" unseen if told it is a standard sixteen-bar blues progression, a framework that fits thousands of songs, with the song's nuances as variations. Musicians past a certain level do not learn pieces note by note but scaffold on earlier pieces and note variations from the standard schema.

Memory for playing thus works much like listening, as in Chapter 4, through standard schemas and expectations. Musicians also use chunking, tying units together and remembering the group, as chess players and athletes do. An everyday example is remembering a long-distance phone number: for a New York number, someone who knows other NYC numbers stores the area code as the single unit 212, and might know Los Angeles is 213, Atlanta 404, and England's country code is 44. Chunking matters because working memory, the contents of present awareness, is severely limited, generally to about nine items, whereas long-term memory has no known practical limit. Treating the area code as one unit plus seven digits helps avoid the limit. Chess players remember boards as groups of pieces in standard, nameable patterns.

Musicians chunk in three ways:

These are what Parncutt uses at the piano. He also knows enough theory and style to bluff through passages he does not truly know, like an actor substituting words when a line is forgotten, replacing an uncertain note or chord with a stylistically plausible one.

Identification memory, the ability most of us have to recognize music heard before, resembles memory for faces, photos, tastes and smells. It varies between individuals and is domain specific: some, like Carla, are especially good with music, others in other senses. Rapidly retrieving a familiar piece is one skill, but quickly attaching a label such as title, artist and year, which Carla did, involves a separate cortical network. It is now thought to involve the planum temporale (associated with absolute pitch) and parts of the inferior prefrontal cortex needed to attach verbal labels to sensory impressions. Why some do this better is unknown, but it may reflect an innate, hardwired predisposition in brain formation, possibly with a partial genetic basis.

When learning note sequences in a new piece, musicians sometimes resort to brute-force repetition, as most of us did as children with the alphabet, the U.S. Pledge of Allegiance or the Lord's Prayer. Rote learning is helped by hierarchical organization, because some words or notes are structurally more important (as in Chapter 4), and learning is organized around them. This plain memorization is what musicians do to learn the muscle movements for a piece, and it is part of why musicians like Kenny cannot start on any note but go to the beginnings of meaningful, hierarchically organized chunks.

Conclusion

Being an expert musician takes many forms: instrumental dexterity, emotional communication, creativity and special mental structures for memory. Being an expert listener, as most of us are by age six, means having absorbed our culture's musical grammar into schemas that produce expectations, the heart of aesthetic experience. How these forms of expertise are acquired remains a neuroscientific mystery. The emerging consensus is that musical expertise is many components, not one, and experts need not have them equally; Irving Berlin lacked even a basic skill, playing an instrument well. Musical expertise is unlikely to differ wholly from expertise elsewhere. It uses some brain structures and circuits other activities do not, but becoming a musical expert, composer or performer, needs the personality traits of other domains: diligence, patience, motivation and stick-to-it-iveness.

Becoming a famous musician is a different matter, which may depend less on intrinsic ability than on charisma, opportunity and luck. Levitin repeats that we are all expert musical listeners, making subtle judgments of what we like even when we cannot say why. Science has something to say about why we like the music we do, and that is the subject of the next chapter, which he introduces as another interesting facet of neurons and notes.


Chapter 8: My Favorite Things — Why Do We Like the Music We Like?

Hearing music before birth

Levitin opens by asking the reader to imagine waking inside the womb: darkness, a distant pulse, and then muffled, rhythmic, wavering sounds that feel familiar and let you vaguely anticipate what comes next, as if heard underwater. A fetus hears its mother's heartbeat, which speeds up and slows down, and it also hears music. The auditory system is working about twenty weeks after conception.

Alexandra Lamont of Keele University in the UK found that a year after birth, children recognise and prefer music they were exposed to in the womb. Mothers played one chosen piece repeatedly during the last three months of pregnancy. The babies also heard everything else in their mothers' lives, filtered through the fluid, such as other music, conversation and ambient noise. The chosen pieces covered several styles:

After birth, the mothers were not allowed to play the piece to their infants. At one year, each baby heard the womb piece alongside a new piece matched for style and tempo. For example, UB40's "Many Rivers to Cross" was paired with Freddie McGregor's "Stop Loving You."

To learn which of two sounds a baby who cannot yet talk prefers, researchers use the conditioned head-turning procedure. Robert Fantz developed it in the 1960s, and John Columbo, Anne Fernald, the late Peter Jusczyk and colleagues refined it. The infant sits, often on a parent's lap, between two loudspeakers. Looking at one speaker starts one sound and looking at the other starts a different sound, so the baby soon learns it controls what plays. Experimenters counterbalance which speaker carries which stimulus, so each comes from either side half the time. Lamont's infants looked longer toward the speaker playing the womb music. A control group of one-year-olds with no prior exposure showed no preference, so the music itself was not what drove the result. She also found that, other things equal, young infants prefer fast, upbeat music to slow music.

Childhood amnesia and the "Mozart Effect"

These results challenge the traditional idea of childhood amnesia, the belief that we have no true memories before about age five. People often claim memories from ages two or three, but these may be memories of being told about an event. The young brain is unfinished: functional specialisation is incomplete, pathways are still forming, and the child cannot yet separate important events from unimportant ones or encode experience systematically. Such a child is very open to suggestion and may take in stories about himself as his own memories. For music, though, even prenatal experience seems to be stored and retrievable without language or conscious awareness.

Levitin then turns to the "Mozart Effect," a study that made the news some years ago. It claimed that ten minutes of Mozart a day made you smarter, specifically that it improved spatial-reasoning performance right after listening, which some journalists took to imply mathematical ability. Politicians passed resolutions, and Georgia's governor funded a Mozart CD for every newborn in the state. Most scientists were uncomfortable. They believe music can help other cognitive skills and want more school music funding, but the study had many flaws and was right for the wrong reasons. Levitin also found the fuss offensive, because it implied music is worth studying only for its side benefits. He says it would sound absurd if someone argued that maths should be funded because it helps music. Music is often the first school program cut, and people justify it by collateral benefits instead of letting it stand on its own rewards.

The flaw was inadequate controls. Research by Bill Thompson, Glenn Schellenberg and others showed the small spatial advantage depended on the control task. Music beat sitting in silence, but there was no advantage over any mild mental stimulation, such as listening to a book on tape or reading. The study also offered no plausible mechanism for how listening could raise spatial performance.

Schellenberg stresses the difference between short-term and long-term effects. The Mozart Effect concerned immediate benefits, but musical activity has long-term effects. Music changes certain neural circuits, including the density of dendritic connections in primary auditory cortex. Gottfried Schlaug of Harvard showed that the front part of the corpus callosum, the fibre bundle linking the two hemispheres, is significantly larger in musicians than non-musicians, especially those who started training early. This supports the idea that musical operations become bilateral with training. Studies also find microstructural changes in the cerebellum after motor-skill learning, including more and denser synapses. Schlaug found musicians tend to have larger cerebellums and more gray matter than non-musicians. Gray matter holds cell bodies, axons and dendrites and handles information processing, whereas white matter handles transmission. Whether these changes improve non-musical abilities is unproven, though music listening and music therapy do help people with a wide range of psychological and physical problems.

Seeds of taste: the womb, acculturation and consonance

Returning to taste, Levitin says Lamont's results show that prenatal and newborn brains store memories and retrieve them over long periods. The environment, even through fluid and the womb, can shape development and preference. But preferences are influenced, not determined, by prenatal exposure, or children would simply favour their mothers' music or whatever plays in Lamaze classes. There is also a long period of acculturation in which the child absorbs the music of her culture. Reports that all infants prefer Western music before they get used to foreign music were not corroborated. What was found is that infants prefer consonance to dissonance. Appreciating dissonance comes later, and people differ in how much they tolerate.

There is probably a neural basis. Consonant and dissonant intervals are handled by separate mechanisms in auditory cortex. Recordings from humans and monkeys exposed to sensory dissonance, meaning dissonance arising from frequency ratios and not from musical context, show that neurons in primary auditory cortex synchronise their firing during dissonant chords but not consonant ones. Why this produces a preference for consonance is unclear.

What infants can process

Infant ears work four months before birth, but the brain needs months or years to reach full auditory processing. Infants recognise transpositions in pitch and in tempo, which means they process relations, something computers still do poorly. Jenny Saffran of the University of Wisconsin and Laurel Trainor of McMaster University found that infants can also use absolute-pitch cues when a task demands it. This shows a flexibility not previously known, with different processing modes, presumably in different circuits, chosen to suit the problem.

Trehub, Dowling and others showed that contour is the most salient musical feature for infants, who detect contour similarities and differences across thirty seconds of retention. Contour is the pattern of ups and downs in a melody regardless of interval size, so someone attending only to contour would encode that a melody rises but not by how much. This parallels sensitivity to linguistic contour, which distinguishes questions from exclamations and belongs to prosody. Fernald and Trehub documented that, across cultures, parents speak to infants more slowly, with a wider pitch range and a higher overall pitch. This is infant-directed speech, or motherese. Mothers, and less so fathers, do it naturally. It seems to draw the baby's attention to the mother's voice and help pick out words.

Levitin gives an example. Instead of saying "This is a ball," a mother produces drawn-out, rising "Seeeee?" and then "See the BAAAALLLL?" with the pitch rising again on ball. The contour signals a question or a statement, and exaggerating it makes two easily distinguished prototypes, one for questions and one for declarations. A scolding creates a third prototype that is short, clipped and flat in pitch, such as "No! Bad!" Babies seem hardwired to track contour in preference to exact pitch intervals.

Trehub also showed that infants encode consonant intervals such as the perfect fourth and fifth more easily than dissonant ones like the tritone. She and colleagues also tested whether the unequal steps of our scale help. Nine-month-olds heard the ordinary seven-note major scale and two invented scales. One divided the octave into eleven equal steps and picked seven tones forming one- and two-step patterns. The other divided the octave into seven equal steps. The task was to detect a mistuned tone. Adults did well on the major scale but poorly on both artificial scales. Infants did equally well on the unequally spaced scales and the equally spaced one. Nine-month-olds are believed not to have a major-scale schema yet, which suggests a general processing advantage for unequal steps, which our major scale has.

Levitin concludes that brains and scales seem to have coevolved. The lopsided major scale is no accident, since melodies are easier to learn with it. It follows from the physics of sound production, the overtone series discussed earlier, because the major-scale tones lie very close to overtone-series tones. Early in childhood, children spontaneously vocalise in ways that can sound like singing, exploring their vocal range and phonetic production in response to what they hear. The more music they hear, the more pitch and rhythm variation appears in their vocalisations.

Childhood development of taste

By age two, children show a preference for their own culture's music, around the time specialised speech processing begins. At first they like simple songs, meaning clear themes (not four-part counterpoint) and chord progressions that resolve directly and predictably. As they mature they tire of predictability and look for challenge.

Mike Posner explains part of this. The frontal lobes and the anterior cingulate, a structure just behind the frontal lobes that directs attention, are not fully formed in children. Children therefore struggle to attend to several things at once and to focus on one stimulus amid distractors. This is why children under about eight have trouble singing rounds like "Row, Row, Row Your Boat." The network linking the cingulate gyrus (the larger structure containing the anterior cingulate) and orbitofrontal regions cannot yet filter out distraction. They face a barrage of sound and are tripped up by the other parts. Posner showed that exercises adapted from NASA attention games can speed up attentional development.

The move from simple to complex is a generalisation. Not all children like music, and some develop unusual tastes by serendipity. Levitin became fascinated with big band and swing at eight, when his grandfather gave him a collection of World War II era 78 rpm records. He was first drawn to novelty songs made for children, including "The Syncopated Clock," "Would You Like to Swing on a Star," "The Teddy Bear's Picnic" and "Bibbidy Bobbidy Boo." Enough exposure to the unusual chords and voicings of the Frank de Vol and Leroy Anderson orchestras became part of his mental wiring. He soon listened to all kinds of jazz, and the children's jazz opened the way for jazz in general.

The teenage years and the critical period

Researchers see the teen years as the turning point. Around ten or eleven, most children take up music as a real interest, even if they did not before. As adults, the music we feel nostalgic for, our "own" music, matches what we heard in these years. Memory loss is an early sign of Alzheimer's disease, which involves changes in nerve cells and neurotransmitters and destruction of synapses, and it worsens over time. Yet many patients can still sing songs they heard at fourteen. Levitin gives two reasons. Those years involved self-discovery and so were emotionally charged, and we remember emotional events because the amygdala and neurotransmitters work together to tag them as important. Also, around fourteen the wiring of the musical brain nears adult completion through neural maturation and pruning.

There is no cutoff for acquiring new tastes, but most people have formed theirs by eighteen or twenty. This is not fully understood, but several studies agree. One reason may be that people grow less open to new experience as they age. In adolescence we discover other ideas, cultures and people and experiment with not being limited to what our parents taught. We also seek out different music. In Western culture, music choice has social consequences. We listen to what our friends listen to and bond with people we want to resemble, externalising the bond by dressing alike, sharing activities and sharing music. This ties to the evolutionary view of music as social bonding and cohesion. Musical preference becomes a mark of personal and group identity and distinction.

Personality may be associated with, and somewhat predictive of, taste, but chance matters a great deal, such as where you went to school, who your friends were and what they happened to play. As a child in northern California, Levitin found Creedence Clearwater Revival huge, since they came from nearby. After he moved to southern California, CCR's country-flavoured sound did not fit the surfer and Hollywood culture that favoured the Beach Boys and theatrical acts like David Bowie.

Brain development also matters. New connections form explosively through adolescence and slow greatly afterward, so this is the formative phase when circuits are structured by experience. New music is assimilated into the framework built then. Critical periods exist for skills such as language. A child who has not learned a language by about six, first or second, will never speak it with a native speaker's ease. Music and mathematics have a longer window, but not an unlimited one. Someone with no music lessons or maths training before about twenty can still learn them but only with great difficulty and will probably never "speak" them like an early learner. The reason is the biological course of synaptic growth. Synapses are programmed to grow and form connections for some years, then pruning removes unneeded ones.

Neuroplasticity is the brain's ability to reorganise itself. Recent demonstrations of reorganisation once thought impossible exist, but adults can reorganise far less than children and adolescents. Individuals differ, as some people heal broken bones faster than others, and some forge new connections more easily. Between about eight and fourteen, pruning begins in the frontal lobes, the seat of reasoning, planning and impulse control. Myelination ramps up in the same period. Myelin is a fatty coating on axons that speeds transmission, which is why older children solve problems faster and tackle harder ones. Myelination of the whole brain is usually done by twenty. Multiple sclerosis is one of several degenerative diseases that attack the myelin sheath.

Simplicity, complexity and schemas

The balance between simplicity and complexity also shapes taste. Studies across painting, poetry, dance and music show an orderly relationship between a work's complexity and how much we like it. Complexity is subjective: what is impenetrable to one person, Levitin's hypothetical Stanley, may fall in the sweet spot for another, Oliver. What one finds hideously simple, another may find hard, depending on background, experience, understanding and cognitive schemas.

Schemas, Levitin says, are everything. They frame understanding and are the system into which we place an aesthetic object's elements and interpretations, shaping models and expectations. He illustrates with Mahler's Fifth. With one schema it can be understood even on first hearing, since it is a symphony in four movements with a main theme, subthemes and repetitions, played by orchestral instruments and not African talking drums or fuzz bass. People who know Mahler's Fourth recognise that the Fifth opens with a variation on that theme, at the same pitch. Those who know Mahler well know he quotes three of his own songs. Musically educated listeners know that symphonies from Haydn to Brahms and Bruckner usually begin and end in the same key, whereas Mahler moves from C-sharp minor to A minor and ends in D major. Someone who cannot keep the key in mind or lacks a sense of a symphony's normal trajectory would find this meaningless. For the seasoned listener, skilful violation of convention is a rewarding surprise. Someone lacking the symphonic schema, or holding another one such as that of an Indian raga fan, may find the work rambling, with ideas melting into each other without boundaries. The schema frames perception, processing and experience.

Music that is too simple strikes us as trivial, and music that is too complex strikes us as unpredictable and ungrounded in anything familiar. Art needs the right balance, and since simplicity and complexity relate to familiarity, and familiarity is another word for schema, taste comes down to schemas.

Levitin offers an operational definition: a piece is too simple when it is trivially predictable, like something already experienced, with no challenge. He uses tic-tac-toe as an analogy. Young children find it fascinating because it has clear rules they can state, an element of surprise since you cannot be sure what the opponent will do, dynamic interaction in which your move depends on theirs, and an uncertain ending (who wins, or a draw) within an outer limit of nine moves. That uncertainty creates tension and expectation, released when the game ends. As children grow they learn strategy, such as that the second player cannot beat a competent opponent and can only draw. Once the sequence and end point are predictable the game loses appeal. Adults can still play with children, but they enjoy the child's pleasure and the multi-year process of the child's brain unlocking the game.

For many adults Raffi and Barney the Dinosaur are the musical tic-tac-toe. When the outcome is too certain and each move from note to note or chord to chord holds no surprise, music seems unchallenging. While listening, especially with focus, the brain thinks ahead about possible next notes, the trajectory, direction and end point. The composer has to lull us into trust and security. We let him take us on a harmonic journey, and he must give enough small rewards, completed expectations, that we feel order and a sense of place.

The hitchhiking analogy and overly complex music

Levitin compares this to hitchhiking from Davis, California, to San Francisco. You want the driver to take the usual route, Highway 80. You may accept shortcuts if the driver is friendly and explains them, for example cutting over Zamora Road to avoid construction. If the driver takes unexplained back roads and you lose all landmarks, your sense of safety is violated. People differ in reacting to unexpected journeys, musical or otherwise. Some panic ("That Stravinsky is going to kill me!") and some feel adventure, as in thinking Coltrane is doing something odd but one can stay and find the way back to musical reality if needed.

Games continue the analogy. Some have rules so complicated that the average person lacks patience, with too many or too unpredictable possibilities per turn for a novice. But unpredictability does not always mean a game will repay persistence. Some games stay unpredictable however much you play, such as dice-driven board games like Chutes and Ladders and Candy Land. Children enjoy the surprise, but adults find them tedious, because the outcome has no structure and no skill can influence it.

Music with too many chord changes or unfamiliar structure sends listeners to the exit or the skip button. Games such as Go, Axiom and Zendo are opaque enough that many novices quit early, since the learning curve is steep and the payoff uncertain. Unfamiliar music is similar. People may tell you Schoenberg is brilliant or Tricky is the next Prince, but if you cannot work out what is going on in the first minute, you wonder whether the effort will pay off. We tell ourselves repeated listening will bring understanding, yet we also recall artists we spent hours on without ever "getting it." Appreciating new music is like a new friendship in that it takes time and sometimes cannot be hurried. Neurally, we need a few landmarks to invoke a schema. With enough hearings of radically new music, some of it becomes encoded and landmarks form. If the composer is skilled, those landmarks will be the ones intended, as his knowledge of composition, perception and memory lets him build hooks that eventually stand out.

Structure, form and jazz

Structural processing is one source of difficulty. Not understanding symphonic form, sonata form or the AABA structure of a jazz standard is like driving a highway without signs, so you never know where you are or when you will arrive. Many people say they do not "get" jazz because it sounds formless, a contest to cram in as many notes as possible. Levitin notes more than half a dozen subgenres under the label: Dixieland, boogie-woogie, big band, swing, bebop, "straight-ahead," acid-jazz, fusion, metaphysical and so on. Straight-ahead, or classic, jazz is the standard form, comparable to the sonata or symphony in classical music, or to a typical Beatles, Billy Joel or Temptations song in rock.

In classic jazz the artist begins by playing the main theme, often a well-known Broadway tune or an earlier hit. Such songs are called standards, for instance "As Time Goes By," "My Funny Valentine" and "All of Me." The artist plays through the full form once, typically two verses and the chorus (also called the refrain), then another verse. The chorus is the part that repeats throughout, while the verses change. This is called AABA, with A for verse and B for chorus, so the order is verse, verse, chorus, verse. Variations exist, and some songs add a C section called the bridge.

"Chorus" also means one run through the entire form, so playing through AABA once is "playing one chorus." Levitin explains that when he plays jazz, "play the chorus" or "go over the chorus" means the section, while "run through one chorus" or "do a couple of choruses" means the whole form.

"Blue Moon" (Frank Sinatra, Billie Holiday) has AABA form. A jazz artist may change the rhythm or feel and embellish the melody. After one pass through the form, which jazz players call the head, band members take turns improvising over the original chord progression and form. Each plays one or more choruses, then the next takes over at the start of the head. Some stay near the original melody and others add distant, exotic harmonic departures. When all have soloed, the band returns to the head, usually played straight, and finishes. Improvisation can run many minutes, so a two- or three-minute song may stretch to ten or fifteen. There is a typical order: horns first, then piano and/or guitar, then bass, and sometimes the drummer, after the bass. Sometimes players split a chorus, each taking four or eight measures and handing off, like a relay race.

To a newcomer it may seem chaotic, but simply knowing that improvisation follows the original chords and form helps orient the listener. Levitin advises new jazz listeners to hum the main tune mentally once improvisation starts, as the improvisers themselves often do, which enriches the experience.

Every genre has its own rules and form, and the more we listen, the more those rules are stored in memory. Unfamiliarity with structure brings frustration or lack of appreciation. Knowing a genre means having a category for it and being able to classify new songs as members, non-members or partial or "fuzzy" members with exceptions.

The inverted-U function

The orderly relation between complexity and liking is called the inverted-U function. Imagine a graph with complexity (to you) on the x-axis and liking on the y-axis. Near the origin sits very simple music that you dislike. As complexity increases, liking rises, until you cross a personal threshold and move from disliking to liking quite a bit. Past some point the music becomes too complex and liking falls, until you cross another threshold and dislike it entirely. The curve forms an inverted U or V.

The hypothesis does not claim complexity is the only reason for liking or disliking, only that it accounts for that variable. Musical elements themselves can be barriers. Music that is too loud or too soft is a problem, but even dynamic range, the gap between loudest and softest parts, can cause rejection. This is especially true for people who use music to regulate mood, whether calming down or being pepped up for a workout. They will not want a piece spanning very soft to very loud, or sad to exhilarating, as Mahler's Fifth does. The dynamic and emotional range is too wide and creates a barrier to entry.

Pitch matters too. Some people cannot stand the low thumping of modern hip-hop, and others dislike the high-pitched whininess of violins. Part of this may be physiological, since different ears may transmit different parts of the spectrum so some sounds seem pleasant and others aversive. Psychological associations with instruments, positive or negative, may also exist.

Rhythm influences appreciation as well. Many musicians love Latin music for its rhythmic complexity. To an outsider it sounds simply "Latin," but to someone who can hear which beats are strong, it is a world of distinct styles: bossa nova, samba, rhumba, beguine, mambo, merengue and tango. Some people enjoy it without telling the styles apart, while others find the rhythms too complicated and unpredictable. Levitin finds that teaching a listener one or two Latin rhythms leads to appreciation, a matter of grounding and having a schema. For others rhythms that are too simple are the dealbreaker. His parents' generation complained that rock and roll, besides being loud, all had the same beat.

Timbre is another barrier, and its influence is likely growing, as Chapter 1 argued. When Levitin first heard John Lennon or Donald Fagen sing, he found the voices unimaginably strange and did not want to like them. Something, perhaps the strangeness, kept him returning, and they became two of his favourite voices, now beyond familiar toward intimate, as though incorporated into who he is, and neurally they have been. After thousands of hours and tens of thousands of playings, his brain has circuitry that picks out their voices from thousands of others, even in unfamiliar songs. It has encoded each vocal nuance and timbral flourish, so on hearing an alternate version, such as the demos on the John Lennon Collection, he can immediately tell how it deviates from the version stored in long-term memory.

Past experience, safety and vulnerability

Musical taste, like other preferences, depends on earlier experiences and whether they were positive or negative. If pumpkin once made you ill, you will be wary of it, but a few mostly positive broccoli experiences might make you try broccoli soup. One positive experience leads to others. Likewise, the sounds, rhythms and textures we like are generally extensions of earlier pleasant musical experiences. Hearing a liked song resembles other pleasant sensory experiences such as chocolate, fresh raspberries, morning coffee, a work of art or a loved one's sleeping face. We enjoy the sensation and find comfort in familiarity and the safety it brings. Looking at and smelling a ripe raspberry, Levitin can expect it to taste good and be safe. A first loganberry shares enough features with a raspberry that he can risk eating it.

Safety matters in choosing music. To an extent we surrender to music, trusting composers and performers with part of our hearts and spirits and letting the music take us outside ourselves. Many feel great music connects us to something bigger, to other people or to God. Even when not transcendent, music changes mood. So we are reluctant to drop emotional defences for just anyone, and will do so only if composers and musicians make us feel safe and our vulnerability will not be exploited.

This is part of why many cannot listen to Wagner. Because of his pernicious anti-Semitism, the vulgarity of his mind (Oliver Sacks's description) and the music's association with the Nazi regime, some do not feel safe hearing it. Levitin says Wagner has always disturbed him, as has even the idea of listening. He is reluctant to be seduced by music from so disturbed a mind and so dangerous or impenetrably hard a heart, fearing he might pick up some of the same ugly thoughts. Listening to a great composer feels like becoming one with him or letting part of him inside. He finds this troubling with pop music too, since some of its purveyors are crude, sexist, racist, or all three.

This vulnerability and surrender is strongest with rock and popular music of the past forty years, which explains the fandom around the Grateful Dead, Dave Matthews Band, Phish, Neil Young, Joni Mitchell, the Beatles, R.E.M. and Ani DiFranco. We let them control our emotions and even our politics, lifting, lowering, comforting and inspiring us. We let them into living rooms and bedrooms when we are alone, and directly into our ears through earbuds and headphones when we are not communicating with anyone else.

Such openness to a total stranger is unusual. We normally have protection against blurting every thought, so when asked "How're ya doin'?" we say "Fine" even if depressed after a fight at home or suffering a minor ailment. Levitin's grandfather said a bore is someone who answers that question truthfully. Even with close friends we hide things like digestive or bowel problems and self-doubt. One reason we open up to favourite musicians is that they often make themselves vulnerable to us, or convey vulnerability through their art. Levitin says whether it is real or represented does not matter for now.

Connecting through art, and the artist's image

Art connects us to each other and to larger truths about being alive and human. When Neil Young sings of an old man looking at his life and living alone in a paradise that makes him think of two, we feel for the writer. Levitin may not live in paradise but can empathise with someone who has material success and no one to share it with, who feels he has "gained the world but lost his soul," as George Harrison sang, echoing both the gospel of Mark and Gandhi.

When Bruce Springsteen sings "Back in Your Arms" about lost love, we resonate with a similar theme from a poet with an "everyman" persona like Young's. Given Springsteen's adoration by millions and his millions of dollars, it is more tragic that he cannot have the one woman he wants.

We hear vulnerability in unlikely places. David Byrne of the Talking Heads is known for abstract, arty, somewhat cerebral lyrics, yet in his solo "Lilies of the Valley" he sings about being alone and scared. Our appreciation is heightened by knowing the artist, or at least his persona as an eccentric intellectual who rarely shows anything so raw.

Connections to the artist, or to what he stands for, can thus shape preference. Johnny Cash cultivated an outlaw image and showed compassion for prisoners by playing many prison concerts. Prisoners may like or grow to like his music for what he represents, apart from purely musical reasons. But fans only go so far in following heroes, as Dylan learned at the Newport Folk Festival. Cash could sing about wanting to leave prison without alienating his audience. Had he said he liked visiting prisons because it helped him appreciate his own freedom, he would have crossed from compassion into gloating, and the inmates would rightly have turned on him.

Exposure, adventurousness and the early foundation

Preferences begin with exposure, and each of us has an "adventuresomeness" quotient for how far from our musical safety zone we will go at a given time. Some people are more open to experiment in all areas of life, music included, and the same person may seek or avoid it at different times. Boredom is generally when we look for new experiences. With Internet radio and personal players spreading, Levitin predicts personalised stations in the next few years, where algorithms play a mix of music we know and like and music we do not know but probably will enjoy. He thinks any such technology should include an "adventuresomeness" knob controlling the mix of old and new, or how far the new music lies from usual listening. This varies between people and even within one person across the day.

Listening builds schemas for genres and forms even when passive and unanalytical. By an early age we know the "legal moves" in our culture's music, and for many people later likes and dislikes follow from the schemas formed by childhood listening. This does not mean childhood music fixes taste for life, since many people study or are exposed to other cultures' music and learn their schemas too. The point is that early exposure is often the most profound and forms the foundation for later understanding.

Finally, musical preferences have a large social component, based on knowledge of the performer, of what family and friends like, and of what the music stands for. Historically, and especially in evolutionary terms, music has been tied to social activity. This may explain why the most common form of musical expression, from the Psalms of David to Tin Pan Alley to today, is the love song, and why love songs are, for most of us, among our favourite things.


Chapter 9: The Music Instinct: Evolution's #1 Hit

Where music came from, and Pinker's challenge

The question of music's evolutionary origin goes back to Darwin, who thought it arose through natural selection as part of mating rituals of humans or their predecessors. Levitin thinks the evidence supports this, though not everyone agrees. Work on the topic was scattered until 1997, when Steven Pinker issued a challenge. About 250 people worldwide study music perception and cognition as their main focus, and they meet once a year. In 1997 the meeting was at MIT and Pinker, who had just finished How the Mind Works but was not yet famous, gave the opening address.

Pinker said language is plainly an adaptation, and that cognitive mechanisms such as memory, attention, categorization and decision making all have clear evolutionary purposes. Sometimes, though, a trait has no evolutionary basis of its own. Evolution propagates an adaptation for one reason and something else comes along for the ride. Stephen Jay Gould called such a by-product a spandrel, borrowing from architecture: the space between the arches that support a dome is unplanned, yet artists fill it with angels and decoration. Feathers evolved for warmth and were co-opted for flight. Many spandrels are used so well that it is hard to tell afterward whether they were adaptations.

Pinker argued that language is the adaptation and music is its spandrel, the least interesting thing for a cognitive scientist to study, an accident riding on language. He called music auditory cheesecake. Humans never evolved a taste for cheesecake as such; they evolved a liking for fats and sugars, which were scarce, so reward centers fire when we eat them. Survival-related activities such as eating and sex are pleasurable because the brain rewards them, but we can short-circuit the original activity and tap the reward system directly, by eating food with no nutrition, having sex without procreating, or taking heroin. The limbic pleasure centers cannot tell the difference. On this view, music exploits existing pleasure channels that evolved to reinforce an adaptive behavior, presumably linguistic communication. Pinker said it pushes buttons for language ability, for the auditory cortex's response to emotional signals in the human voice, and for the motor system that puts rhythm into walking and dancing. In The Language Instinct he wrote that music is biologically useless and could vanish without changing the rest of our lives much.

Coming from so respected a scientist, this made Levitin and colleagues re-examine an assumption they had never questioned. Others hold similar views: cosmologist John Barrow said music has no survival role, and psychologist Dan Sperber called it an evolutionary parasite, exploiting a capacity for processing complex pitch and duration patterns that had evolved for real, prelinguistic communication. Ian Cross summarized that for these three, music exists purely for pleasure.

Darwin, genes and sexual selection

Levitin thinks Pinker is wrong. He begins with Darwin and notes that "survival of the fittest," spread by Herbert Spencer, oversimplifies the theory. Evolution rests on several assumptions. Phenotype (appearance, physiology, some behaviors) is encoded in genes passed between generations; genes direct protein manufacture and act in specific cells, so eye cells do not grow skin. Genotype gives rise to phenotype. Natural genetic variability exists among individuals. Mating combines genetic material, half from each parent. And spontaneous mutations occasionally occur and may be inherited.

The genes in you now are those that reproduced successfully in the past; everyone alive descends from winners of a long genetic competition. Living long does not pass on genes, reproducing does. Once an organism has reproduced and its young are protected, there is no strong evolutionary reason for it to live longer, and some birds and spiders die during or after mating. Extra years matter only if used to protect offspring, secure resources or help them find mates. Genes succeed when the organism mates and its offspring survive to do likewise.

Darwin saw this and proposed sexual selection: traits that attract mates become encoded in the genome. If square jaws and big biceps attract, men with them reproduce more and the genes spread. Nurturing genes could also spread if the children of nurturers fare better. Darwin thought music was involved. In The Descent of Man he wrote that musical notes and rhythm were first acquired by male or female ancestors to charm the opposite sex, which tied them to strong passions. He thought music preceded speech as a courtship device and likened it to the peacock's tail, a feature with no direct survival purpose except making the bearer attractive.

Geoffrey Miller and the fitness display

Geoffrey Miller linked this to modern society. He notes that Jimi Hendrix had liaisons with hundreds of groupies, kept two long-term relationships and fathered at least three children in the US, Germany and Sweden, and would have had more before birth control. Robert Plant recalled that on Led Zeppelin's seventies tours, wherever he went he was heading toward a great sexual encounter. Top rock stars, such as Mick Jagger, have hundreds of times the partners of an ordinary man, and looks seem not to matter.

Animals advertise the quality of their genes, bodies and minds, and human conversation, music, art and humor may have evolved largely to advertise intelligence. Miller suggests that since music and dance were intertwined through most of history, performing them signaled fitness in two ways. Singing and dancing showed stamina and physical and mental health. Skill showed that one had enough food and shelter to spend time on a useless pursuit. The peacock's tail works the same way: its size tracks age, health and fitness, and signals spare metabolism and resources. Humans do the same with elaborate houses and expensive cars. Levitin adds that men near the poverty line often buy old Cadillacs and Lincolns, and that bling can be seen this way; desire for cars and jewelry peaks in adolescence, when men are most sexually potent. Music making shows an array of physical and mental skills, and time spent developing it suggests resources.

Interest in music also peaks in adolescence: far more nineteen-year-olds than forty-year-olds start bands or chase new music, despite the older group's longer chance to develop. Miller argues music evolved, and still works, as a courtship display, mostly broadcast by young males.

This is plausible given how some hunter-gatherers hunted. Persistence hunting meant throwing projectiles and chasing prey for hours until it collapsed. Tribal dance, if like today's, lasts hours and demands aerobic effort, with stamping, high-stepping and jumping using the biggest muscles, so it would show fitness to take part in or lead a hunt. Schizophrenia and Parkinson's, among other illnesses, impair rhythmic performance, so rhythmic music and dance guarantee health and perhaps reliability and conscientiousness, since expertise needs focus (Chapter 7).

Evolution may also have selected creativity generally. Improvisation and novelty would signal cognitive flexibility and cunning on the hunt. Wealth is a strong attractor for females because it promises food, shelter and protection, so music might seem unimportant. But Miller and Martie Haselton of UCLA showed creativity beats wealth, at least in women. Their idea is that wealth predicts a good dad, whereas creativity may predict the best genes.

In their study, women at different points in their menstrual cycle (peak, minimum and between) rated men described in written vignettes. One man was a creative artist, poor through bad luck; another had average creative intelligence but was rich through good luck. The vignettes made creativity look like an inherent, heritable trait and wealth look accidental. At peak fertility women preferred the poor creative artist as a short-term mate or for a brief encounter; at other times they did not. Levitin stresses that preferences are largely hardwired and not overridden by conscious thought, and reliable birth control is too new to affect them. The best caregivers are not necessarily the best genetic contributors. People do not always marry those they are most attracted to, half of both sexes report extramarital affairs, and far more women want to sleep with rock stars and athletes than marry them. A European study found 10 percent of mothers reported children being raised by men who wrongly believed them their own. Innate preferences are hard to separate from cultural tastes.

Huron's test, age of music, and evolutionary lag

David Huron asks what advantage musical individuals gain over nonmusical ones. If music were a nonadaptive pleasure behavior, as the cheesecake view says, it should not last long. Huron compares it to heroin: users neglect health, die young and neglect offspring, which cuts gene transmission. So music lovers should be at a disadvantage, and music should be recent, since low-value activities do not persist or take much time and energy.

The evidence says otherwise. Instruments are among the oldest human artifacts. The Slovenian bone flute, about fifty thousand years old and made from the femur of an extinct European bear, is a prime example. Music predates agriculture, and there is no evidence that language came before music; physical evidence suggests the reverse. Flutes were probably not the first instruments: drums, shakers and rattles likely came thousands of years earlier, as seen in modern hunter-gatherers and in European invaders' reports of Native American cultures, and singing probably came before flutes. The archaeological record shows music wherever humans lived, in every era.

Mutations that help an organism live to reproduce become adaptations, and an adaptation needs at least about fifty thousand years to appear throughout the human genome. This is evolutionary lag, the delay between an adaptation first appearing in a few individuals and becoming widespread. Adaptations therefore answer conditions of fifty thousand or more years ago, not today. Our ancestors lived very differently, and many modern ills, from cancer and heart disease to perhaps the high divorce rate, come from bodies designed for that earlier life. Levitin jokes that by about the year 52,006 we may have adapted to crowding, pollution, video games, polyester, glazed doughnuts and global inequality, perhaps tolerating close quarters, processing carbon monoxide, radioactive waste and refined sugar, and using now-unusable resources.

So asking about music's origins means thinking of music fifty thousand years ago, not Britney or Bach. Instruments, cave and stoneware paintings, and isolated contemporary hunter-gatherer groups all inform this. In every known society, music and dance are inseparable.

Music as embodied and universal

Arguments against adaptation treat music as disembodied sound performed by experts for an audience. Music has been a spectator activity only for about five hundred years, and the link between sound and movement weakened only in the last hundred. John Blacking says the inseparability of movement and sound marks music across cultures and eras. We would be shocked if symphony audiences stood, clapped and whooped, yet that is normal at a James Brown concert, which is closer to our nature. Polite, purely cerebral listening, with emotions felt inwardly, runs counter to our history. Children sway and shout at classical concerts until trained to be "civilized."

A widely distributed trait is taken to be genomic whether adaptation or spandrel. Blacking argues that universal music-making in African societies shows musical ability is a general human characteristic, not a rare talent. Cross adds that musical ability is not just production; nearly everyone can listen to and understand music.

Other ways music may have been selected

Besides sexual selection, there are further proposals. One is social bonding and cohesion. Humans are social, and group music may have fostered togetherness, synchrony and practice at turn-taking. Singing around a campfire might have kept people awake, warded off predators and built coordination and cooperation.

Levitin's work with Ursula Bellugi on Williams syndrome (WS) and autism spectrum disorders (ASD) offers evidence. WS is genetic, causes abnormal neural and cognitive development and intellectual impairment, yet people with it are especially musical and sociable. ASD, often also with intellectual impairment (whose genetic basis is disputed), is marked by difficulty empathizing and reading others' emotions, usually including inability to appreciate art and music aesthetically. Some people with ASD play music, some with technical skill, but they do not report being moved; anecdotal evidence suggests they are drawn to structure. Temple Grandin, an autistic professor, finds music "pretty" but doesn't get why people react as they do.

These form a double dissociation: a gregarious, highly musical group and an antisocial, less musical one. This hints at a gene cluster influencing both outgoingness and musicality, so deviations in one would accompany deviations in the other. The brains also differ complementarily: Allan Reiss showed the neocerebellum is larger than normal in WS and smaller in ASD, fitting the cerebellum's role in music. An unidentified genetic abnormality presumably causes the neural abnormality, leading to enhanced musical behavior in one and diminished in the other. Julie Korenberg speculates that WS lacks normal inhibition genes, so musical behavior is uninhibited. Reports on 60 Minutes, in a film narrated by Oliver Sacks, and in newspapers say people with WS are immersed in music. Levitin's lab scanned WS brains during listening and found far wider neural activation than in others, with the amygdala and cerebellum notably stronger. Their brains were humming.

A third argument is that music promoted cognitive development, preparing ancestors for speech and representational flexibility. Singing and playing may have refined motor skills for the fine control speech or signing needs. Trehub suggests music prepares infants for mental life and lets them practice speech perception in a separate context.

Language is not learned by memorization, for two reasons. Empirically, children overextend rules, saying "goed," "buyed," "swimmed" and "eated," because the developing brain forms and prunes connections and instantiates rules; brighter children make these errors earlier, and since few adults do, children are not mimicking. Logically, we all say sentences we've never heard, so language is generative and infinite; for example, "I don't believe" can be prefixed to any sentence repeatedly. Music is generative too: any phrase can be extended by adding a note at the start, middle or end.

Cosmides and Tooby argue music exercises the child's brain for language and social interaction. Lacking specific referents, it is a safe, nonconfrontational way to express mood. It may pave the way to prosody before phonetics are processable, and is a kind of play building exploratory competence leading to babbling and complex speech. Mother-infant music universally involves singing with rocking or caressing. In the first six months, as in Chapter 7, the senses are not separated and the future auditory, sensory and visual cortices are undifferentiated, so Simon Baron-Cohen says the infant lives in psychedelic splendor without drugs.

Cross concedes that music today differs from music fifty thousand years ago, but ancient music was likely heavily rhythmic, which explains why rhythm moves us. Rhythm stirs the body; tonality and melody stir the brain; together they link the cerebellum with the cerebral cortex. Examples are Ravel's Bolero, Charlie Parker's "Koko" and the Rolling Stones' "Honky Tonk Women." Rock, metal and hip-hop have been the world's most popular genres for twenty years. Mitch Miller of Columbia Records predicted in the early sixties that rock and roll would die; in 2006 it has not. Classical music from about 1575 to 1950, Monteverdi to Stravinsky and Rachmaninoff, is no longer being written. Composers such as Philip Glass and John Cage, and lesser-known ones, write art music rarely played by orchestras, unlike Copland and Bernstein, whose works the public enjoyed. Contemporary "classical" music lives mostly in universities, is heard by almost no one, deconstructs harmony, melody and rhythm, is purely intellectual, and is not danced to.

Other species

A fourth argument is that other species use music similarly, but one must avoid anthropomorphizing. A dog rolling in grass looks happy to us, but male dogs do it to cover themselves in a pungent smell, preferably of a dead animal, to seem skilled hunters. Likewise birdsong that sounds joyful may not be meant or heard so. Still, birdsong holds a special place; Aristotle and Mozart thought it as musical as human compositions.

Birds, whales, gibbons and frogs vocalize for various purposes. Chimpanzees and prairie dogs give predator-specific alarms: chimps have one call for an eagle (hide under something) and another for a snake (climb a tree). Male birds claim territory by voice, and robins and crows have a call for predators such as cats and dogs. Other calls relate to courtship. Male songbirds usually sing, and in some species a bigger repertoire attracts mates more readily. In a study playing songs to female birds over loudspeakers, they ovulated sooner with a large repertoire than a small one. Some males sing courtship songs until they die of exhaustion. Several species build songs generatively from basic sounds, and the most elaborate singer usually mates best, an analogue of music's role in sexual selection.

The summary case

Levitin concludes music's evolutionary origin is established. It is universal in humans, meeting the biologists' widespread-trait criterion. It is old, refuting the cheesecake idea. It uses specialized brain structures, including memory systems that can work when others fail. It has analogues in other species. Rhythm excites recurrent neural networks in mammal brains, including loops among motor cortex, cerebellum and frontal regions. Tonal systems, pitch transitions and chords build on the auditory system's properties, products of the physical world and the nature of vibrating objects; the auditory system develops around the link between scales and the overtone series. Musical novelty attracts attention and fights boredom, aiding memory. Just as the discovery of genes and DNA's structure revolutionized natural selection, a new revolution may concern evolution's dependence on social behavior and culture.

Mirror neurons and cultural evolution

One of neuroscience's most cited discoveries of the last twenty years is mirror neurons, found by Giacomo Rizzolatti, Leonardo Fogassi and Vittorio Gallese while studying reaching and grasping in monkeys. They recorded a single neuron while a monkey reached for food. When Fogassi reached for a banana, a movement-related neuron fired though the monkey did not move. Rizzolatti recalls first suspecting equipment failure, but everything checked out and it repeated. A decade of work has shown primates, some birds and humans have neurons that fire both when performing and when watching an action.

Their presumed purpose is training the organism to make new movements. They have been found in Broca's area, tied to speaking and learning to speak, and may explain how infants imitate parents' faces. They may also explain why rhythm moves us emotionally and physically; some neuroscientists speculate they fire when we watch or hear musicians, as the brain works out how the sounds are made in order to echo them back as signaling. Many musicians can replay a part after hearing it once, likely with mirror neurons' help. Levitin suggests that as genes pass protein recipes, mirror neurons, now aided by sheet music, CDs and iPods, may carry music across people and generations, enabling cultural evolution, through which beliefs, obsessions and all art arise.

Why courtship display in a social species?

For solitary species, a ritualized courtship display makes sense because a pair may meet for minutes. Why would highly social humans, who watch one another over long periods in many situations, need stylized singing and dancing to show fitness? Primates live in groups with long-term relationships and social strategies, and hominid courtship was probably long-term. Memorable music would lodge in a prospective mate's mind, making her think of her suitor while he was away hunting and favor him on return. Rhythm, melody and contour reinforce one another and make songs stick, which is why ancient myths, epics and the Old Testament were set to music for oral transmission. Music is weaker than language at evoking specific thoughts but better at arousing feelings. Combining the two, best seen in a love song, makes the best courtship display of all.


Back Matter: Appendices A and B, Bibliographic Notes

Appendix A: This Is Your Brain on Music

The first appendix is a pair of labelled diagrams of the brain. Its opening point is that music processing is spread throughout the brain rather than housed in one spot. The figures show the main computational centres involved. The first is a side view with the front of the brain at the left. The second is an inside (cut-through) view from the same angle. Both are redrawn from illustrations Mark Tramo published in Science in 2001, with newer findings added. The OCR of the labels is partly garbled, but the roles listed are as follows.

Side view:

Inside view:

Appendix B: Chords and Harmony

Building chords in a key. Within the key of C, the only "legal" chords are those built from the notes of the C major scale. Because the spacing between scale tones is uneven, some of these chords come out major and some minor. A standard three-note chord, a triad, is built by starting on a scale tone, skipping one tone, taking the next, skipping another and taking the next. Starting on C gives C-E-G. The first interval, C to E, is a major third, so this is a major chord, specifically C major. Starting on D gives D-F-A. The first interval, D to F, is a minor third, so this is a D minor chord.

Major and minor chords sound clearly different. Most non-musicians cannot name a chord or label it major or minor, but they can tell the two apart when they hear them back to back. Several studies have shown that non-musicians' bodies respond differently to major and minor chords, and to major and minor keys.

The seven chords of a major key. Built this way on the seven degrees of the major scale, three chords are major (degrees one, four and five) and three are minor (degrees two, three and six). The seventh-degree chord is diminished and consists of two stacked minor thirds. The key is still called C major despite the three minor chords because the root chord, the one the music points toward and that feels like home, is C major.

Harmony and typical progressions. Composers generally use chords to set a mood, and the use of chords and the way they are strung together is called harmony. The word is also used for two or more singers or players performing different notes together. The author says the two senses are conceptually the same idea. Some chord sequences are used so often that they come to define a genre.

Seventh chords. A "7" after a chord name means a tetrad, a four-note chord, formed by adding a fourth note on top of a triad. In the dominant seven, such as G7 ("G seven" or "G dominant seven"), the added top note is a minor third above the chord's third. Tetrads allow much richer tonal variety. Rock and blues mostly use only the dominant seven, but two other types are common, each with its own emotional flavour:

Why the dominant seven wants to resolve. The dominant seven occurs naturally within a key (diatonically) when built on the fifth degree of the major scale. In C, G7 uses only white notes. It contains the formerly banned tritone, and it is the only chord in a key that does. The tritone is the most unstable interval in Western harmony, so it carries a strong perceptual urge to resolve. The chord also contains the most unstable scale tone, the seventh degree (B in C). For both reasons it "wants to" return to the root, C. That is why the V7 chord (G7 in C) is the most typical, standard and clichéd chord just before a piece ends on its root. G7 to C major pairs the single most unstable chord with the single most stable one, which gives the maximum tension and resolution possible. When some Beethoven symphonies seem to end on and on, the composer is repeating this two-chord move over and over until the piece finally settles on the root.

Bibliographic Notes

This section is a reading list, not a continuation of the argument. The author says it is incomplete and holds additional sources most relevant to the points in the book. He wrote the book for non-specialists and simplified topics without oversimplifying them. Fuller accounts are in these readings and in the works they cite. An asterisk marks the more technical entries, mostly primary research papers plus a few graduate textbooks. Each entry carries a short comment on what it supports. The notes follow the book's chapters.

Introduction. The notes cover the philosophy of mind (Paul Churchland, whose introduction the author borrowed from for the passage on curiosity solving great mysteries). They also cover evolutionary psychology (Cosmides and Tooby; Deaner and Nunn on evolutionary lag; Geoffrey Miller on sexual selection and music; Robert Sapolsky on stress and evolutionary lag) and Roger Shepard's papers on the evolution of mind. Other entries are Karl Pribram's papers from the course the author took, an interview source for a Paul Simon quote, and the Rolling Stone Encyclopedia of Rock & Roll, cited for how little space it gave U2 in 1983 compared with Adam and the Ants.

Chapter 1 (pitch, timbre, scales and sound). Entries here cover:

Chapter 2 (rhythm, melody and grouping). The recommended reading covers:

Chapter 3 (the brain as machine and the neuroanatomy of music). This is the most heavily cited chapter in the notes. The references include:

Chapter 4 (expectation, schemas and brain organisation). Reading covers:

Chapter 5 (memory, categories and absolute pitch). The notes cover:

Chapter 6 (rhythm, timing, the cerebellum and emotion). Reading covers:

Chapter 7 (expertise and talent). The notes cover:

Chapter 8 (musical preference and development). Reading covers:

Chapter 9 (evolution of music). Reading covers:

Acknowledgments and Index

The book ends with an acknowledgments section, where the author thanks the engineers, producers, musicians, scientists, collaborators, students, colleagues and editors who helped him, and a subject and name index. Neither is summarised here.