Ultra-Long-Term Music Backup

“Solon of Athens heard his nephew sing a song of Sappho’s over the wine. And since he liked the song so much, he asked the boy to teach it to him. When someone asked why, he said, ‘So that I may learn it and then die.'”1

The song was lost 1,000 years ago, like all songs written by Sappho. If the song really existed, it would have had its own mood, feelings, a general spirit, and musical details like melody, rhythm, and tempo. None of them were preserved.

We say however that she’s a poet, since all that remains are Sappho’s lyrics. And those also are mostly gone. From the eight volumes of her poems common in the Roman and Greek world, we have one single poem remaining intact. The fragments which lasted look like this, fractured and destroyed:

Papyrus fragment of Sappho’s lyrics. This is a big fragment.

Humans are about 300,000 years into our existence. From our initial 295,000 years, all music is lost. All writing is also lost.

When people began music-making is not something we know. But if these prehistoric people were like the people we know today, in nearly every culture, they sang and danced.

Subtlety, beauty, and thoughtfulness can be seen in art of separate peoples and many periods and places. They’re not modern creations. Our glimpses of this in the cave paintings in Chauvet left 30,000 years ago:

Animal sketches from the Chauvet Cave, made 30,000 years ago.

The music they made is gone now. We’ll never have a way of recovering it.

When we find ancient music scores, they are hard (or impossible) to decipher.

In 1,400 BC, ancient Hurrians notated a melody in cuneiform. It’s our earliest score now surviving.2 Decoding it hasn’t been easy. So far, experts have published five different interpretations. These are noble attempts, but without knowing anything about this, I’m guessing that four of them are wrong. So the music is in fact lost unless one of them is right.

Luckily, we may have earlier melodies preserved in the Vedic Sanskrit manuscripts of the Samaveda. They were notated around 1,200 B.C., but the melodies were likely older:

Manuscript of the Samaveda, originally from 1,200 B.C. The melodies were likely much older.

The texts accompanying the melodies were the new parts, and were added at some long after the melodies were composed. We don’t know how old the melodies truly are.

We have to wait another 1,300 years until we see another (complete) melody preserved—the Epitaph of Seikilos, in Greece:

Epitaph of Seikilos image
The Epitaph of Seikilos: music and lyrics preserved on a grave stone. The lyrics are the big letters on each line. The notes are the small letters above the big letters.

This is the melody in modern notation:

These snippets are extraordinary because they are so rare. They are the only music preserved from our first 299,000 years. Music made for this entire time, apart from these snippets, is now dust.

Preservation Methods and their Points of Failure

Humanity may last for another 300,000 years, but no music paper will last that long. No computers will last that long. Our current media, such as paper, pigments, and electronics, were not designed to endure for such time periods. Electronic media is made to last no more than a decade. This is a point of failure that most of us deal with, at some point, with our personal files.

We therefore rely on copies, and copies of copies, and so on. This is trivial for us. But copying—both analogue and digital—requires a vast infrastructure that is invisible to a layman. Copying is trivial only when our infrastructure is finely tailored to make it trivial.

Floppy diskettes were omnipresent in 2002 and unreadable in 2012. Their infrastructure fell apart in 10 years. Their widespread copying stopped, and much of the information on those discs will last only as long as the discs themselves. Relying on copying, therefore, is relying on the longevity of a skillset and infrastructure. Both are potential points of failure in this system.

Digital scores (kept in personal files or found on the internet) rely on data migrations across hard drives. This infrastructure has subtle cracks that are hard to see. Copying files is easy on a day-to-day level, when we stay within the same operating system. But if you copy to a different system, it gets more risky; the more data, and the more dissimilar the systems, the more risk that not all data will be copied.

Data migrations seem trivial until they glitch and lose data—maybe the new operating system has different filename requirements, different security requirements, different character sets allowed. File systems have hundreds of these parameters, any of which can be incompatible.

MySpace lost nearly all its music during a server migration—everything from its heyday, from 2003 to 2015. That’s 53 million songs recorded by 14 million musicians. Data migration plagues even well-established companies and organizations, from major banks to the BBC.

Major organizations routinely lose massive amounts of data. Ransomware is the newest threat, but even simple programming errors can destroy data. Kyoto University lost 77 Terabytes of research data when a backup script accidentally overwrote some original data during a program update.

Even if you keep 99% of your data with each migration, any losses will compound quickly. After 10 migrations you will have kept only 90.4% of your data. After 100 migrations, you’ll have 36.6% remaining. After 458 migrations, you will be down to 1%.

# of Data MigrationsData Remaining after 1% loss during Each Migration
199%
10036.6%
1,0000.004%

After 300,000 years of error-prone copying, almost no original data will remain. If we need a data migration every 5-10 years, we require 30,000-60,000 migrations. Even if we manage to cut the loss down to only 0.01% loss of our data, 30,000 migrations gives us only about 5% remaining data. And 60,000 gets us a pathetic 0.2% remaining.

# of Data MigrationsData Remaining after .01% loss during Each Migration
199.99%
30,0004.978%
60,0000.248%
Compounding errors: the evil twin of compounding interest.

Backing up Institutions

Institutions like libraries and museums are our best hope for preserving items. They’re imperfect, but have the best track record in many cultures—especially Eastern Asia, India, Islamic cultures, and Western countries. However imperfect, they have the knowledge, skills, manpower, and funding to preserve things.

In private hands, by contrast, an item will change hands many times; with each change, there is a new, non-zero chance that it will be damaged or lost. With each new generation it will change owners. It may pass to heirs who don’t see its value. Or, if they see its value, they may not have enough money to protect it. Or know how to maintain and arrest damage from humidity, light, or microbial growth. They may lack the security forces to protect it from theft. It may be lost during a move, or buried somewhere in a heap of detritus.

This was the fate of many of J.S. Bach’s works. Half of his manuscripts were passed to his son Wilhelm Friedemann Bach. He valued the collection, but didn’t have enough money to maintain it. He sold it off, and nearly all of those works are lost. We lost about 100 of Bach’s cantatas. Many works mentioned in Bach’s obituary did not survive. We only have copies of many of his keyboard works because his students made copies of them.

Over a thousand years, a valuable item may change hands 4 times per century. That means it could change hands 40 times—making the likelihood of one of these misfortunes far higher than if it were owned by a single institution. Each one of these 40 moves is a potential point of failure.

But institutions—and their collections—still face several threats:

  • Funding cuts.
  • Reliance on an electrical grid, which has been mostly reliable, but vulnerable to black-swan events.
  • Data migration mistakes.
  • Ransomware and other computer-related threats.
  • Physical security threats.
  • War, famine, disease.
  • Ideological attacks.

Libraries and museums also have finite life spans. They have a staff to pay, buildings to maintain, heating and cooling to maintain, supplies to purchase. They rely on money from other well-funded institutions. If they lose this money, they fall apart. A recent example is the looting of the Baghdad museum around 2003. They also attract attention during times of ideological fervor, making them subject to censorship, book burnings, and destruction.

But all media relies on infrastructure to keep it safe. This includes any institution that funds a library, buildings that house the books, the expertise to maintain them, and a reliable electrical grid. A long term migration schedule should therefore also be based on the time span over which any major institution or infrastructure may become unreliable.

Libraries depend on funding and, more recently, on an electrical grid. If either is absent for, say, a year, irreparable damage can be done to a collection. If they are absent for 10 years, a collection may be effectively destroyed. Damage can come from weak security and theft, burst pipes during a freeze, rain water, fire from lack of fire-suppression systems, the inability to maintain the structure, and loss of personnel due to travel hardship.

New threats will appear to libraries and museums; in the last five years, institutions around the world have been threatened by ransomware. This kind of threat was unimaginable even 20 years ago.

This may seem like a distant worry. After all, almost all of the works of, say, Beethoven and Brahms have been preserved. None of them are under threat, or have ever been under threat.

But 200 years is no test of long-term preservation. Sappho’s poetry was popular 1,000 years after she died, but disappeared within another 1,000 years. Her esteem in the Greek-speaking world had been roughly equivalent to where we might place Brahms or Milton. That anecdote at the beginning of this post (even if it is apocryphal) was recorded 1,000 years after her death, when her reputation was still strong. She was called the 10th Muse and simply “the Poetess” (a counterpart to Homer, “The Poet,” who occupied a place roughly akin to the Christian Bible in Western Europe). By 1500, only her name survived—Renaissance writers had no examples of her poetry to read. (Fragments were slowly discovered since then, mostly in trash heaps and archaeological finds.) So it seems that even 1,000 years of prestige and widespread copying is no guarantee that a work will survive 2,000 years.

The cave paintings of Lascaux lasted for 17,000 years, but were nearly destroyed within 20 years after opening to tourists. They had been well preserved in their original environment. After they were rediscovered in 1940, they degraded quickly from fungal and lichen growth. They might have been destroyed entirely if the caves hadn’t been closed.

When compared with our gaps in cultural artifacts, 200 years is microscopic. Lion-Man, one of the oldest surviving (unambiguous) cultural artifacts, is 40,000 years old. Humanity existed 260,000 years before that. That is a gap of 65,000 generations of people. We have no idea of what those people thought, how they felt. We know that they made cultural artifacts, because that’s what people do in every culture. But they were made on perishable media, such as textiles.

We have to wait another 10,000 years until we see cave paintings in Chauvet, and 13,000 more years for the paintings in Lascaux. Gaps of 5,000 years, 10,000 years, and even 100,000 years are the rule. We may assume that because modern society is so robust, we will be less likely to leave such gaps. But are we really building anything now that will last, unambiguously, for 100,000 years?

Data Migration Schedules

All backups require a data migration schedule: a plan to copy data to new media before the old media degrades. For example, a spinning magnetic disc has a typical lifespan around 10 years, but often starts degrading after 5 years or so. So its data should be migrated around year 4. Papyrus had a lifespan of around 50 years in European climates, so anything not copied in that time frame was lost.

Libraries and museums have finite lives. A library is, in a way, a data storage center with a lifespan measured in centuries. Some institutions have surpassed that, such as the University of al-Qarawiyyin in Fez, Morocco (founded in the 9th century), or the University of Paris (founded in 1150). But, just as I have 15-year-old hard drives that work, we should not assume that they will last forever, and that their collections are forever safe.

Although copying is necessary to preserve collections, it is risky. You have to translate one system’s data to another. It is a potential point of failure. With digital items, the migration may need to cover multiple drives, multiple operating systems, multiple directory structures, at a fairly strict schedule. It’s a major chore, often a headache, will never be done as often as it’s needed. It needs a professional staff, and even then, with huge resources, can lead to disaster (e.g., the MySpace losing most of its files).

The more data migration you need, the more points of failure you introduce. Even if you lose only 1% of data with each migration, after 10 migrations you will have kept only 90.4% of your data. After 100 migrations, you’ll have 36.6% remaining. After 458 migrations, you will be down to 1% of your original data.

With analogue copies, the rate of loss with each migration is even higher. They generally lack the precision of digital copies. A medieval scribe, for example, was more likely to introduce errors in books than a simple copy & paste operation. The reel-to-reel tape copies of music in the 60s changed their sound sources so much that they were, in practice, a distortion effect. 

Their centralized nature means that we are putting all our eggs in one basket. The best approach is to diversify in as many ways as possible:

  • Geographic diversity;
  • Diversity of media, since we don’t know what kind of conditions our collections will be forced to endure;
  • Diversity of ownership, such as private ownership.

Every institution is susceptible to black-swan events: a never-seen-before, highly unlikely event that causes major economic or infrastructure damage.

The likelihood of any given black-swan event is low; so low that, individually, they are impossible to predict. But the likelihood of any black-swan event at all is high. Future conditions will not perfectly match past conditions.

Furthermore, there is no way to prepare for all of them. This does not mean that we should not prepare at all.

The challenge here is that we don’t know how long our infrastructure will last. It could be decades, centuries, or millennia. But we should not assume that it will be millennia; almost no institutions have lasted that long. And our electrical grid is new and famously fragile.

In some ways, we’re living in a golden age of written music. Physical scores are well-preserved in libraries and university collections, especially critical editions of major composers. Digital copies are distributed by the million at will, and are easily accessible from sites like IMSLP. Dover has reprinted millions of high-quality scores on acid-free paper, including critical editions that are out of copyright. Notation software allows for fast creation of clearly-notated scores.

Longterm backups might seem like they should be a low priority, given that they are already backed up by the million.

But every backup system has points of failure. With our system, some points of failure are:

  • Reliance on financial support from well-funded institutions.
  • Reliance on a dependable electrical grid. When electricity is unavailable, data migrations (for digital media) are impossible. Controlling heat and humidity will be hard, making paper collections more fragile.
  • Reliance on a dependable internet connection. When the web goes down, it cannot be used as a backup and data migration tool.

Everything is subject to the whims of economics, logistics, and taste. And to the laws of physics. Nothing can withstand their pressure forever.

Digital collections rely on electronic media, which are fragile. To preserve files stored on spinning hard drives, for example, they must be migrated to new drives once every 5 years. For logistical reasons, this is unlikely in many cases.

Here are some cultural artifacts that were once reproduced by the thousand or even million, but which disappeared or are on the verge of disappearing:

  • In 764, the Japanese Empress Shōtoku funded the creation of one million small wooden pagodas. Each had a printed prayer inside. 1,300 years later, about 45,000 wooden pagodas survive, and 4,000 printed prayers survive. For the pagodas, 95.5% were lost, but only 300 are unbroken. For the printed prayers, 99.6% were lost. At that rate, 766 were lost every year. This is particularly worrisome since the wooden pagodas and printed prayers were treated as treasured objects by many generations. The passing of time bodes ill for their long-term survival. Most are held at a single temple in Nara, Japan—a temple that had major fires in 604 and 1949. If it experiences another, and the printed prayers are lost, 99.95% of all the prayers would be gone.
  • The printed texts, however, are indecipherable. It was early in the adoption of the Chinese characters for the Japanese language, the system they used was not yet standard. The Chinese characters were used phonetically, adapted from the sounds of the Chinese language, but to represent sounds of Sanskrit as it slowly made its way through the non-Sanskrit speakers living in China, then Korea, and finally Japan.
  • Although printing began in China, these printed prayers predate any known examples from China. This is likely due to two factors: their high quantity and their storage in (relatively stable) temples. The earliest known Chinese text, for example, was found in the Mogao caves, that had been walled-up for 900 years.
  • Once-common, widely-copied works of literature are now lost. For example, Homer’s comedy Margites, two-thirds of Aristotle’s works (including a second book of Poetics, on comedy), Augustus’s autobiography, almost all of Sappho, about most plays of Aeschylus (about 80 out of 90), most plays of Aristophanes (29 out of), everything by Pythagoras, Confucius’ book on music.

Sappho is a troubling example. Along with Homer, she was considered the major poet by ancient Greeks; she was nicknamed the Tenth Muse. Her works were ubiquitous, so common that no one thought it would be crucial to preserve.

If a piece follows the “rate of loss” of the Shotoku prayers, it will take 1,300 years for 99.6% of printed materials to be lost. For a work like Beethoven’s Fur Elise, let’s pretend there are one million printed copies. If there were no more copies made, but the music were treated as treasured objects, then about 4,000 copies would remain. For a masterpiece like Franck’s final organ chorales, I would be shocked if there were more than 10,000.  If you add up its top 10 search results, Beethoven’s Fur Elise has a total of 200,000,000 views on YouTube; the Franck Chorales have about 100,000. By this crude measure, Fur Elise is 2,000 times more popular than the chorales. By this crude estimation, there would only be 2 copies remaining of these masterpieces—and that’s assuming they were to be treated as treasured objects!

Libraries now are well-funded; this may not always be the case. Libraries rely on patronage to flourish. They go through periods of solid and poor funding. The Library of Alexandria was the largest in the world, but faced a period of slow decline. All of its works were lost. There were dozens of libraries in Rome alone; but they did nothing to transmit the works of the Greek and Roman antiquities to modern Europe.

Libraries are common targets of ideological fervor. Wikipedia has a long list of book burnings; destruction of heretical works is a favorite tool. During the Middle Ages, there was a heavy bias against non-Christian works. Major works from the period of English music from which 3rds and 6ths became dominant are now lost, due to the destruction of English monasteries after Henry VIII. Libraries in the Soviet Union and Nazi Germany suffered from extreme culling against heretical works.

If this seems overly pessimistic, consider the case of Greek music.

The Borders bookstores from which I bought most of my scores are gone. The ones that remain do not sell music scores. Amazon.com is now the largest seller of sheet music—and a great variety of them—but we don’t know how long that will last. Amazon has many businesses, and selling scores is low on their list of priorities. They may one day change owners, change focus, and lose interest in selling scores. It is a single point of failure.

  • Paper scores are heavy and take up space. When people donate their books to libraries, most are sold or destroyed. They are costly to make and store. They also decay with age. There will always be an economic argument against creating and storing them.
  • Digital copies are easy to copy and take little space. But they have their own problems. Tiny magnetic polarities change easily. This makes it a fragile, ephemeral medium.  In 1996, five billion floppy disks were being used. (https://web.archive.org/web/20080706151833/http://www.businessweek.com/1996/33/b3488114.htm) They were mostly unreadable fifteen years later.
  • Digital copies are easily spread via the web, but the web suffers from rapid changes. A quarter of all links in NYTimes articles are now defunct. This will only increase.

My e-reader lasted 10 years, an eternity in electronics time. Its rubber shell is breaking down, turning into a gooey, sticky mess.

The forces of economics seem indomitable, and in a way, they are. But, like a river during a storm, they move vast amounts of material but are chaotic and often irrational. Bach was out of favor for 75 years; half of his manuscripts went to his son, who died penniless. All of those manuscripts were destroyed.

And I wonder how long they will last. The forces of economics and fashion that brought them to me will not last forever.

Printed material is tougher than electronics, but it too is subject to loss.

Some entertainment can make money; most won’t. Most art loses money. It requires some form of sponsoring, patronage, or philanthropy.

Notes:

  1. Stobaios, Florilegium 3.29.58, from “If Not, Winter: Fragments of Sappho,” by Anne Carson, p. xiii.
  2. This is one of 36 clay tablets found nearby with notated works. Incredibly, some of them have their composers’ names included. Unfortunately, we only have small pieces of the tablets and melodies.

Are MP3s Perfect? Sometimes Yes; Sometimes No.

Summary

In my tests, the fidelity of MP3 encoding/playback varied according to a few factors. These factors included the density and noisiness of the audio source and (more surprisingly) which media playback program was being used.

  • For clipped, distorted, music (e.g., Top-40 pop from 2010), the audio was so full of digital noise that a bit of extra MP3 noise was inaudible.
  • For quiet, transparent audio (e.g., solo piano), high-bitrate MP3 encoding hardly changed the audio signal. The algorithm was nearly ideal there.
  • For music in between (loud, but not as ultra-clipped as Top-40 music from 2010, and not as clean as a solo piano recording), even a high-quality MP3 encoding sounded slightly noisier and more distant.
  • But even when the algorithm was (nearly) perfect, consumer audio programs often messed with MP3 sound quality in a way that they didn’t with WAV files. Some audio playback programs distorted all MP3 audio. No matter how well-encoded, MP3s often sounded noisier and more distant than WAV files. This was the biggest surprise to me.
  • Some playback programs distorted all audio, including both MP3s and uncompressed WAV files.

Testing

Many confounding variables affect audio quality, and they affect each other in ways that are hard to isolate. For example, changing one parameter (e.g., volume) might affect another parameter (e.g., distortion), for which I might try to compensate by changing another parameter (e.g., EQ settings). Any of these factors in turn may add their own distortion, loudness, etc., to different frequencies.

Other variables affecting our sound perception are just part of being a fallible human, such as selective attention to different details, tiredness, and the placebo effect.

Below are variables I’ve been trying to isolate in my own testing. A few surprised me by how much (or little) they affect the perceived audio:

  • Different music when testing—many audio processing algorithms work better on some material than others. For example, a compressor or reverb that sounds great on drums may sound less great on the voice. If the MP3 algorithm follows this pattern, we should test it on a bunch of different sounds. (For example, ABX Digitalfeed uses ultra-processed pop music from around 2013 as its test material. Low-bitrate MP3s perform well in these comparisons, but not as well with other material.)
  • Different audio playback programs—this is more of an issue than I thought. To my surprise, many consumer audio programs (like Windows Media Player and VLC) play MP3s differently from uncompressed audio. I assumed decoding MP3s is simple, that programs would output an identical stream of 1’s and 0’s sent to our audio interface/speakers. Turns out this is wrong. (Luckily, I haven’t found these differences when they’re imported into DAWs.)
  • Different MP3 encoding—luckily, we have consensus that LAME MP3 encoding at 320 kbps is the best to use for testing. Unluckily, streaming services often use much lower bitrates. They’re improving, but I sometimes encounter tracks on these services that have obvious swirlies in strings and cymbals (compression artifacts that are absent in lossless formats).
  • Different bitdepths—these are easy to separate. But, as a practical matter, it’s not part of my routine. So I often end up comparing my uncompressed 24 bits vs. compressed 16 bits. Like many musicians/engineers, I render my own music in 24 bits. MP3s are 16 bits. I have to go out of my way to isolate this variable.
  • Different audio equipment and listening environments—we have different needs from our audio when listening on a noisy bus with $20 earbuds vs. listening at home with $10,000 of clunky equipment.
  • Different testing methods—it’s hard to build an audio test, and most online tests were not made by audio researchers or engineers. Some have weird interfaces, barrage the listener with a bunch of different sounds, and make it hard to isolate all the different variables listed here.
  • Different levels of familiarity with the quirks of human hearing—some well-known quirks of human hearing can make comparisons hard. For example, if you lengthen the time between playback of different audio sources, you make it harder to judge their differences. And simply making a track .5 dB louder will bias a listener in favor of it (even an experienced producer or audio engineer). This is a challenge for evaluating any audio processing.
  • Focus on different aspects of audio—some people want a powerful bass, some want a liquid-smooth high-end, some want clear vocals, some want pleasant background music.

I’ll look at the first two variables—different kinds of audio and different audio programs.

MP3 Playback in Common Audio/Media Programs

For some reason, many consumer audio programs play MP3s differently from uncompressed audio. This was a surprise to me.

Here’s an example of a waveform altered by playback in VLC, a freeware media program. The MP3 waveform is red, and the uncompressed waveform is blue. The MP3 is slow to reach full volume, and slow to decrease volume, turning it into more of a sawtooth wave:

Waveforms from Brahms’ Intermezzo Op. 118, No.6. This is output from VLC, a consumer audio program: MP3 is in red vs. uncompressed audio in blue (WAV). Identical portions of the waveform are purple.

Here are the same audio sources inside a DAW. The MP3 is still in red and the uncompressed is in blue, but it’s hard to tell because they are nearly identical (making a purple line). The only difference is at the very peak of the waveform:

Brahms’ Intermezzo Op. 118, No.6: inside Sonar (a professional DAW), the exact same MP3 vs. uncompressed audio (WAV).

Here is a spectrogram of that same audio source, as played in VLC (consumer audio program). This lets you see some differences between the MP3 and the WAV. Some parts get brighter (louder), some get darker (quieter). Overall, the MP3 is a bit louder/noisier in this section. I circled a few areas where VLC adds audible noise:

Brahms’ Intermezzo Op. 118, No.6: spectrogram of a recording made from VLC, a consumer audio program. The MP3 playback adds noise in the circled areas. It also removes noise from other areas.

These aren’t artifacts of my recording setup. To make sure, I summed a WAV and a (re-recorded) phase-flipped WAV. This generated faint noise, nearly inaudible even when the gain was jacked up. I then repeated this process, but substituted an MP3 for one of the WAV recordings. If the MP3 playback were affected the same as WAV playback, it would look similar to the noise generated by the WAV + WAV-flipped-Ø.

Windows Media Player had fewer artifacts than VLC, but still generated very audible artifacts:

Brahms’ Intermezzo Op. 118, No.6: major audible differences when summing an MP3 with a WAV-flipped-Ø.

I did a similar process inside a DAW, using the same MP3 and uncompressed files. (Brahms, Intermezzo Op. 118, No. 6.) There were no audible differences at all when an MP3 was summed with a WAV-flipped-Ø. Just one inaudible blip at -84 dB:

Brahms’ Intermezzo Op. 118, No.6: inside a DAW, there is almost no difference between this MP3 waveform and the WAV waveform. This is the result when the MP3 was summed with a WAV-flipped-Ø. This is for a recording of Brahms’ Intermezzo Op. 118, No. 6. Using other audio sources, however, you’ll see much bigger differences.

Overall, the distortion added by these programs give MP3s a subtle lo-fi sound. This isn’t always a bad thing—lots of people love the lo-fi sound. But, in general, you don’t want it added automatically to your audio. The musician/producer/engineer should be calling the shots on this, adding it where desired. In my experience, this is most of the reason MP3s can sound different from WAVs.

As a non-programmer, I assumed that decoding MP3s is easy. I thought that, as long as you have code that decodes an MP3, it decodes them identically to other programs, and outputs an identical stream of 1’s and 0’s sent to our audio interface/speakers. Surprisingly, some part of this assumption is wrong.

Luckily, this doesn’t affect playback in ProTools or Sonar, which are Digital Audio Workstations (DAWs) meant for professional editing. Those programs don’t usually play MP3s directly; to play an MP3 file, the program needs to import it and convert it into something uncompressed.

But programs like VLC and Windows Media Player have a noticeable effect. They add noise here, remove noise there, and mess with transients. They do this to MP3s much more than to uncompressed files like WAV.

The MP3 Algorithm on Different Sources

In my tests, MP3 waveforms were visibly different from their original audio waveforms. But the changes were smaller than I expected. In fact, for some sources (like solo piano music with little processing and a lot of headroom), the MP3 algorithm produced a (near) 1:1 match with the original audio.

There are situations when MP3’s were indistinguishable from WAVs:

  • When the music texture is sparse (e.g., solo piano music or solo vocals). A high-quality encoding can reproduce a WAV with a nearly 1:1 match.
  • When the sounds are so processed/noisy/distorted that more distortion is hard to hear. I suspect this is true when the distortion comes from plugins (e.g., compression, limiting, aural excitation, EQing) circa 2010. Pop/Rock/Dance at the height of the loudness wars are good examples here.

To test this, I encoded the the sparse opening of Brahms’ Intermezzo Op. 118, No.6. This turns out to be pretty ideal use for the algorithm at high settings. It’s a single piano line with no other pitches. It has only a little processing and a lot of headroom. (The MP3 was encoded using LAME MP3, at 320 kbps, variable bitrate.)

Zooming in so that we can see each sample (there are 44,100 samples in a second), we can see that the MP3 and WAV values are nearly identical. The only difference here is where the MP3 waveform (red) is peaking out over the uncompressed waveform (blue). Where they are identical, the line turns purple. But the whole thing is nearly purple, because it’s nearly identical:

Brahms’ Intermezzo Op. 118, No.6: the uncompressed waveform (blue) is nearly identical to the MP3 waveform (red). Encoded with LAME MP3, at 320 kbps, variable bitrate

Eyes (and waveform rendering) can be deceiving, however. So I flipped the polarity of one track and summed the signals. The result was nearly total silence. A nearly pure flat tone — slight noise at -95 dB, with an occasional blip of -84 dB. No one’s going to hear this—these waveforms are basically identical:

Brahms’ Intermezzo Op. 118, No.6: when working inside a DAW, the MP3 is almost identical to an uncompressed WAV. This is the result when the MP3 is summed with a flipped-phase WAV, they cancel each other out (almost) perfectly. The MP3 was encoded with LAME MP3, at 320 kbps, variable bitrate.

It’s also hard to hear MP3 distortion in tracks that are highly processed and are full of digital-sounding noise. Here, for example, is the crash cymbal in a pop-rock track released in 2012. I couldn’t hear any difference between the uncompressed version and an MP3 encoded at 320 kbps. Here’s the spectrogram compared against the pure, random noise. It’s also compared against a natural, unprocessed crash cymbal:

The crash cymbal in a pop-rock track from 2012 is more similar to random noise than it is to a natural, unprocessed crash cymbal.

The commercial cymbal from 2012 is similar to the random noise. It’s highly uniform, missing all of the horizontal lines in the natural cymbal (the small resonances you hear in a natural body). It ends up sounding (and looking) like digitally-produced noise.

Here’s the waveform of that commercial cymbal. It has a lot of tiny peaks at irregular intervals, resulting in random-sounding noise at high frequencies. There is also little match of these high-frequency peaks between the left and the right ear, making it sound even more random:

A commercial cymbal, with a huge number of tiny peaks at irregular intervals. It ends up sounding like random noise.

Here’s a closer look at the waveform. The smaller peaks are mostly irregular and are basically a jumble of random frequencies:

Commercial cymbal: a closer look at those irregular peaks and their perceived frequencies.

Here is the waveform of a natural, unprocessed crash cymbal. There is still plenty of high-frequency noise, but it’s more regular, matching what the ear expects of a naturally resonating body. The peaks are also more correlated in the Right and Left channels:

A natural, unprocessed crash cymbal: plenty of noise, but more regular peaks than what you see with the commercial cymbal. The peaks are also more correlated in the Right and Left channels. It matches what the ear expects of a resonating body.

As a result, the noise/distortion you get with the MP3 algorithm is almost impossible to hear on this source. When it’s most noticeable, the MP3 algorithm produces a bit of digital noise (often called “the swirlies”), reduces the sense of depth, and shrinks the sound stage. But you can’t really hear this if the source is already packed full of digital noise and has a sound stage to match.

MP3s can be less perfect in other modern pop and dance tracks. These tracks may have plenty of transients and harmonics from vocals, drums, vocal harmonies, instrumental harmonies, synths, sound effects, etc. But they aren’t maxed out in the way that the earlier pop-rock track was, nor do they have effects that are indistinguishable from random noise.

Here’s a spectrogram of this kind of track. The light parts are sound; the dark parts are silence. As you can see, there aren’t many dark parts. But it’s not quite as maxed out as the earlier commercial cymbal, which was similar to random noise:

Compare this with the Brahms Intermezzo:

The MP3 flattens some transients, resulting in a less punchy sound. Here, all the blue peaks have been flattened by the MP3 algorithm:

While making some waveforms slightly louder:

Moreover, the left and right channels are processed separately. The gains & reductions you get in the right channel will be different from those in the left. Not surprisingly, this removes some info that the brain uses to sense position and space. Here’s what happens when you sum a flipped phase WAV vs. MP3:

There’s almost no correlation between the Left and Right channels, resulting a little reduction in sound stage. However, it’s really minor compared to the differences created by the consumer audio playback programs.

Conclusions

My take-aways from these tests:

  • Some consumer audio programs (at least VLC and Windows Media Player) change the sound of MP3s. They add a subtle lo-fi distortion. Luckily, professional DAWs don’t.
  • The MP3 algorithm is extremely faithful to sparse textures, such as a solo piano. (Playing it directly in a consumer audio program, however, will still reveal some minor distortion.) It’s less faithful to full textures with a lot of different pitches and timbres.
  • The MP3 algorithm also does justice to sounds that are already noisy and distorted, especially if that noise/distortion has a digital feel.

I only tip-toed into the shallow end of this pool. But testing a few audio programs and a few sound sources yielded a few big surprises. Other variables — e.g., bitdepth, audio equipment, different MP3 encoders — may yield their own surprises.