How Avery Wang's Shazam Algorithm Identified Songs Without AI
Aakash Gupta explains that Shazam's 2002 system, invented by Stanford PhD Avery Wang, used spectrogram peaks, constellation maps and hash lookups instead of AI, and was published openly in 2003. Apple acquired the company in 2018.
Original post · 2 min read
The inventor was Avery Wang, a Stanford PhD in audio signal processing. His problem was brutal. Match a short clip recorded on a 2002 cell phone mic, in a noisy bar, against a database of a million songs, in seconds, over a phone call.
His solution treated music as geometry instead of sound.
The algorithm converts audio into a spectrogram, a picture of the song, then throws almost all of it away. It keeps only the peaks, the loudest frequency points at each moment in time. Bar chatter and blown-out speakers can wreck most of a recording. The peaks survive. Shazam only ever needed the peaks.
Those surviving points form what Wang called a constellation map, because it looks like a star field. Pairs of peaks get converted into hash numbers, and identifying a song becomes a dictionary lookup rather than an audio comparison. That made it fast enough to search a million tracks on 2002 hardware.
Wang published the full method openly in 2003 in a paper called "An Industrial-Strength Audio Search Algorithm." Anyone could read exactly how the magic worked. The moat was the database and the deals with carriers, never the secret.
Apple bought the company in 2018 for a reported $400 million. People have tagged over 100 billion songs since the very first one, Jeepster by T. Rex, during the beta in April 2002.
One deterministic signal-processing trick, written before most people had heard the phrase machine learning, and it's still so good that in 2026 everyone assumes it must be AI.
Nathan Ruff @TheNathanRuffDude, how did Shazam work 15 years ago without AI!?


