Bart van Merriënboer, Vincent Dumoulin, Jenny Hamer, Lauren Harrell, Andrea Burns, Tom Denton (Google DeepMind · Google Research) — arXiv:2508.04665, August 2025 (v2, January 2026)

The Bittern Lesson

The field kept betting that self-supervised audio models would beat supervised ones, the way they did in vision. Perch 2.0 is a 12-million-parameter convolutional network trained to name 14,795 species from five seconds of sound, and with two tricks, a teacher that lives inside the model and a task that treats every recording as its own class, it beats every self-supervised bioacoustics model on the public benchmarks, including on whales it barely heard during training.

Prerequisites: what a spectrogram is (sound as a picture of frequency over time) + what a classifier and a cross-entropy loss do. Mixup, distillation, prototypes, linear probes and ROC-AUC are all built from zero here.
10
Chapters
11
Simulations
14,795
Classes, one head
0.908
BirdSet ROC-AUC, no fine-tuning

Chapter 0: The Problem

Somewhere in a wetland in New Zealand a microphone has been recording for six months. Nobody has listened to it. Nobody could: it is four thousand hours of wind, rain, frogs, insects, distant traffic, and, on maybe forty of those hours, the deep booming call of an Australasian bittern, one of the most elusive endangered birds on Earth. The ecologist who placed the microphone needs to know on which nights the bittern called, how often, and whether there were two of them. That is the whole field of bioacoustics in one sentence: an ocean of unlabeled sound, a handful of labels, and a question about a rare animal.

Machine learning replaced template matching for this job years ago, and the standard tool is a model trained on labeled bird recordings that either scores every five-second window for thousands of species out of the box, or produces an embedding, a fixed-length vector per window, that a practitioner can cluster, search, or train a tiny classifier on. Perch 1.0 and BirdNET are the models people actually run. This paper is the second Perch, and its title is a joke with a thesis inside it.

The joke: bitterns are herons, and the bittern lesson is a pun on Richard Sutton's essay The Bitter Lesson, which argues that progress in AI comes from simple, general methods that scale rather than from hand-built cleverness. The thesis: in bioacoustics, the simple general method that keeps winning is supervised classification. Not masked autoencoders, not contrastive learning, not the self-supervised recipes that took over vision and language. A plain classifier, trained to name species, with good augmentations and two auxiliary losses, is very hard to beat.

Four Thousand Hours, Forty Labels

A synthetic six-month recording as a strip of five-second windows. Most are noise; a few contain the target call. The ecologist can afford to label only a handful. Drag the label budget and watch what each approach can do with it: a model that scores windows out of the box, an embedding you can search by example, and a linear probe trained on your few labels. The numbers are illustrative; the constraints are the paper's.

Three Constraints the Paper Refuses to Relax

Read the paper's stated design goals and you find a set of constraints that most machine-learning papers would quietly drop. First, the model must be small. Practitioners run these things on laptops over terabytes of audio, so the paper picks EfficientNet-B3, twelve million parameters, and treats that as a feature. Second, the embedding must be frozen. Every evaluation in the paper, including the state-of-the-art results, is done with the embedding model untouched: only a linear classifier or a prototype classifier is trained on top. Embed once, reuse forever. Third, the evaluation must mimic deployment: soundscapes rather than the clean focal recordings the model trained on, tasks other than species identification, and taxa the model never heard, like bats, mosquitoes and whales.

Why be so strict? Because the failure mode in this field is a model that tops a benchmark by fine-tuning on the benchmark's own training split, then falls apart on the next microphone in the next forest. The paper's whole evaluation apparatus, which gets its own chapter, exists to make that failure visible before release.

The mismatch that makes bioacoustics hard, and which the model has to survive: training data is focal: someone pointed a microphone at a bird and uploaded the clip with the species name. Deployment data is a soundscape: a microphone on a tree, recording everything at once, with the target species quiet, distant, and overlapping three others. And the label on a training clip is weak: it says the species is somewhere in a two-minute recording, not in which five seconds. Every design choice in the next chapters is downstream of one of those three facts.

What Was Tried, and Why It Kept Losing

Vision and language moved to self-supervised foundation models: train on unlabeled data with a pretext task, adapt later. Bioacoustics tried to follow. Bird-MAE masks spectrogram patches and reconstructs them. BirdAVES adapts a HuBERT-style speech recipe. SimCLR-style contrastive models learn invariances from augmentation. The authors report that they tried all three families themselves and that none consistently beat a well-trained supervised model, which matches independent comparisons in the literature. Chapter 9 gives their hypotheses for why. For now, the important thing is the shape of the claim: the paper is not saying self-supervision cannot work. It is saying that in this domain, with this much labeled data, supervision is the stronger baseline, and here is how to make it stronger still.

Every headline result in Perch 2.0 is reported with the embedding model frozen, training only a linear or prototype classifier on top. What does that constraint protect against?

Chapter 1: The Key Insight

Strip the paper to its argument and it is this: a harder classification problem makes a better embedding. The authors say it almost in passing, "we generally find that increasing the difficulty of the classification problem increases the overall quality of the embedding model," and then every design decision turns out to be a way of making classification harder in a useful direction.

Count the ways. They expand the label space from birds to 14,795 classes across birds, frogs, insects, mammals and general sound events, so the classifier must separate more things. They mix up to five recordings into one window with a multi-hot target, so the classifier must find every species in a crowd rather than the loudest one. They add a source-prediction head that treats each of the 1.5 million recordings as its own class, the most fine-grained classification problem imaginable. And they add a prototype-based teacher that gives the classifier soft targets richer than one-hot labels. Each one is a way of asking the embedding to carry more information, on the theory that the information a classifier needs to tell a Wilson's warbler from a Hooded warbler is the same information a whale researcher needs to tell an orca ecotype from another.

Granularity Is the Pretraining Signal

The paper's own ablation, Table 7, as a dial. The same Xeno-Canto data, the same architecture, but the training labels collapsed to a coarser level of the taxonomy. Turn the dial from order (41 classes) to species (10,906 classes) and watch transfer performance on the cross-taxa BEANS benchmark rise with the number of classes the model had to tell apart.

The dial is the cleanest evidence in the paper for the insight. Train on Xeno-Canto with species labels and the BEANS transfer accuracy is 0.837. Collapse the same labels to genus, 2,398 classes, and it drops to 0.820. To family, 249 classes: 0.761. To order, 41 classes: 0.608. Nothing about the audio changed. Only how finely the model was forced to carve it. The paper cites vision work showing the same effect, that the gap between self-supervised and supervised models widens as the supervised labels get finer, and it argues that bioacoustics is unusually well placed to exploit it, because species are an exceptionally fine-grained natural label and there are 1.5 million labeled recordings to learn from.

A way to hold the idea: a label is a question the model must be able to answer from the sound. "Is this a bird?" can be answered from a few coarse features. "Is this a Hermit thrush or a Swainson's thrush?" can only be answered by resolving fine structure in the song. An embedding that resolves that structure is, as a side effect, an embedding that can resolve the fine structure in a gibbon's call or a whale's. Coarse questions build coarse embeddings.

The Map of the Paper

PieceWhat it makes harder or betterChapter
Four datasets, 14,795 classes, multiple taxaMore things to tell apart; weak labels handled by window selection2
Generalized mixup, up to 5 sources, multi-hot targetsFind every species in a crowd, not the loudest3
EfficientNet-B3 with three headsSmall, linearly separable embedding; prototype head; source head4
Species loss + self-distillation + source prediction, two phasesSoft targets from an in-model teacher; the finest possible classification task5
Vizier search, 100 models per stage, two stages per phaseHyperparameters chosen on transfer, not on training accuracy6
Nineteen validation datasets, three task types, geometric means, frozen embeddingsModel selection that mimics deployment7
BirdSet, BEANS, marine transferState of the art, including on whales8
In the label-granularity ablation, the audio and the architecture are identical across runs and only the label coarseness changes. What does the steady drop from species to order tell you?

Chapter 2: The Data

Four sources, one shared label space, and a problem hiding in the lengths of the recordings.

SourceBirdsAmphibiansInsectsMammalsOtherTotal recordings
Xeno-Canto (citizen science)860,7012,26031,9711,3230896,255
iNaturalist (research-grade audio via GBIF)480,23051,45030,5359,074409571,698
Tierstimmenarchiv (Berlin museum archive)26,6221,3418604,9924433,859
FSD50K (general sound events)000040,96640,966
Total1,367,55355,05163,36615,38941,4191,542,778

Three of the sources use different taxonomies, so Xeno-Canto and the Berlin archive were mapped by hand onto iNaturalist's species names. Bats were removed entirely, because their echolocation lives above the frequencies the spectrogram front end keeps. The result is 14,795 classes: 14,597 species plus 198 general sound-event classes from FSD50K, which teach the model what a car door and a dog are so that it does not confuse them with wildlife.

Look at the proportions before moving on. Birds are 89% of the data. Mammals are 1%. The paper is honest that this is still mostly a bird model, and Chapter 8's marine results are surprising precisely because the training set contains, in the authors' words, "a few dozen cetacean recordings," most of them phone recordings made above water.

The Weak-Label Problem

Recordings run from under a second to over an hour, mostly five seconds to two and a half minutes. The model consumes exactly five seconds. So the label "Australasian bittern" on a ninety-second recording is weak: the bird is in there somewhere, but which five-second window? Pick the wrong window and you have trained the model to call wind a bittern.

Which Five Seconds?

A synthetic ninety-second recording with the labeled species calling in a few places and noise elsewhere. Two strategies for choosing a training window: uniformly random, or the Perch 1.0 heuristic that finds energy peaks with a wavelet detector, takes a six-second window around the strongest, and samples five seconds inside it. Draw windows repeatedly and count how often each strategy lands on the bird.

The peak detector is worth knowing in detail, because it is the kind of hand-built heuristic the bitter lesson warns about, and the paper ends up not needing it. Build a mel spectrogram with an 80 ms window and 10 ms hop. Denoise it twice: per frequency bin, compute mean and standard deviation of the log magnitudes over time, throw away everything above the mean plus 1.5 standard deviations, recompute, and keep only what lies above the new mean plus 0.75 standard deviations. Sum the surviving magnitudes across frequency. Find peaks between half a second and two seconds wide with a bank of ten wavelet filters. Discard any peak whose 600 ms neighborhood holds less than 1.5 times the recording's mean energy. Keep the top five. If nothing is found, take the first six seconds.

That machinery assumes the labeled species is the loudest thing in the recording. Often true for a focal recording, not always. And here is the paper's first surprise: with the full training recipe, random windows perform on par with peak selection. Contrary to the Perch 1.0 experience and to the literature. The authors' hypothesis is that the self-distillation phase, coming in Chapter 5, absorbs the label noise that random windows introduce, so the heuristic becomes unnecessary. A simpler, more general method won. The bittern lesson, in miniature.

What the data pipeline actually guarantees: every training example is a five-second, 32 kHz, single-channel window with a set of species labels attached, and after mixup (next chapter) possibly several. Everything the model ever sees has that exact shape, which is why the front end can be a fixed function rather than a learned one.
Perch 1.0 relied on energy-peak window selection. Perch 2.0 finds random windows perform equally well. What is the authors' explanation?

Chapter 3: Mixing Sources

A soundscape never contains one animal. A dawn chorus has a dozen species overlapping, and the rare one you care about is under all the others. But almost every training clip is focal: one species, foregrounded. If you train only on those, the model learns to name the loudest sound and stops. The classic fix is mixup: blend two training examples into one and blend their labels the same way. The paper generalizes it, and the generalization is worth deriving because the obvious version is wrong for this problem.

Original Mixup, and Its Flaw Here

Original mixup takes two inputs and two one-hot labels, draws a mixing weight, and forms a weighted sum of both. The audio is seventy percent recording A and thirty percent recording B, and the target says seventy percent species A, thirty percent species B. That makes sense for images, where a blend of a cat and a dog really is a bit of each. It is wrong for sound. If a bittern booms at thirty percent volume under a louder frog, the bittern is still fully present. A detector that reports "thirty percent bittern" has half-failed. Every species in the window should be recognized with high confidence regardless of its loudness.

The Generalization

Three changes. First, mix more than two sources. Draw the number of components from a beta-binomial distribution, then add one, so you always have at least one recording and sometimes as many as five. Second, draw the mixing weights from a symmetric Dirichlet distribution, so they are positive and sum to one, with a concentration parameter that controls whether one source tends to dominate or all are roughly equal. Third, after taking the weighted sum of the waveforms, divide by the square root of the sum of squared weights, so the composite signal has the same overall gain as a single recording. Without that normalization, a five-source mix would be much quieter than a single-source clip, and the model would learn that quiet means crowded.

N ~ BetaBinomial(n, α, β) + 1    w ~ SymmetricDirichlet(N, ω)
xmix = ( w1x1 + … + wNxN ) / √( w1² + … + wN² )    target = multi-hot over all N label sets

And the target is multi-hot, not a weighted average. Every species present in any of the components gets a one. The weights shape the audio; they do not shape the label.

Mixing a Chorus

Five synthetic species, each with its own spectrogram signature. Draw a mix: the number of components comes from the beta-binomial, the weights from the Dirichlet. The left panel shows the mixed spectrogram, the right panel the target vector, first as original mixup would write it (weights) and then as Perch writes it (multi-hot). Move the concentration slider to see how it changes the weights; low concentration lets one source dominate, high concentration spreads them evenly.

Two of the hyperparameters are worth remembering for Chapter 6, because the search found something strange about them. In the first training phase, the best models mixed between two and five sources with a concentration between ten and thirty: crowded windows, fairly even weights. In the second phase, the self-distillation phase, the search turned mixup almost entirely off. The paper's final second-phase configuration has no mixing at all. The augmentation that makes the classifier robust to crowds is wanted while the classifier is learning to hear, and unwanted while it is learning to imitate its teacher.

# generalized mixup for one training window (the paper's Section 2.1, in code)
N = beta_binomial(n, alpha, beta) + 1            # how many recordings to blend, 1..n+1
w = dirichlet([omega] * N)                         # positive weights summing to 1
xs, ys = sample_windows(N)                         # N x [160000] waveforms, N label sets
x = sum(w[i] * xs[i] for i in range(N)) / sqrt(sum(w[i]**2 for i in range(N)))   # gain preserved
y = zeros(14795)
for labels in ys: y[labels] = 1                    # multi-hot: present is present, loud or quiet
The teaching point under the augmentation: a label should encode what is true about the window, not what is prominent in it. Original mixup confuses the two because in images they coincide. In a soundscape they do not, and the multi-hot target is the paper writing that distinction into the loss.
Perch mixes up to five recordings with Dirichlet weights but builds a multi-hot target instead of the weighted-average target of original mixup. Why?

Chapter 4: The Architecture

Three parts: a fixed front end that turns sound into a picture, a convolutional embedding model, and three output heads that exist only during training. Follow the shapes, because the shapes are where the design decisions live.

Five Seconds, Shape by Shape

The paper's Figure 2 with the tensor shapes filled in. Click a stage to see what it does and why. The three heads share one embedding; the stop-gradient on the prototype head is the detail that makes Chapter 5 work.

Click a stage.

Front End: Sound to Picture

Input is 5 seconds of single-channel audio at 32 kHz: 160,000 samples. A short-time Fourier transform with a 20 ms window (640 samples) and a 10 ms hop (320 samples) gives 500 frames. Each frame's magnitudes are pooled onto 128 mel-scaled frequency bins from 60 Hz to 16 kHz, then log-compressed with a floor of one part in a hundred thousand and scaled by 0.1. The result is a 500 by 128 image. The mel scale spaces bins the way a mammalian ear does, densely at low frequencies, sparsely at high, and the 16 kHz ceiling is why bats had to leave the training set: their calls live above it.

Embedding Model: EfficientNet-B3

A convolutional residual network with 12 million parameters, using depthwise convolutions to keep the parameter count low. Perch 1.0 used the smaller B1 at 7.8 million; the step up reflects the larger dataset, and B3 is still tiny by modern standards, which is the point. Its output is a spatial embedding of shape 5 by 3 by 1536: five positions along time, three along frequency, 1,536 features at each. Averaging over the five-by-three grid gives a single mean embedding of 1,536 numbers. That mean embedding is what practitioners download, store, cluster and search. Everything else in this chapter is scaffolding around it.

Three Heads

HeadInputOutputWhy it exists
Linear classifierMean embedding, 153614,795 class logitsForces the embedding to be linearly separable by species, which is exactly what a linear probe later needs
Prototype classifier (ProtoPNet)Spatial embedding, 5×3×153614,795 class scores; four learned prototypes per class, prediction is the max activation across themA second, differently shaped classifier that becomes the teacher in self-distillation; also a strong probe for detection tasks
Source predictorMean embedding, 1536, through a rank-512 projectionOne logit per source recording, over 1.5 millionThe self-supervised source-prediction loss; low rank because a full 1536-by-1.5-million matrix would dwarf the model

The prototype head deserves a picture in your mind. Instead of one weight vector per class, it learns four prototypes per class, each a vector in the 1,536-dimensional feature space. To score a class, it compares every one of the fifteen spatial positions against each of the four prototypes and takes the maximum activation. A prototype can therefore latch onto a species' characteristic sound occurring anywhere in the window, at any of the five time positions and three frequency bands, which is why the paper finds prototype probes beat linear probes on detection tasks. An orthogonality loss pushes the four prototypes of a class apart so they capture different aspects of the call rather than the same one four times.

Read the low-rank source head as a budget decision. The source-prediction task needs a classifier over 1.5 million classes. A dense 1536-by-1,542,778 weight matrix is about 2.4 billion parameters, two hundred times the embedding model. Projecting the embedding to 512 dimensions first cuts the head to about 790 million, still large but stored only during training and never shipped. The head exists to shape the embedding, not to be used.
# forward pass, shapes as the paper gives them
x = audio                                  # [160000]         5 s at 32 kHz, mono
S = log_mel(x, win=640, hop=320, mels=128)   # [500, 128]       fixed function, no parameters
E_S = efficientnet_b3(S)                    # [5, 3, 1536]     spatial embedding
E_A = E_S.mean(axis=(0, 1))                  # [1536]           mean embedding (what users keep)
logits = W_lin @ E_A                        # [14795]          linear head
proto  = max_over_positions_and_prototypes(stop_gradient(E_S))   # [14795]  4 prototypes per class
source = W_src @ (P_512 @ E_A)              # [1542778]        rank-512 projection first
The prototype head reads the 5×3×1536 spatial embedding while the linear head reads the 1536-dimensional mean. What does the prototype head gain from the spatial input?

Chapter 5: Three Losses

The model is trained with three objectives and in two phases, and the interplay between them is the paper's method. Take them one at a time, then watch them run together.

Loss 1: Species Cross-Entropy, With a Twist

The linear head is trained with softmax and cross-entropy. That is unremarkable except for one choice. A window can contain several species, so the natural loss is a sigmoid per class with binary cross-entropy: each species present or absent, independently. The paper does not do that. Following a large-scale weakly-supervised vision result, it uses a softmax over all 14,795 classes with a target that spreads probability evenly across the k species present, one over k each. The authors found this trains faster than the traditional multi-label sigmoid. The intuition: softmax makes classes compete, so pushing up the present species automatically pushes down the 14,790 absent ones, a much stronger gradient signal per example than 14,795 independent yes-or-no decisions.

Loss 2: Self-Distillation, With a Stop-Gradient

Distillation normally needs a bigger teacher. Self-distillation uses the model itself, and the paper's version has an architectural elegance to it. The prototype classifier is trained for species classification with its own softmax cross-entropy plus the orthogonality loss. Its predictions are then used as soft targets for the linear classifier: instead of learning "this window is one hundred percent Hermit thrush," the linear head learns "eighty percent Hermit thrush, fifteen percent Swainson's thrush, a little Veery," the confusion structure the prototype head has discovered. Teacher and student share the same embedding model.

The critical detail is a stop-gradient between the embedding model and the prototype head. The prototype head reads the spatial embedding but never sends gradients back into it. Only the linear head and the source head shape the embedding. Why? Because if the teacher could reshape the embedding to make its own job easier, teacher and student would collapse toward each other and the soft targets would carry no new information. Freezing the teacher's influence on the shared representation keeps the two views genuinely different, which is what makes the distillation worth anything.

The Self-Distillation Loop

A toy with six species. Watch the two-phase schedule run. In phase one the prototype head learns from hard labels and the linear head learns from hard labels; the prototype head's soft predictions sharpen over time but are not used. In phase two the linear head's target switches to the prototype head's soft distribution. Toggle the stop-gradient off to see what happens when the teacher is allowed to reshape the shared embedding: the two heads collapse toward agreement and the soft targets lose their structure.

Loss 3: Source Prediction

The third objective is called DIET in the vision paper that proposed it, and it is almost embarrassingly simple: give every recording in the dataset its own class, and train a classifier to predict which recording a window came from. There are over 1.5 million recordings, so this is a 1.5-million-way classification problem, and it is self-supervised in the sense that no human label is involved: the identity of the recording is the label.

What makes it work is data augmentation, and in audio the augmentation is free. A long recording yields many non-overlapping five-second windows, and the model must map all of them to the same source. To do that it has to learn what is stable across a recording: the individual bird's voice, the microphone, the habitat's background, the reverberation of that particular clearing. Those are exactly the fine-grained features that a species label alone would never force it to keep, and the paper notes that a similar objective has been used to identify individual animals. The authors also offer a second reading: source prediction is just an extremely fine-grained supervised classification problem, which is the Chapter 1 insight taken to its limit.

Two Phases

Phase 1 · up to 300,000 steps
Linear head, prototype head and source head all train on hard labels. Mixup on (2 to 5 sources), source-prediction weight around 0.1, dropout about 0.5. The prototype head's predictions are computed but not yet used as targets. Output: a strong classifier and a trained teacher.
↓
Phase 2 · up to 400,000 steps, self-distillation
Continue from the best phase-1 model. The linear head now learns from the prototype head's soft predictions, weighted about 4 (heavier than the species loss itself). Mixup off, dropout 0, source prediction off, learning rate two hundred times smaller. The model is polishing, not exploring.

Look at the phase-2 settings again. Every source of noise that phase one wanted, mixed windows, dropout, the source-prediction pressure, is switched off. The only signal is the teacher. And the teacher, remember, was trained with all that noise. Phase two is the model distilling into its own linear head everything its noisier, prototype-shaped self learned, at a learning rate small enough not to disturb it. This is also the authors' explanation for why random windows suddenly work: a soft target from a teacher that has seen the whole distribution is far more forgiving of a window that missed the bird than a hard label is.

# one training step, both phases (the paper's Section 2.3)
E_S, E_A = embed(mixed_window)
L_species = softmax_xent(W_lin @ E_A, target=uniform_over_present(y))       # 1/k on each present class
p_proto   = proto_head(stop_gradient(E_S))                                # teacher never shapes E
L_proto   = softmax_xent(p_proto, target=uniform_over_present(y)) + orthogonality(prototypes)
L_source  = softmax_xent(W_src @ (P @ E_A), target=source_id)             # 1.5M-way
L_distill = softmax_xent(W_lin @ E_A, target=softmax(p_proto)) if phase == 2 else 0
loss = L_species + L_proto + lam_src * L_source + lam_distill * L_distill      # lam_src 0.11 then 0; lam_distill 0 then 4.22
Why three losses and not one: each pulls the embedding toward a different kind of separability. Species loss: linearly separable by class. Source loss: separable by individual recording, which forces fine detail. Distillation: shaped by a second classifier's view of the confusions. The stop-gradient is what stops the second and third from collapsing into the first.
The prototype teacher reads the embedding through a stop-gradient. What would go wrong without it?

Chapter 6: The Search

A recipe with this many knobs, mixup's four parameters, three loss weights, dropout, learning rate, cannot be tuned by hand, and it must not be tuned on training accuracy, because the whole point is transfer. So the paper runs a black-box optimizer, Vizier, against the validation suite of Chapter 7, and the settings it chose are themselves evidence about what the method is doing.

How the Search Ran

Each phase gets two stages of one hundred models. Stage one trains a hundred configurations; based on their validation scores Vizier proposes a hundred more; the best is kept. Each model took twenty to thirty hours on an eight-chip TPU pod, the range set mostly by how many mixup sources it blended. That is roughly four hundred training runs for the two phases, and the paper ran the whole search twice, once with random windows and once with peak selection. Phase one searches learning rate, dropout on the embedding before the heads, the source-prediction loss weight, and the mixup parameters. Phase two adds the self-distillation weight.

What the Optimizer Preferred

The ranges Vizier converged on for each phase, with the final chosen values marked. The two phases wanted opposite things, and reading the difference is the fastest way to understand the method. Hover or tap a knob for the paper's reading.

KnobPhase 1 preferred rangePhase 1 finalPhase 2 preferred rangePhase 2 final
Mixup sources2 to 5n = 2 (up to 3 sources)mostly 1 (no mixing)none
Mixup concentration ω10 to 301––
Beta-binomial α, βα > β91.3, 100––
Source-prediction weight0.1 to 0.90.110 to 0.40
Dropout0.3 to 0.60.490 to 0.50
Self-distillation weight––1.5 to 4.54.22
Learning rate–6.41 × 10−4small3.20 × 10−6

Two things stand out. The source-prediction loss has a much larger scale than the species loss, because it is a 1.5-million-way softmax, so a weight of 0.11 is not "barely on"; the paper footnotes exactly this. And the phase-two learning rate is two hundred times smaller than phase one's. The optimizer, given the freedom, chose to make phase two a gentle polishing pass with a single strong signal, the teacher, and nothing else. Nobody designed that schedule. The validation suite selected it.

What makes this search trustworthy: it is scored on frozen-embedding transfer to nineteen datasets, most of which are not bird classification, using a geometric mean that punishes any single collapse. A configuration that overfits the training species would score badly. So the preferences above are preferences for generalization, which is why the phase-two choices, no augmentation, tiny learning rate, are worth taking seriously rather than reading as a quirk.
Phase two's chosen settings turn off mixup, dropout and source prediction and cut the learning rate by two hundred times. What is the model doing in that phase?

Chapter 7: Measuring Transfer

The validation design is the part of this paper most worth stealing for other fields. It answers a question every foundation-model team faces: how do you choose a model when the tasks it will be used for do not exist yet?

Three Kinds of Use, Three Kinds of Score

The authors ask how practitioners actually use Perch, and they find three patterns. Some run the pretrained classifier out of the box. Some embed everything and search by example: find me every window that sounds like this one. Some train a small classifier on a few dozen labels. So validation has three task types, each scored with ROC-AUC, the probability that a random positive outscores a random negative, which is stable across class imbalance in a way accuracy is not.

Task typeProcedureDatasets
Pretrained classificationScore every 5 s window with a 2.5 s stride using the pretrained head; a prediction counts if it overlaps an annotationPowdermill and Caples: two fully annotated soundscapes whose species are a subset of the training classes
One-shot retrievalPick a random example of a species, rank all windows by cosine similarity to it, compute ROC-AUC treating every other species as negative; a proxy for search and clusteringBEANS detection sets, Weldy call types, Powdermill, Caples
Linear transferSixteen random examples per class, each embedded as the average over its 5 s windows; train a logistic-regression probe for 10,000 steps; ROC-AUC on the restBEANS classification sets, Godwit calls, Yellowhammer dialects, bats, and three marine sets: DCLDE, NOAA, ReefSet

Within each type the datasets are combined with a geometric mean, and the three type scores are combined with another geometric mean. That choice is deliberate and worth understanding.

Why the Geometric Mean

Two candidate models, each scored on six datasets. Drag any bar. The arithmetic mean rewards a model that is brilliant on five datasets and collapses on one; the geometric mean does not. Try dropping one score toward zero and watch the two means diverge. That gap is why model selection here favors consistency over peaks.

A geometric mean of six numbers is the sixth root of their product. If any one of them goes to zero, so does the product. So a model that scores 0.95 on five bird datasets and 0.40 on the bat dataset loses to a model scoring 0.88 everywhere. The authors' earlier evaluation paper argued this is the right bias for bioacoustics: a practitioner picking up the model for a new taxon needs it to not fail, more than they need it to excel on birds. Nineteen datasets in total feed the validation score, and the embedding is frozen throughout.

Test Benchmarks

Validation chooses the model; two public benchmarks test it, and neither was used for selection. BirdSet has six fully annotated soundscape datasets from the continental United States, Hawai'i, Peru and Colombia. It provides training splits for fine-tuning, and the paper pointedly does not use them: it applies the prototype classifier's predictions directly. BEANS, the Benchmark of Animal Sounds, has twelve cross-taxa tasks covering birds, land and marine mammals, frogs and insects, split into classification and detection. There, the paper trains linear and prototype probes on each task's training split with the embedding frozen. The two together cover the same kinds of domain shift the validation suite was designed around, while allowing a fair comparison with published models.

The distinction to keep straight: a linear probe is a logistic regression on the 1,536-dimensional mean embedding. A prototypical probe is the Chapter 4 prototype head, trained fresh on the task's data over the frozen spatial embedding. Both leave the twelve million embedding parameters untouched. "Fine-tuning," which competing models use and Perch does not, updates those parameters on the benchmark's own training data.
Why combine the nineteen validation datasets with a geometric mean rather than an arithmetic mean?

Chapter 8: Results

Three tables carry the paper's claims. The first compares against published bioacoustics models on both benchmarks; the second is the label-granularity ablation you met in Chapter 1; the third is the marine transfer result, the one that surprised the authors.

The Benchmark Board

Table 3 as bars. Pick a metric. Methods are labeled by how they touch the benchmark: Pre applies a pretrained head, LP and PP are linear and prototypical probes on a frozen embedding, FT fine-tunes the whole model on the benchmark's training split. Every Perch 2.0 row is Pre, LP or PP. Watch which competitors needed FT to get close.

MethodHowBirdSet AUROCBirdSet cmAPBirdSet top-1BEANS acc.BEANS mAP
Audio ProtoPNet-5Pre0.8960.4230.623––
BirdMAE-LFT0.8860.4400.601––
AVES-BioFT–––0.8170.398
BioLingualFT––––0.479
Perch 1.0Pre / LP0.8390.3560.6130.8090.353
Perch 2.0, phase 1 onlyPre / LP / PP0.9020.4310.6420.8350.426 / 0.499
Perch 2.0, peak-selected windowsPre / LP / PP0.9070.4300.6190.839 / 0.8410.426 / 0.504
Perch 2.0, random windowsPre / LP / PP0.9080.4310.6650.838 / 0.8400.415 / 0.502

Read it in three passes. First, the headline: on BirdSet ROC-AUC, the metric the authors argue is the most stable and informative, Perch 2.0 reaches 0.908 with no fine-tuning, against 0.896 for the best prior model and 0.886 for BirdMAE-L, which was fine-tuned on BirdSet's own training splits and whose score is additionally inflated by having tuned hyperparameters on one of the test subsets. On BEANS, the prototype probe's 0.504 mAP beats BioLingual's fine-tuned 0.479. Second, the honest column: on BirdSet cmAP, BirdMAE-L's fine-tuned 0.440 edges Perch's 0.431. The paper wins the metric it argues for and loses one it argues against, and reports both. Third, the two comparisons inside the Perch rows: random windows match peak selection, and the prototype probe beats the linear probe on detection by a wide margin (0.502 versus 0.415), which is why the prototype head, originally added as a teacher, is also recommended as a probe.

The Marine Result

Perch 2.0's training data has almost no underwater audio: a few dozen whale recordings, mostly made on phones above the surface. Yet on few-shot transfer to marine datasets, sixteen labeled examples per class and a linear probe on the frozen embedding, it beats two models built specifically for the sea.

Model (16-shot linear probe, ROC-AUC)DCLDE speciesDCLDE ecotypeDCLDE known speciesNOAA whalesReefSet
Multispecies Whale (Google)0.9140.8210.9540.917*0.855
Surf Perch (trained on ReefSet)0.9470.9030.9840.8990.986*
Perch 1.00.9680.9310.9810.9050.970
Perch 2.00.9770.9450.9890.9240.981

The asterisks matter: the whale model was trained on much of the NOAA data and Surf Perch was trained on ReefSet, so those two cells are in-distribution for the specialists and still only barely ahead. On orca ecotypes, telling apart locally adapted killer-whale populations by their calls, a bird model with a sixteen-shot probe scores 0.945 against the whale model's 0.821. And there is a sharp lesson in one more number the appendix gives: applying the whale model's own pretrained classifier directly to DCLDE scores 0.612, while a probe on that same model's embeddings scores 0.954. The embedding knew far more than the head exposed.

Why would a bird model hear whales? The authors offer two hypotheses. Birdsong is extraordinarily diverse, so an embedding that resolves fourteen thousand kinds of it has covered an enormous range of acoustic structure. And there are universal mechanisms of sound production shared across terrestrial vertebrates, so the features that separate two thrushes are, physically, close to the features that separate two whales. The bittern lesson has a second clause: the fine-grained supervised task did not just beat self-supervision on birds; it produced an embedding general enough to cross the waterline.
The Multispecies Whale model's own classifier scores 0.612 ROC-AUC on the DCLDE species task, but a sixteen-shot linear probe on its embeddings scores 0.954. What does that gap demonstrate?

Chapter 9: Why Supervision

The discussion section is titled "Why supervision?" and it is the most useful two pages in the paper for anyone outside bioacoustics, because it tries to explain a result that contradicts the field's expectations rather than just report it. Four hypotheses.

Hypothesis 1: There Is Not Enough Unlabeled Data Yet

The self-supervised models that work in vision are trained on enormous corpora and are themselves enormous. DINOv2 saw 142 million images. Xeno-Canto and iNaturalist together are two orders of magnitude smaller. Self-supervision's advantage is that it can use unlabeled data, and bioacoustics does not yet have a hundred million diverse unlabeled recordings organized for training. Meanwhile, vision results suggest that once a supervised model has hundreds of thousands of labeled examples, it becomes very hard to beat, and Perch has 1.5 million.

Hypothesis 2: The Augmentations Are Domain-Specific

Self-supervised objectives depend on data augmentations that encode what should not matter, cropping, color jitter, flipping, and performance is highly sensitive to choosing them well. Those were tuned for photographs over a decade. Nobody knows the right invariances for a spectrogram of a frog yet. The authors note that general audio benchmarks are still dominated by supervised and semi-supervised models, and that methods developed for vision have transferred to bioacoustics worse than hoped in their own earlier work on domain adaptation.

Hypothesis 3: Fine-Grained Labels Are Unusually Powerful Here

This is the Chapter 1 argument with its citations. The gap between self-supervised and supervised models widens as the supervised labels get finer. Species labels are about as fine as natural labels get: over fifteen thousand classes, with the hardest distinctions between species in the same genus. Table 7 showed the effect directly, transfer degrading monotonically as labels were coarsened to genus, family and order.

Where the Signal Goes

A picture of the four hypotheses as pressures on one embedding. Each pressure, when active, spreads the synthetic classes apart along a different axis; remove a pressure and its axis collapses. Toggle them and watch a probe's separability on a held-out task, the number at the bottom, respond. Illustrative, but the ordering follows the paper's ablations.

Hypothesis 4: Birdsong Is Diverse, and Vocal Physics Is Shared

The transfer to bats, mosquitoes, gibbons and whales is the part that looks least likely on paper. The authors' explanation is that birdsong spans a huge range of acoustic structure, and that the physical mechanisms of sound production are shared across terrestrial vertebrates, so a model that has learned to resolve the full variety of birdsong has learned much of what any vocalization can do.

Limits and Future Work

Stated limitImplication
Still 89% bird data; mammals 1%; bats excluded by the 16 kHz ceilingThe taxa with the fewest labels are where the model is weakest, and the paper proposes source prediction as the semi-supervised route to fill those gaps with unlabeled recordings
Evaluation still does not fully match deploymentThe authors say more work is needed on benchmarks that reflect real use; the geometric-mean suite is their best current answer, not their final one
Only species labels are used as supervisionRecordings carry metadata such as time of day, season and location; the success of source prediction suggests those could become additional classification tasks
One retrieval result per species, sixteen-shot probesPractical for practitioners, but the numbers describe a specific low-label regime, not full fine-tuning

The Cheat Sheet

ThingOne lineNumber
DataXeno-Canto + iNaturalist + Tierstimmenarchiv + FSD50K, one taxonomy, bats removed1,542,778 recordings, 14,795 classes
Window5 s at 32 kHz; random or energy-peak selection; random works as well after distillation160,000 samples → 500×128 log-mel
MixupN ~ BetaBinomial + 1 sources, Dirichlet weights, gain-normalized sum, multi-hot target2 to 5 sources in phase 1, none in phase 2
ModelEfficientNet-B3 → spatial 5×3×1536 → mean 153612M parameters
HeadsLinear (14,795), ProtoPNet (4 prototypes per class, max), source (rank-512, 1.5M classes)Training only
LossesSoftmax species CE with 1/k targets; self-distillation from the stop-gradient prototype teacher; source prediction (DIET)Phase 2 distill weight 4.22
SearchVizier, 2 stages × 100 models per phase, 20 to 30 h per model on TPUv3-8Scored on frozen-embedding transfer
ValidationClassify, retrieve, linear-transfer; geometric means; 19 datasetsEmbedding frozen
ResultsBirdSet AUROC 0.908 (prior best 0.896); BEANS mAP 0.504 with a prototype probe; DCLDE species 0.977 vs 0.947 for Surf PerchNo fine-tuning anywhere
GranularitySpecies 0.837 → genus 0.820 → family 0.761 → order 0.608 BEANS accuracyFiner labels, better transfer

Connections on This Site

The spectrogram front end and why the mel scale exists are in Audio Representations. The self-supervised competitors Perch beats are worked through in the Audio-MAE Veanor and in Classical Audio Classification. Distillation in general, and why a teacher's soft targets carry more than a hard label, is Knowledge Distillation. The contrastive alternative this paper argues against is CLAP, and the transformer alternative to EfficientNet for spectrograms is the Audio Spectrogram Transformer. For the other half of the argument, a case where the same lab does put a reasoning layer above frozen embeddings, see the Earth AI Veanor, whose Population Dynamics and AlphaEarth embeddings are used with exactly the frozen-embedding, cheap-probe philosophy this paper defends.

What to Try Yourself

Perch 2.0 is released, and the whole evaluation philosophy is reproducible on a laptop: download the model, embed a public soundscape dataset once, and train a logistic-regression probe on sixteen examples per class. Then re-run with the labels collapsed to genus and family and watch the probe's ROC-AUC fall. If it falls the way Table 7 says, you have reproduced the bittern lesson's central mechanism in an afternoon. If it does not, you have found something the authors would want to hear about.

Which of the authors' hypotheses would be most directly tested by the arrival of a hundred-million-recording unlabeled bioacoustics corpus?