The field kept betting that self-supervised audio models would beat supervised ones, the way they did in vision. Perch 2.0 is a 12-million-parameter convolutional network trained to name 14,795 species from five seconds of sound, and with two tricks, a teacher that lives inside the model and a task that treats every recording as its own class, it beats every self-supervised bioacoustics model on the public benchmarks, including on whales it barely heard during training.
Somewhere in a wetland in New Zealand a microphone has been recording for six months. Nobody has listened to it. Nobody could: it is four thousand hours of wind, rain, frogs, insects, distant traffic, and, on maybe forty of those hours, the deep booming call of an Australasian bittern, one of the most elusive endangered birds on Earth. The ecologist who placed the microphone needs to know on which nights the bittern called, how often, and whether there were two of them. That is the whole field of bioacoustics in one sentence: an ocean of unlabeled sound, a handful of labels, and a question about a rare animal.
Machine learning replaced template matching for this job years ago, and the standard tool is a model trained on labeled bird recordings that either scores every five-second window for thousands of species out of the box, or produces an embedding, a fixed-length vector per window, that a practitioner can cluster, search, or train a tiny classifier on. Perch 1.0 and BirdNET are the models people actually run. This paper is the second Perch, and its title is a joke with a thesis inside it.
The joke: bitterns are herons, and the bittern lesson is a pun on Richard Sutton's essay The Bitter Lesson, which argues that progress in AI comes from simple, general methods that scale rather than from hand-built cleverness. The thesis: in bioacoustics, the simple general method that keeps winning is supervised classification. Not masked autoencoders, not contrastive learning, not the self-supervised recipes that took over vision and language. A plain classifier, trained to name species, with good augmentations and two auxiliary losses, is very hard to beat.
A synthetic six-month recording as a strip of five-second windows. Most are noise; a few contain the target call. The ecologist can afford to label only a handful. Drag the label budget and watch what each approach can do with it: a model that scores windows out of the box, an embedding you can search by example, and a linear probe trained on your few labels. The numbers are illustrative; the constraints are the paper's.
Read the paper's stated design goals and you find a set of constraints that most machine-learning papers would quietly drop. First, the model must be small. Practitioners run these things on laptops over terabytes of audio, so the paper picks EfficientNet-B3, twelve million parameters, and treats that as a feature. Second, the embedding must be frozen. Every evaluation in the paper, including the state-of-the-art results, is done with the embedding model untouched: only a linear classifier or a prototype classifier is trained on top. Embed once, reuse forever. Third, the evaluation must mimic deployment: soundscapes rather than the clean focal recordings the model trained on, tasks other than species identification, and taxa the model never heard, like bats, mosquitoes and whales.
Why be so strict? Because the failure mode in this field is a model that tops a benchmark by fine-tuning on the benchmark's own training split, then falls apart on the next microphone in the next forest. The paper's whole evaluation apparatus, which gets its own chapter, exists to make that failure visible before release.
Vision and language moved to self-supervised foundation models: train on unlabeled data with a pretext task, adapt later. Bioacoustics tried to follow. Bird-MAE masks spectrogram patches and reconstructs them. BirdAVES adapts a HuBERT-style speech recipe. SimCLR-style contrastive models learn invariances from augmentation. The authors report that they tried all three families themselves and that none consistently beat a well-trained supervised model, which matches independent comparisons in the literature. Chapter 9 gives their hypotheses for why. For now, the important thing is the shape of the claim: the paper is not saying self-supervision cannot work. It is saying that in this domain, with this much labeled data, supervision is the stronger baseline, and here is how to make it stronger still.
Strip the paper to its argument and it is this: a harder classification problem makes a better embedding. The authors say it almost in passing, "we generally find that increasing the difficulty of the classification problem increases the overall quality of the embedding model," and then every design decision turns out to be a way of making classification harder in a useful direction.
Count the ways. They expand the label space from birds to 14,795 classes across birds, frogs, insects, mammals and general sound events, so the classifier must separate more things. They mix up to five recordings into one window with a multi-hot target, so the classifier must find every species in a crowd rather than the loudest one. They add a source-prediction head that treats each of the 1.5 million recordings as its own class, the most fine-grained classification problem imaginable. And they add a prototype-based teacher that gives the classifier soft targets richer than one-hot labels. Each one is a way of asking the embedding to carry more information, on the theory that the information a classifier needs to tell a Wilson's warbler from a Hooded warbler is the same information a whale researcher needs to tell an orca ecotype from another.
The paper's own ablation, Table 7, as a dial. The same Xeno-Canto data, the same architecture, but the training labels collapsed to a coarser level of the taxonomy. Turn the dial from order (41 classes) to species (10,906 classes) and watch transfer performance on the cross-taxa BEANS benchmark rise with the number of classes the model had to tell apart.
The dial is the cleanest evidence in the paper for the insight. Train on Xeno-Canto with species labels and the BEANS transfer accuracy is 0.837. Collapse the same labels to genus, 2,398 classes, and it drops to 0.820. To family, 249 classes: 0.761. To order, 41 classes: 0.608. Nothing about the audio changed. Only how finely the model was forced to carve it. The paper cites vision work showing the same effect, that the gap between self-supervised and supervised models widens as the supervised labels get finer, and it argues that bioacoustics is unusually well placed to exploit it, because species are an exceptionally fine-grained natural label and there are 1.5 million labeled recordings to learn from.
| Piece | What it makes harder or better | Chapter |
|---|---|---|
| Four datasets, 14,795 classes, multiple taxa | More things to tell apart; weak labels handled by window selection | 2 |
| Generalized mixup, up to 5 sources, multi-hot targets | Find every species in a crowd, not the loudest | 3 |
| EfficientNet-B3 with three heads | Small, linearly separable embedding; prototype head; source head | 4 |
| Species loss + self-distillation + source prediction, two phases | Soft targets from an in-model teacher; the finest possible classification task | 5 |
| Vizier search, 100 models per stage, two stages per phase | Hyperparameters chosen on transfer, not on training accuracy | 6 |
| Nineteen validation datasets, three task types, geometric means, frozen embeddings | Model selection that mimics deployment | 7 |
| BirdSet, BEANS, marine transfer | State of the art, including on whales | 8 |
Four sources, one shared label space, and a problem hiding in the lengths of the recordings.
| Source | Birds | Amphibians | Insects | Mammals | Other | Total recordings |
|---|---|---|---|---|---|---|
| Xeno-Canto (citizen science) | 860,701 | 2,260 | 31,971 | 1,323 | 0 | 896,255 |
| iNaturalist (research-grade audio via GBIF) | 480,230 | 51,450 | 30,535 | 9,074 | 409 | 571,698 |
| Tierstimmenarchiv (Berlin museum archive) | 26,622 | 1,341 | 860 | 4,992 | 44 | 33,859 |
| FSD50K (general sound events) | 0 | 0 | 0 | 0 | 40,966 | 40,966 |
| Total | 1,367,553 | 55,051 | 63,366 | 15,389 | 41,419 | 1,542,778 |
Three of the sources use different taxonomies, so Xeno-Canto and the Berlin archive were mapped by hand onto iNaturalist's species names. Bats were removed entirely, because their echolocation lives above the frequencies the spectrogram front end keeps. The result is 14,795 classes: 14,597 species plus 198 general sound-event classes from FSD50K, which teach the model what a car door and a dog are so that it does not confuse them with wildlife.
Look at the proportions before moving on. Birds are 89% of the data. Mammals are 1%. The paper is honest that this is still mostly a bird model, and Chapter 8's marine results are surprising precisely because the training set contains, in the authors' words, "a few dozen cetacean recordings," most of them phone recordings made above water.
Recordings run from under a second to over an hour, mostly five seconds to two and a half minutes. The model consumes exactly five seconds. So the label "Australasian bittern" on a ninety-second recording is weak: the bird is in there somewhere, but which five-second window? Pick the wrong window and you have trained the model to call wind a bittern.
A synthetic ninety-second recording with the labeled species calling in a few places and noise elsewhere. Two strategies for choosing a training window: uniformly random, or the Perch 1.0 heuristic that finds energy peaks with a wavelet detector, takes a six-second window around the strongest, and samples five seconds inside it. Draw windows repeatedly and count how often each strategy lands on the bird.
The peak detector is worth knowing in detail, because it is the kind of hand-built heuristic the bitter lesson warns about, and the paper ends up not needing it. Build a mel spectrogram with an 80 ms window and 10 ms hop. Denoise it twice: per frequency bin, compute mean and standard deviation of the log magnitudes over time, throw away everything above the mean plus 1.5 standard deviations, recompute, and keep only what lies above the new mean plus 0.75 standard deviations. Sum the surviving magnitudes across frequency. Find peaks between half a second and two seconds wide with a bank of ten wavelet filters. Discard any peak whose 600 ms neighborhood holds less than 1.5 times the recording's mean energy. Keep the top five. If nothing is found, take the first six seconds.
That machinery assumes the labeled species is the loudest thing in the recording. Often true for a focal recording, not always. And here is the paper's first surprise: with the full training recipe, random windows perform on par with peak selection. Contrary to the Perch 1.0 experience and to the literature. The authors' hypothesis is that the self-distillation phase, coming in Chapter 5, absorbs the label noise that random windows introduce, so the heuristic becomes unnecessary. A simpler, more general method won. The bittern lesson, in miniature.
A soundscape never contains one animal. A dawn chorus has a dozen species overlapping, and the rare one you care about is under all the others. But almost every training clip is focal: one species, foregrounded. If you train only on those, the model learns to name the loudest sound and stops. The classic fix is mixup: blend two training examples into one and blend their labels the same way. The paper generalizes it, and the generalization is worth deriving because the obvious version is wrong for this problem.
Original mixup takes two inputs and two one-hot labels, draws a mixing weight, and forms a weighted sum of both. The audio is seventy percent recording A and thirty percent recording B, and the target says seventy percent species A, thirty percent species B. That makes sense for images, where a blend of a cat and a dog really is a bit of each. It is wrong for sound. If a bittern booms at thirty percent volume under a louder frog, the bittern is still fully present. A detector that reports "thirty percent bittern" has half-failed. Every species in the window should be recognized with high confidence regardless of its loudness.
Three changes. First, mix more than two sources. Draw the number of components from a beta-binomial distribution, then add one, so you always have at least one recording and sometimes as many as five. Second, draw the mixing weights from a symmetric Dirichlet distribution, so they are positive and sum to one, with a concentration parameter that controls whether one source tends to dominate or all are roughly equal. Third, after taking the weighted sum of the waveforms, divide by the square root of the sum of squared weights, so the composite signal has the same overall gain as a single recording. Without that normalization, a five-source mix would be much quieter than a single-source clip, and the model would learn that quiet means crowded.
And the target is multi-hot, not a weighted average. Every species present in any of the components gets a one. The weights shape the audio; they do not shape the label.
Five synthetic species, each with its own spectrogram signature. Draw a mix: the number of components comes from the beta-binomial, the weights from the Dirichlet. The left panel shows the mixed spectrogram, the right panel the target vector, first as original mixup would write it (weights) and then as Perch writes it (multi-hot). Move the concentration slider to see how it changes the weights; low concentration lets one source dominate, high concentration spreads them evenly.
Two of the hyperparameters are worth remembering for Chapter 6, because the search found something strange about them. In the first training phase, the best models mixed between two and five sources with a concentration between ten and thirty: crowded windows, fairly even weights. In the second phase, the self-distillation phase, the search turned mixup almost entirely off. The paper's final second-phase configuration has no mixing at all. The augmentation that makes the classifier robust to crowds is wanted while the classifier is learning to hear, and unwanted while it is learning to imitate its teacher.
# generalized mixup for one training window (the paper's Section 2.1, in code) N = beta_binomial(n, alpha, beta) + 1 # how many recordings to blend, 1..n+1 w = dirichlet([omega] * N) # positive weights summing to 1 xs, ys = sample_windows(N) # N x [160000] waveforms, N label sets x = sum(w[i] * xs[i] for i in range(N)) / sqrt(sum(w[i]**2 for i in range(N))) # gain preserved y = zeros(14795) for labels in ys: y[labels] = 1 # multi-hot: present is present, loud or quiet
Three parts: a fixed front end that turns sound into a picture, a convolutional embedding model, and three output heads that exist only during training. Follow the shapes, because the shapes are where the design decisions live.
The paper's Figure 2 with the tensor shapes filled in. Click a stage to see what it does and why. The three heads share one embedding; the stop-gradient on the prototype head is the detail that makes Chapter 5 work.
Input is 5 seconds of single-channel audio at 32 kHz: 160,000 samples. A short-time Fourier transform with a 20 ms window (640 samples) and a 10 ms hop (320 samples) gives 500 frames. Each frame's magnitudes are pooled onto 128 mel-scaled frequency bins from 60 Hz to 16 kHz, then log-compressed with a floor of one part in a hundred thousand and scaled by 0.1. The result is a 500 by 128 image. The mel scale spaces bins the way a mammalian ear does, densely at low frequencies, sparsely at high, and the 16 kHz ceiling is why bats had to leave the training set: their calls live above it.
A convolutional residual network with 12 million parameters, using depthwise convolutions to keep the parameter count low. Perch 1.0 used the smaller B1 at 7.8 million; the step up reflects the larger dataset, and B3 is still tiny by modern standards, which is the point. Its output is a spatial embedding of shape 5 by 3 by 1536: five positions along time, three along frequency, 1,536 features at each. Averaging over the five-by-three grid gives a single mean embedding of 1,536 numbers. That mean embedding is what practitioners download, store, cluster and search. Everything else in this chapter is scaffolding around it.
| Head | Input | Output | Why it exists |
|---|---|---|---|
| Linear classifier | Mean embedding, 1536 | 14,795 class logits | Forces the embedding to be linearly separable by species, which is exactly what a linear probe later needs |
| Prototype classifier (ProtoPNet) | Spatial embedding, 5×3×1536 | 14,795 class scores; four learned prototypes per class, prediction is the max activation across them | A second, differently shaped classifier that becomes the teacher in self-distillation; also a strong probe for detection tasks |
| Source predictor | Mean embedding, 1536, through a rank-512 projection | One logit per source recording, over 1.5 million | The self-supervised source-prediction loss; low rank because a full 1536-by-1.5-million matrix would dwarf the model |
The prototype head deserves a picture in your mind. Instead of one weight vector per class, it learns four prototypes per class, each a vector in the 1,536-dimensional feature space. To score a class, it compares every one of the fifteen spatial positions against each of the four prototypes and takes the maximum activation. A prototype can therefore latch onto a species' characteristic sound occurring anywhere in the window, at any of the five time positions and three frequency bands, which is why the paper finds prototype probes beat linear probes on detection tasks. An orthogonality loss pushes the four prototypes of a class apart so they capture different aspects of the call rather than the same one four times.
# forward pass, shapes as the paper gives them x = audio # [160000] 5 s at 32 kHz, mono S = log_mel(x, win=640, hop=320, mels=128) # [500, 128] fixed function, no parameters E_S = efficientnet_b3(S) # [5, 3, 1536] spatial embedding E_A = E_S.mean(axis=(0, 1)) # [1536] mean embedding (what users keep) logits = W_lin @ E_A # [14795] linear head proto = max_over_positions_and_prototypes(stop_gradient(E_S)) # [14795] 4 prototypes per class source = W_src @ (P_512 @ E_A) # [1542778] rank-512 projection first
The model is trained with three objectives and in two phases, and the interplay between them is the paper's method. Take them one at a time, then watch them run together.
The linear head is trained with softmax and cross-entropy. That is unremarkable except for one choice. A window can contain several species, so the natural loss is a sigmoid per class with binary cross-entropy: each species present or absent, independently. The paper does not do that. Following a large-scale weakly-supervised vision result, it uses a softmax over all 14,795 classes with a target that spreads probability evenly across the k species present, one over k each. The authors found this trains faster than the traditional multi-label sigmoid. The intuition: softmax makes classes compete, so pushing up the present species automatically pushes down the 14,790 absent ones, a much stronger gradient signal per example than 14,795 independent yes-or-no decisions.
Distillation normally needs a bigger teacher. Self-distillation uses the model itself, and the paper's version has an architectural elegance to it. The prototype classifier is trained for species classification with its own softmax cross-entropy plus the orthogonality loss. Its predictions are then used as soft targets for the linear classifier: instead of learning "this window is one hundred percent Hermit thrush," the linear head learns "eighty percent Hermit thrush, fifteen percent Swainson's thrush, a little Veery," the confusion structure the prototype head has discovered. Teacher and student share the same embedding model.
The critical detail is a stop-gradient between the embedding model and the prototype head. The prototype head reads the spatial embedding but never sends gradients back into it. Only the linear head and the source head shape the embedding. Why? Because if the teacher could reshape the embedding to make its own job easier, teacher and student would collapse toward each other and the soft targets would carry no new information. Freezing the teacher's influence on the shared representation keeps the two views genuinely different, which is what makes the distillation worth anything.
A toy with six species. Watch the two-phase schedule run. In phase one the prototype head learns from hard labels and the linear head learns from hard labels; the prototype head's soft predictions sharpen over time but are not used. In phase two the linear head's target switches to the prototype head's soft distribution. Toggle the stop-gradient off to see what happens when the teacher is allowed to reshape the shared embedding: the two heads collapse toward agreement and the soft targets lose their structure.
The third objective is called DIET in the vision paper that proposed it, and it is almost embarrassingly simple: give every recording in the dataset its own class, and train a classifier to predict which recording a window came from. There are over 1.5 million recordings, so this is a 1.5-million-way classification problem, and it is self-supervised in the sense that no human label is involved: the identity of the recording is the label.
What makes it work is data augmentation, and in audio the augmentation is free. A long recording yields many non-overlapping five-second windows, and the model must map all of them to the same source. To do that it has to learn what is stable across a recording: the individual bird's voice, the microphone, the habitat's background, the reverberation of that particular clearing. Those are exactly the fine-grained features that a species label alone would never force it to keep, and the paper notes that a similar objective has been used to identify individual animals. The authors also offer a second reading: source prediction is just an extremely fine-grained supervised classification problem, which is the Chapter 1 insight taken to its limit.
Look at the phase-2 settings again. Every source of noise that phase one wanted, mixed windows, dropout, the source-prediction pressure, is switched off. The only signal is the teacher. And the teacher, remember, was trained with all that noise. Phase two is the model distilling into its own linear head everything its noisier, prototype-shaped self learned, at a learning rate small enough not to disturb it. This is also the authors' explanation for why random windows suddenly work: a soft target from a teacher that has seen the whole distribution is far more forgiving of a window that missed the bird than a hard label is.
# one training step, both phases (the paper's Section 2.3) E_S, E_A = embed(mixed_window) L_species = softmax_xent(W_lin @ E_A, target=uniform_over_present(y)) # 1/k on each present class p_proto = proto_head(stop_gradient(E_S)) # teacher never shapes E L_proto = softmax_xent(p_proto, target=uniform_over_present(y)) + orthogonality(prototypes) L_source = softmax_xent(W_src @ (P @ E_A), target=source_id) # 1.5M-way L_distill = softmax_xent(W_lin @ E_A, target=softmax(p_proto)) if phase == 2 else 0 loss = L_species + L_proto + lam_src * L_source + lam_distill * L_distill # lam_src 0.11 then 0; lam_distill 0 then 4.22
A recipe with this many knobs, mixup's four parameters, three loss weights, dropout, learning rate, cannot be tuned by hand, and it must not be tuned on training accuracy, because the whole point is transfer. So the paper runs a black-box optimizer, Vizier, against the validation suite of Chapter 7, and the settings it chose are themselves evidence about what the method is doing.
Each phase gets two stages of one hundred models. Stage one trains a hundred configurations; based on their validation scores Vizier proposes a hundred more; the best is kept. Each model took twenty to thirty hours on an eight-chip TPU pod, the range set mostly by how many mixup sources it blended. That is roughly four hundred training runs for the two phases, and the paper ran the whole search twice, once with random windows and once with peak selection. Phase one searches learning rate, dropout on the embedding before the heads, the source-prediction loss weight, and the mixup parameters. Phase two adds the self-distillation weight.
The ranges Vizier converged on for each phase, with the final chosen values marked. The two phases wanted opposite things, and reading the difference is the fastest way to understand the method. Hover or tap a knob for the paper's reading.
| Knob | Phase 1 preferred range | Phase 1 final | Phase 2 preferred range | Phase 2 final |
|---|---|---|---|---|
| Mixup sources | 2 to 5 | n = 2 (up to 3 sources) | mostly 1 (no mixing) | none |
| Mixup concentration ω | 10 to 30 | 1 | – | – |
| Beta-binomial α, β | α > β | 91.3, 100 | – | – |
| Source-prediction weight | 0.1 to 0.9 | 0.11 | 0 to 0.4 | 0 |
| Dropout | 0.3 to 0.6 | 0.49 | 0 to 0.5 | 0 |
| Self-distillation weight | – | – | 1.5 to 4.5 | 4.22 |
| Learning rate | – | 6.41 × 10−4 | small | 3.20 × 10−6 |
Two things stand out. The source-prediction loss has a much larger scale than the species loss, because it is a 1.5-million-way softmax, so a weight of 0.11 is not "barely on"; the paper footnotes exactly this. And the phase-two learning rate is two hundred times smaller than phase one's. The optimizer, given the freedom, chose to make phase two a gentle polishing pass with a single strong signal, the teacher, and nothing else. Nobody designed that schedule. The validation suite selected it.
The validation design is the part of this paper most worth stealing for other fields. It answers a question every foundation-model team faces: how do you choose a model when the tasks it will be used for do not exist yet?
The authors ask how practitioners actually use Perch, and they find three patterns. Some run the pretrained classifier out of the box. Some embed everything and search by example: find me every window that sounds like this one. Some train a small classifier on a few dozen labels. So validation has three task types, each scored with ROC-AUC, the probability that a random positive outscores a random negative, which is stable across class imbalance in a way accuracy is not.
| Task type | Procedure | Datasets |
|---|---|---|
| Pretrained classification | Score every 5 s window with a 2.5 s stride using the pretrained head; a prediction counts if it overlaps an annotation | Powdermill and Caples: two fully annotated soundscapes whose species are a subset of the training classes |
| One-shot retrieval | Pick a random example of a species, rank all windows by cosine similarity to it, compute ROC-AUC treating every other species as negative; a proxy for search and clustering | BEANS detection sets, Weldy call types, Powdermill, Caples |
| Linear transfer | Sixteen random examples per class, each embedded as the average over its 5 s windows; train a logistic-regression probe for 10,000 steps; ROC-AUC on the rest | BEANS classification sets, Godwit calls, Yellowhammer dialects, bats, and three marine sets: DCLDE, NOAA, ReefSet |
Within each type the datasets are combined with a geometric mean, and the three type scores are combined with another geometric mean. That choice is deliberate and worth understanding.
Two candidate models, each scored on six datasets. Drag any bar. The arithmetic mean rewards a model that is brilliant on five datasets and collapses on one; the geometric mean does not. Try dropping one score toward zero and watch the two means diverge. That gap is why model selection here favors consistency over peaks.
A geometric mean of six numbers is the sixth root of their product. If any one of them goes to zero, so does the product. So a model that scores 0.95 on five bird datasets and 0.40 on the bat dataset loses to a model scoring 0.88 everywhere. The authors' earlier evaluation paper argued this is the right bias for bioacoustics: a practitioner picking up the model for a new taxon needs it to not fail, more than they need it to excel on birds. Nineteen datasets in total feed the validation score, and the embedding is frozen throughout.
Validation chooses the model; two public benchmarks test it, and neither was used for selection. BirdSet has six fully annotated soundscape datasets from the continental United States, Hawai'i, Peru and Colombia. It provides training splits for fine-tuning, and the paper pointedly does not use them: it applies the prototype classifier's predictions directly. BEANS, the Benchmark of Animal Sounds, has twelve cross-taxa tasks covering birds, land and marine mammals, frogs and insects, split into classification and detection. There, the paper trains linear and prototype probes on each task's training split with the embedding frozen. The two together cover the same kinds of domain shift the validation suite was designed around, while allowing a fair comparison with published models.
Three tables carry the paper's claims. The first compares against published bioacoustics models on both benchmarks; the second is the label-granularity ablation you met in Chapter 1; the third is the marine transfer result, the one that surprised the authors.
Table 3 as bars. Pick a metric. Methods are labeled by how they touch the benchmark: Pre applies a pretrained head, LP and PP are linear and prototypical probes on a frozen embedding, FT fine-tunes the whole model on the benchmark's training split. Every Perch 2.0 row is Pre, LP or PP. Watch which competitors needed FT to get close.
| Method | How | BirdSet AUROC | BirdSet cmAP | BirdSet top-1 | BEANS acc. | BEANS mAP |
|---|---|---|---|---|---|---|
| Audio ProtoPNet-5 | Pre | 0.896 | 0.423 | 0.623 | – | – |
| BirdMAE-L | FT | 0.886 | 0.440 | 0.601 | – | – |
| AVES-Bio | FT | – | – | – | 0.817 | 0.398 |
| BioLingual | FT | – | – | – | – | 0.479 |
| Perch 1.0 | Pre / LP | 0.839 | 0.356 | 0.613 | 0.809 | 0.353 |
| Perch 2.0, phase 1 only | Pre / LP / PP | 0.902 | 0.431 | 0.642 | 0.835 | 0.426 / 0.499 |
| Perch 2.0, peak-selected windows | Pre / LP / PP | 0.907 | 0.430 | 0.619 | 0.839 / 0.841 | 0.426 / 0.504 |
| Perch 2.0, random windows | Pre / LP / PP | 0.908 | 0.431 | 0.665 | 0.838 / 0.840 | 0.415 / 0.502 |
Read it in three passes. First, the headline: on BirdSet ROC-AUC, the metric the authors argue is the most stable and informative, Perch 2.0 reaches 0.908 with no fine-tuning, against 0.896 for the best prior model and 0.886 for BirdMAE-L, which was fine-tuned on BirdSet's own training splits and whose score is additionally inflated by having tuned hyperparameters on one of the test subsets. On BEANS, the prototype probe's 0.504 mAP beats BioLingual's fine-tuned 0.479. Second, the honest column: on BirdSet cmAP, BirdMAE-L's fine-tuned 0.440 edges Perch's 0.431. The paper wins the metric it argues for and loses one it argues against, and reports both. Third, the two comparisons inside the Perch rows: random windows match peak selection, and the prototype probe beats the linear probe on detection by a wide margin (0.502 versus 0.415), which is why the prototype head, originally added as a teacher, is also recommended as a probe.
Perch 2.0's training data has almost no underwater audio: a few dozen whale recordings, mostly made on phones above the surface. Yet on few-shot transfer to marine datasets, sixteen labeled examples per class and a linear probe on the frozen embedding, it beats two models built specifically for the sea.
| Model (16-shot linear probe, ROC-AUC) | DCLDE species | DCLDE ecotype | DCLDE known species | NOAA whales | ReefSet |
|---|---|---|---|---|---|
| Multispecies Whale (Google) | 0.914 | 0.821 | 0.954 | 0.917* | 0.855 |
| Surf Perch (trained on ReefSet) | 0.947 | 0.903 | 0.984 | 0.899 | 0.986* |
| Perch 1.0 | 0.968 | 0.931 | 0.981 | 0.905 | 0.970 |
| Perch 2.0 | 0.977 | 0.945 | 0.989 | 0.924 | 0.981 |
The asterisks matter: the whale model was trained on much of the NOAA data and Surf Perch was trained on ReefSet, so those two cells are in-distribution for the specialists and still only barely ahead. On orca ecotypes, telling apart locally adapted killer-whale populations by their calls, a bird model with a sixteen-shot probe scores 0.945 against the whale model's 0.821. And there is a sharp lesson in one more number the appendix gives: applying the whale model's own pretrained classifier directly to DCLDE scores 0.612, while a probe on that same model's embeddings scores 0.954. The embedding knew far more than the head exposed.
The discussion section is titled "Why supervision?" and it is the most useful two pages in the paper for anyone outside bioacoustics, because it tries to explain a result that contradicts the field's expectations rather than just report it. Four hypotheses.
The self-supervised models that work in vision are trained on enormous corpora and are themselves enormous. DINOv2 saw 142 million images. Xeno-Canto and iNaturalist together are two orders of magnitude smaller. Self-supervision's advantage is that it can use unlabeled data, and bioacoustics does not yet have a hundred million diverse unlabeled recordings organized for training. Meanwhile, vision results suggest that once a supervised model has hundreds of thousands of labeled examples, it becomes very hard to beat, and Perch has 1.5 million.
Self-supervised objectives depend on data augmentations that encode what should not matter, cropping, color jitter, flipping, and performance is highly sensitive to choosing them well. Those were tuned for photographs over a decade. Nobody knows the right invariances for a spectrogram of a frog yet. The authors note that general audio benchmarks are still dominated by supervised and semi-supervised models, and that methods developed for vision have transferred to bioacoustics worse than hoped in their own earlier work on domain adaptation.
This is the Chapter 1 argument with its citations. The gap between self-supervised and supervised models widens as the supervised labels get finer. Species labels are about as fine as natural labels get: over fifteen thousand classes, with the hardest distinctions between species in the same genus. Table 7 showed the effect directly, transfer degrading monotonically as labels were coarsened to genus, family and order.
A picture of the four hypotheses as pressures on one embedding. Each pressure, when active, spreads the synthetic classes apart along a different axis; remove a pressure and its axis collapses. Toggle them and watch a probe's separability on a held-out task, the number at the bottom, respond. Illustrative, but the ordering follows the paper's ablations.
The transfer to bats, mosquitoes, gibbons and whales is the part that looks least likely on paper. The authors' explanation is that birdsong spans a huge range of acoustic structure, and that the physical mechanisms of sound production are shared across terrestrial vertebrates, so a model that has learned to resolve the full variety of birdsong has learned much of what any vocalization can do.
| Stated limit | Implication |
|---|---|
| Still 89% bird data; mammals 1%; bats excluded by the 16 kHz ceiling | The taxa with the fewest labels are where the model is weakest, and the paper proposes source prediction as the semi-supervised route to fill those gaps with unlabeled recordings |
| Evaluation still does not fully match deployment | The authors say more work is needed on benchmarks that reflect real use; the geometric-mean suite is their best current answer, not their final one |
| Only species labels are used as supervision | Recordings carry metadata such as time of day, season and location; the success of source prediction suggests those could become additional classification tasks |
| One retrieval result per species, sixteen-shot probes | Practical for practitioners, but the numbers describe a specific low-label regime, not full fine-tuning |
| Thing | One line | Number |
|---|---|---|
| Data | Xeno-Canto + iNaturalist + Tierstimmenarchiv + FSD50K, one taxonomy, bats removed | 1,542,778 recordings, 14,795 classes |
| Window | 5 s at 32 kHz; random or energy-peak selection; random works as well after distillation | 160,000 samples → 500×128 log-mel |
| Mixup | N ~ BetaBinomial + 1 sources, Dirichlet weights, gain-normalized sum, multi-hot target | 2 to 5 sources in phase 1, none in phase 2 |
| Model | EfficientNet-B3 → spatial 5×3×1536 → mean 1536 | 12M parameters |
| Heads | Linear (14,795), ProtoPNet (4 prototypes per class, max), source (rank-512, 1.5M classes) | Training only |
| Losses | Softmax species CE with 1/k targets; self-distillation from the stop-gradient prototype teacher; source prediction (DIET) | Phase 2 distill weight 4.22 |
| Search | Vizier, 2 stages × 100 models per phase, 20 to 30 h per model on TPUv3-8 | Scored on frozen-embedding transfer |
| Validation | Classify, retrieve, linear-transfer; geometric means; 19 datasets | Embedding frozen |
| Results | BirdSet AUROC 0.908 (prior best 0.896); BEANS mAP 0.504 with a prototype probe; DCLDE species 0.977 vs 0.947 for Surf Perch | No fine-tuning anywhere |
| Granularity | Species 0.837 → genus 0.820 → family 0.761 → order 0.608 BEANS accuracy | Finer labels, better transfer |
The spectrogram front end and why the mel scale exists are in Audio Representations. The self-supervised competitors Perch beats are worked through in the Audio-MAE Veanor and in Classical Audio Classification. Distillation in general, and why a teacher's soft targets carry more than a hard label, is Knowledge Distillation. The contrastive alternative this paper argues against is CLAP, and the transformer alternative to EfficientNet for spectrograms is the Audio Spectrogram Transformer. For the other half of the argument, a case where the same lab does put a reasoning layer above frozen embeddings, see the Earth AI Veanor, whose Population Dynamics and AlphaEarth embeddings are used with exactly the frozen-embedding, cheap-probe philosophy this paper defends.
Perch 2.0 is released, and the whole evaluation philosophy is reproducible on a laptop: download the model, embed a public soundscape dataset once, and train a logistic-regression probe on sixteen examples per class. Then re-run with the labels collapsed to genus and family and watch the probe's ROC-AUC fall. If it falls the way Table 7 says, you have reproduced the bittern lesson's central mechanism in an afternoon. If it does not, you have found something the authors would want to hear about.