The bench

How we measure a headphone.

Every graph on this site came out of the procedure below. We are publishing it in full — the rig, the seating rule, the capture count, and the repeatability numbers off our own bench — so you can decide what our measurements are worth instead of taking our word for it.

FixtureminiDSP EARS Pro, both capsules, per-unit calibration at 1/24 octave.
CouplerIEC 60318-4 — the ear-simulator standard the industry rigs use.
Analog gain0 dB, and it stays there. Moving it invalidates the calibration files.
Sample rate48 kHz.
DriveThe fixture’s own headphone output, set once at 94 dB SPL and never touched between left and right.
Captures19 per side. 38 per headphone. 60–90 minutes.
Published bandFull. Nothing truncated at 8 kHz.

We measure it seated well, then prove we can find that seating again

This is the choice everything else follows from, so it goes first.

We seat the headphone the way you would. Shift the cups, move the headband, settle the pads, listen, adjust again. Keep going until it seals properly and sounds right. Then measure it there, and prove we can come back to the same spot.

That is a departure from the usual practice, and the departure is on purpose. The convention headphone reviewing settled on — five placements per side, deliberately randomised, then averaged — was designed to answer a different question: where does a careless user end up? That is a fair question. It is also one nobody asks about their own headphones, because nobody listens at a random seating. You fiddle until it sounds right, and then you listen for three hours at that position.

So the response worth publishing is the one you converge on. Our grid of nine positions is a search for that spot rather than a survey of how bad the misses are.

We never average the reseats

Averaging has a real problem underneath it. A seal either holds or it breaks, so the reseats come out in two clumps. Average four good seals with one broken one and you publish a bass shelf that no listener will ever hear, then build a correction on top of it. The average of a two-humped distribution is a value that never happened.

What we publish instead is the medoid: of six real reseats at the winning coordinates, the one closest to the middle of the group. It is an actual placement that actually occurred, picked by a stated rule rather than by taste. Nothing is blended.

The spread across those six gets published too. Nobody else publishes it, and it turns out to be the more interesting number.

The procedure, per side

Fit. Set the jig spacing for even pad pressure on both sides, write it down, and use that same setting every time we measure that model. Ever.

Nine on a grid. A 3×3 pattern at a fixed step, with plate coordinates recorded for every position. Each one is normalised to 0 dB at 1 kHz and resampled to a log grid before anything is computed, so that a plain level shift between seatings doesn’t read as instability and the top octave doesn’t swamp the metric. The winner is the position that deviates least from its eight neighbours. If the winner lands on an edge of the grid, we guessed the starting point wrong, so we recentre and run it again.

Six reseats. At the winning coordinates, lift the headphone off and seat it again, six times. The medoid of those six is the published response. Their spread is the repeatability figure, and if it comes out wide it means the coordinates aren’t capturing the fit and the grid needs rerunning.

Four perturbations. Headband too tight, headband too loose, and the seal broken with a thin eyeglass arm and a thick one. Glasses wreck a seal, and how much depends on the frame.

Nineteen captures a side, thirty-eight a headphone. The sweeps themselves take seconds. All the time goes into moving the headphone.

Four bands, because they fail differently

20–200 HzSeal. How forgiving the headphone is about fit. Below 100 Hz the coupler is outside its verified range — see the limits.
200 Hz–2 kHzThe control band. Scatter here should be near zero. If it isn’t, our measurement is wrong and we go again.
2–10 kHzPlacement on the pinna. How fussy the headphone is about alignment.
Above 10 kHzReported, and outside the standard’s simulated range. Read it gently.

Spread is reported as a 10th–90th percentile range, which survives one bad seating and reads as a plain ± figure.

What our own bench does

We ran this on ourselves before we ran it on anyone’s headphones, and the numbers are published here whether or not they flatter us.

The instrument contributes 0.002 dB through the midrange — thirty captures with nothing touched between them, both capsules agreeing. So the instrument is never the thing that limits a measurement here. Seating is.

And seating repeatability belongs to the headphone. That was the surprise. Measured across three headphones and five sides, seat-to-seat spread through the midrange ran from0.059 dB on an HD 800 to 0.316 dB on an Oppo PM-3. A five-fold range on one bench, one operator, one afternoon.

We could quote you the 0.059 and let you assume it applies to everything. It would be a lie of the most convenient kind, and we have made a version of it before: this figure has been stated wrongly twice in our own notes, once too pessimistic and once too flattering. There is no single number that describes how repeatably this rig seats a headphone, because the headphone gets a vote.

The two Sennheisers are opposites, which is the clearest way to see it. The HD 800 seals beautifully and hates being moved. The HD 650 is the reverse. So “open circumaurals behave like this” is not a rule, and no shortcut by category survives contact with the bench.

The good news is that the hard ones are still measurable. They need more seatings. Against our own derived tolerance, an HD 800 needs one seating per unit and a PM-3 needs eight. The method absorbs the difference by working harder, which is what a per-model tolerance is for.

Where our measurements stop being trustworthy

The standard our coupler is built to covers 100 Hz to 10 kHz. That is the range where IEC 60318-4 says the thing behaves like an occluded human ear. Outside it, the standard is blunt about what you are holding: above 10 kHz it “does not simulate a human ear, but can be used as an acoustic coupler” out to 16 kHz, and below 100 Hz it “has not been verified to simulate a human ear,” usable as a coupler down to 20 Hz.

So both ends of our graphs sit outside the simulated range, and the two ends fail differently. A reading below 100 Hz is a real, repeatable acoustic measurement of a real cavity. What it stops being is a claim about the pressure at your eardrum. Up top, the same caveat plus the harder problem that ear canals differ more between people than headphones differ from each other, so the top octave is partly about whose ear you are asking.

We publish both regions anyway. Truncating at 8 kHz throws away pinna interaction and driver resonance, and that is most of what people mean by treble character. Cutting off at 100 Hz would throw away the sub-bass, which is half of why people buy headphones. Better to show them and say where the ground gets soft.

One unit is one unit. A single measurement of a single sample is exactly that. Unit variation is real and we have no way to tell you where in the distribution yours fell.

Left and right are always measured separately. Channels can differ by several dB. We keep both and never average them into one line.

Our repeatability figures are about our bench. They describe how precisely we can repeat ourselves, on one rig, with one operator. They are not error bars on your headphone and they never appear as a band around a published curve.

Things we record that often go unrecorded

Pads dominate the response, and pad condition is the variable most commonly left out of a published measurement. So we write down the pad type, its age and condition, the jig spacing, and the clamping force if it has been modified. The headphone comes up to room temperature before it goes on the rig, because cold pads seal differently.

The rig itself gets the same treatment. Capsule serial numbers, calibration files, gain setting, drive level. Every one of those is part of the measurement, and a change to any of them means measurements taken either side of it are not comparable — which we say out loud when it happens.

One thing is still open and worth naming: the fixture’s manual does not state the output impedance of its headphone output. A high output impedance interacts with a headphone’s impedance curve and would put some of the amplifier into the result. We have asked. When we have an answer it goes here, and if it comes back high we will revisit the decision and say what changed.

Two rigs never share an axis

We also hold an older archive measured on a Head Acoustics HMS II.3, from the HeadRoom years. It is a research reference and we treat it as one. It never shares a graph with anything measured here, and no count on this site ever adds the two together. Different fixture, different coordinate system, and no honest way to merge them.

Send us a headphone and hold us to it.

Everything above is what happens to it. If we depart from any of it, the departure gets recorded on the measurement page.

Warren Labs

Get the occasional update

Attune is free on the App Store. Leave your email for the occasional honest update: new measurements, what we learn on the rig, and what ships next.

No spam. Unsubscribe anytime.