This became one of my favorite engineering sayings. Not because customers are misleading, but because evaluation samples rarely represent reality.
The typical situation
A customer sends:
- Their best samples
- Their worst samples
What is almost always missing:
- Normal production variation - the boring middle that makes up 99% of real output.
Why the middle is the dangerous part
The best and worst samples are easy: any threshold separates them with a huge margin. The trap is that the threshold gets placed in the empty space between them - and production then fills that empty space with parts.
A made-up but typical picture in numbers:
evaluation set: good parts score 10-20, defects score 80-95
chosen threshold: 50 (comfortable margin both ways)
production: good parts score 10-60 across batches and seasons
result: every good part scoring 50-60 becomes a false reject
Nothing was wrong with the threshold given the data. The data was wrong about the world.
How samples lie in practice
- Survivor bias - sample boxes are hand-picked, handled gently, and stored flat. Production parts arrive with fixture marks, handling scuffs, and a week in changing humidity.
- Frozen time - samples capture one moment of a process that drifts with tool wear, seasons, and supplier lots.
- Defect theater - the defect samples are the dramatic ones someone could find by eye. The marginal defects, the ones the system actually exists for, are precisely the ones nobody could collect.
My current questions
Before trusting any sample set, I ask:
- How many suppliers?
- How many production lines?
- How many years of process history?
- What recently changed?
- And the most useful one: how were these samples selected, and by whom? The selection process tells you which lie this particular box is telling.