Every sample lies - not because anyone is misleading, but because evaluation samples rarely represent reality.
Common issues with sample sets
- Sample size too small.
- Missing normal production variation.
- Supplier differences not represented.
- Only a handful of defect examples.
My evaluation process
- Prove feasibility - can it work at all?
- Gather variation - get samples across suppliers, lines, and time.
- Stress-test assumptions - actively try to break it.
- Deploy carefully - ramp slowly and watch for surprises.
What each phase looks like
Phase 1: feasibility (days)
One question only: is the defect information present in the image at all? Best samples, lab conditions, manual everything. If feasibility fails honestly here, you have saved everyone months. Resist the urge to declare victory here - this phase can only say "maybe", never "yes".
Phase 2: variation (weeks)
The phase everyone wants to skip. Ask for samples by date code, not by quality: parts from different weeks, lines, and suppliers. A request for "more samples" gets you more of the same box. A request for "ten parts from each of the last six months" gets you reality.
Phase 3: stress-testing (days)
Deliberately try to break the setup: worst-case part positioning, dirty samples, lighting aged by tape over part of the diffuser, focus offset by half the tolerance. Each failure found here is one that will not be found at 2 AM in production. Write down the breaking point of every parameter - that list becomes the operating window specification.
Phase 4: deployment (weeks)
Run in shadow mode first: the system records decisions but does not act on them. Compare against the existing process for a few thousand parts. Every disagreement is either a system bug or - surprisingly often - a case where the existing process was wrong. Both are worth knowing before going live.
Reporting honestly
The evaluation report should contain a section titled "Where it stops working". If that section is empty, the evaluation is not finished. A customer who hears "it fails beyond 15 degrees of part tilt" trusts the rest of the report far more than one who hears only good news.
The lesson
A successful evaluation proves where the solution stops working, not just where it works.