Proficiency Testing Failure Interpreter
Proficiency Testing Failure Interpreter
An EQA sample has come back outside the acceptance limit. Before you touch the instrument, answer four questions: did one analyte fail or many, were the deviations all in one direction or scattered, was the internal quality control acceptable at the time, and was the sample handled exactly like a patient sample? Those four answers separate a bias from an imprecision problem from a blunder, and those are three different investigations.
Bias, imprecision or blunder?
Failure pattern → where to investigateOne analyte failed on every sample in the round. All five results were low, by amounts between 1.2 and 1.9 times the scheme’s acceptance limit — the worst was 1.6 times on the sample being reviewed. Internal quality control for that analyte was in control before, during and after the event, and the sample was handled exactly as a patient specimen.
The four discriminating questions
2. Which direction? All one way → systematic error, meaning bias. Scattered both ways → random error, meaning imprecision.
3. What was the internal QC doing? Shift or trend → bias, confirmed. Increased scatter → imprecision, confirmed. Clean → the cause is invisible to QC, which is itself informative.
4. Was the sample handled as a patient sample? If not, nothing above can be interpreted, and it is a regulatory finding in its own right.
- why QC being clean is a finding
- internal quality control limits are usually derived from the laboratory’s own historical data, so a bias that was present when the means were established is baked into the chart. QC answers "is the method where it has been?", not "is the method where it should be?". A one-directional EQA deviation with clean QC is the classic presentation of a long-standing calibration bias, or of proficiency material that is not commutable on your method
- commutability
- whether a processed proficiency material behaves on a given method the way a patient sample of the same concentration would. Non-commutable material produces method-group-specific deviations that are a property of the material, not a defect in the laboratory, which is why many schemes grade against a method-group peer mean rather than an all-method consensus. If your result sits with your peers and away from the overall target, this is the first explanation to consider
- expressing the deviation scheme-neutrally
- as a multiple of the scheme’s own acceptance limit, because the schemes do not agree. CLIA, the RCPA allowable limits of performance, the German RiliBAK limits and the commercial EQAS programmes set different limits for the same analyte, derived in different ways. Any number quoted has to name its scheme
- what counts as failing, under CLIA
- for chemistry, haematology and immunology analytes, failure to attain at least 80% acceptable responses for an analyte in a testing event is unsatisfactory performance for that analyte; unsuccessful participation is failure in two consecutive testing events or two out of three consecutive events. So one isolated miss is a signal to monitor, and the second one is the one with consequences
Worked example
One analyte failed on every sample in the round. All five results were low, by amounts between 1.2 and 1.9 times the scheme's acceptance limit — the worst was 1.6 times on the sample being reviewed. Internal quality control for that analyte was in control before, during and after the event, and the sample was handled exactly as a patient specimen.
Handling was routine, so the analytical evidence can be read
The failure is confined to one analyte but spans every sample in the round → a property of that method, not a discrete event and not an instrument-wide fault
Every deviation is in the same direction → systematic error, not imprecision
Internal QC was clean throughout → the bias is invisible to the control chart, which happens when the control means were established while the bias was already present
So: compare your result with the method-group peer mean as well as the all-method target. Sitting with your peers and away from the consensus points at a method-group calibration or commutability issue rather than at your laboratory
Change one answer — QC showed a shift at the same time — and the verdict becomes a bias with a date attached, and the investigation starts at the calibration and reagent lot records instead
What the pattern points at
| Spread | Direction | Internal QC | Most likely |
|---|---|---|---|
| One analyte, one sample | n/a | Clean | Blunder — transposition, transcription, dilution, or a chance excursion |
| One analyte, many samples | All one way | Shift | Bias with a date — calibration, calibrator or reagent lot |
| One analyte, many samples | All one way | Clean | Long-standing calibration bias, or non-commutable PT material |
| One analyte, many samples | Scattered | More scatter | Imprecision — probe, fluidics, mixing, temperature |
| Many analytes, one instrument | Any | Shift | Instrument-wide — water, temperature, optics, pipetting |
| Many analytes, one instrument | Any | Clean | The specimen — reconstitution, storage, matrix effect |
| Many analytes, several instruments | Any | Any | Not analytical — clerical, units, submission mapping, shared utility failure |
Causes of proficiency failure in a large published series
| Category | Failures | Share |
|---|---|---|
| Technical | 2,949 | 37.4% |
| Equipment | 1,859 | 23.6% |
| Methodological | 1,277 | 16.2% |
| Unexplained | 903 | 11.5% |
| Clerical | 802 | 10.2% |
| PT evaluation | 41 | 0.5% |
| Other | 52 | 0.7% |
| Total | 7,883 | 100% |
Acceptance limits differ by scheme — name the scheme with the number
| Analyte | CLIA limit (effective 2024, implemented by PT providers January 2025) |
|---|---|
| Potassium | Target value ± 0.3 mmol/L |
| Sodium | Target value ± 4 mmol/L |
| Glucose | Target value ± 6 mg/dL or ± 8%, whichever is greater |
| Calcium, total | Target value ± 1.0 mg/dL |
| Creatinine | Target value ± 0.2 mg/dL or ± 10%, whichever is greater |
| Cholesterol, total | Target value ± 10% |
| ALT | Target value ± 15% or ± 6 U/L, whichever is greater |
| AST | Target value ± 15% or ± 6 U/L, whichever is greater |
| TSH | Target value ± 20% or ± 0.2 mIU/L, whichever is greater |
| Haemoglobin | Target value ± 4% |
A scattered outlier with clean QC is a different problem from a one-way shift
When an external quality assessment result comes back outside the acceptance limit, the instinct is to look at the instrument. That is often the wrong place, and it is almost always the wrong place to look first. The proficiency report itself, read alongside the internal quality control record for the same period, usually identifies which of three error types you are dealing with before anybody opens a service manual — and bias, imprecision and blunders have almost nothing in common as investigations.
Systematic error announces itself through direction. If several samples at different concentrations all deviated the same way, the method was measuring consistently and consistently wrongly, and there will be a cause with a date: a calibration, a calibrator or reagent lot, a service event, a parameter change. Random error announces itself through scatter. Deviations falling on both sides of the target, especially alongside wider-than-usual control results, mean the method has become imprecise, and imprecision cannot be corrected by recalibration — it has to be repaired. Blunders announce themselves through isolation. A single badly wrong result in a round where everything else passed, with a clean control chart, is very unlikely to be an analytical property of anything, and the yield from hunting for one is poor.
The internal quality control record is what turns these from guesses into findings, and the most instructive case is the one where the control chart is clean while the proficiency result is not. Control limits are usually derived from a laboratory’s own historical performance, which means a bias that existed when those means were established has been absorbed into the chart and is invisible to it. Internal QC answers whether the method is where it has been; only external assessment answers whether it is where it should be. A consistent one-directional external deviation with untroubled internal control is the classic presentation of a long-standing calibration or standardisation bias — or of proficiency material that is not commutable on your method, which is a property of the material rather than a defect in the laboratory, and which is why many schemes grade against a method-group peer mean rather than an all-method consensus.
Two things bracket the whole exercise. First, the sample must have been handled exactly as a patient sample: same operators, same run, same number of measurements, one report, no consultation with another laboratory. Where that did not happen the analytical evidence cannot be interpreted, and the handling is itself a regulatory finding. Second, the number you failed by is meaningless without naming the scheme it was judged against, because CLIA, the RCPA allowable limits of performance, the German RiliBAK limits and the commercial programmes set different limits for the same analyte by different reasoning. A result can pass under one and fail under another, and the corrective action should be proportionate to the clinical consequence rather than to which scheme you happen to be enrolled in.
Frequently asked questions
My EQA result failed but internal QC was perfect. How is that possible?
Very easily, and it is the most informative combination in the whole exercise. Internal control limits are usually derived from the laboratory’s own historical data, so a bias that was already present when the control means were assigned is built into the chart and cannot be seen on it. Internal QC tells you whether the method is where it has been; external assessment tells you whether it is where it should be. A one-directional external deviation with clean internal QC points at a long-standing calibration or standardisation bias, or at proficiency material that is not commutable on your method.
How do I tell bias from imprecision from the EQA report alone?
By direction and spread. Deviations that are all the same way across several samples and concentrations are systematic error — bias — and there will be a cause with a date attached. Deviations that scatter either side of the target are random error — imprecision — and it cannot be recalibrated away. A single isolated failure in a round where everything else passed is usually neither, and should be investigated as a discrete event: transposition, transcription, dilution, or a genuine chance excursion.
Several analytes failed on one instrument but QC was fine. What does that mean?
Usually something about the proficiency specimen rather than the analyser, because internal control material and proficiency material are different matrices and a matrix effect is invisible to one and obvious to the other. Check the reconstitution first — one wrong diluent volume moves every analyte in that vial in the same direction — then storage, then commutability. Compare your result against the method-group peer mean as well as the overall target; sitting with your peers and away from the consensus is the signature of a method-group matrix effect.
What counts as failing, and when does it become serious?
It depends on the scheme, and the number has to be quoted with the scheme’s name. Under CLIA, for chemistry, haematology and immunology analytes, failure to reach at least 80% acceptable responses for an analyte in a testing event is unsatisfactory performance for that analyte, and unsuccessful participation is failure in two consecutive events or two out of three consecutive events. So one isolated miss is a reason to investigate and monitor; the second consecutive one has regulatory consequences.
Is it acceptable to repeat a proficiency sample and report the mean?
No. CLIA requires that proficiency samples be examined in the same manner as patient specimens — the routine method, the routine number of measurements, the routine operators, a single report — and prohibits sending them to another laboratory or discussing them with one. Repeating and averaging, running duplicates when patients are run singly, or assigning the sample to the most experienced operator all breach that and also make the result unrepresentative of what a patient would have received.
Related calculators
References
- 42 CFR Part 493 Subpart H — Participation in Proficiency Testing for Laboratories Performing Nonwaived Testing. §§ 493.801, 493.837, 493.841, 493.843, 493.845.
- Centers for Medicare & Medicaid Services. Clinical Laboratory Improvement Amendments of 1988 (CLIA) Proficiency Testing Regulations Related to Analytes and Acceptable Performance. Final rule. Fed Regist. 2022;87(131):40946-41014. Effective 11 July 2024.
- Li T, Zhao H, Zhang C, et al. Reasons for proficiency testing failures in routine chemistry analysis in China. Lab Med. 2019;50(1):103-110.
- Miller WG, Myers GL, Rej R. Why commutability matters. Clin Chem. 2006;52(4):553-554.
- Miller WG, Myers GL, Ashwood ER, et al. Specimen materials, target values and commutability for external quality assessment (proficiency testing) schemes. Clin Chim Acta. 2003;327(1-2):25-37.
- Jones GRD, Sikaris K, Gill J. ‘Allowable limits of performance’ for external quality assurance programs — an approach to application of the Stockholm criteria by the RCPA Quality Assurance Programs. Clin Biochem Rev. 2012;33(4):133-139.
Medical Disclaimer: The tools and content provided here are for educational and reference purposes only. They are not intended to substitute for professional medical advice, diagnosis, or treatment. Clinical decisions should always be based on the comprehensive assessment of a qualified healthcare professional.
