Proficiency Testing Failure Interpreter

Proficiency Testing Failure Interpreter

An EQA sample has come back outside the acceptance limit. Before you touch the instrument, answer four questions: did one analyte fail or many, were the deviations all in one direction or scattered, was the internal quality control acceptable at the time, and was the sample handled exactly like a patient sample? Those four answers separate a bias from an imprecision problem from a blunder, and those are three different investigations.

Bias, imprecision or blunder?

Failure pattern → where to investigate
This is asked first because a difference in handling invalidates everything downstream, and because under CLIA the laboratory is required to test PT samples in the same manner as patient specimens. Repeating and averaging, or referring the sample elsewhere, are deficiencies in their own right.
The single most informative question. An analytical fault lives on an analyte and an instrument; a handling or clerical fault does not respect those boundaries, and something that crosses instruments is almost never a method problem.
Direction is what separates systematic error from random error. A consistent one-directional deviation across concentrations is a shift or a scaling problem; deviations that scatter either side of the target are imprecision, or a single mishap.
Internal QC is the control group for this investigation. QC that confirms the EQA finding tells you the method was genuinely misbehaving; QC that was clean while EQA failed is itself a finding, and usually points at the sample rather than the method.
Divide the absolute difference between your result and the scheme’s target by the scheme’s acceptance limit for that analyte. A value of 1.0 sits exactly on the limit. Expressing it this way keeps the page scheme-neutral, because CLIA, RCPA, RiliBAK and the commercial EQAS programmes do not use the same limits.
A consistent one-directional deviation that the internal quality control did not see — bias the QC cannot detectExample

One analyte failed on every sample in the round. All five results were low, by amounts between 1.2 and 1.9 times the scheme’s acceptance limit — the worst was 1.6 times on the sample being reviewed. Internal quality control for that analyte was in control before, during and after the event, and the sample was handled exactly as a patient specimen.

The four discriminating questions

1. How far did it spread? One analyte and one sample → a discrete event. One analyte, many samples → a property of that method. Many analytes, one instrument → a property of that analyser, or of the specimen. Many analytes, many instruments → not analytical at all.
2. Which direction? All one way → systematic error, meaning bias. Scattered both ways → random error, meaning imprecision.
3. What was the internal QC doing? Shift or trend → bias, confirmed. Increased scatter → imprecision, confirmed. Clean → the cause is invisible to QC, which is itself informative.
4. Was the sample handled as a patient sample? If not, nothing above can be interpreted, and it is a regulatory finding in its own right.
why QC being clean is a finding
internal quality control limits are usually derived from the laboratory’s own historical data, so a bias that was present when the means were established is baked into the chart. QC answers "is the method where it has been?", not "is the method where it should be?". A one-directional EQA deviation with clean QC is the classic presentation of a long-standing calibration bias, or of proficiency material that is not commutable on your method
commutability
whether a processed proficiency material behaves on a given method the way a patient sample of the same concentration would. Non-commutable material produces method-group-specific deviations that are a property of the material, not a defect in the laboratory, which is why many schemes grade against a method-group peer mean rather than an all-method consensus. If your result sits with your peers and away from the overall target, this is the first explanation to consider
expressing the deviation scheme-neutrally
as a multiple of the scheme’s own acceptance limit, because the schemes do not agree. CLIA, the RCPA allowable limits of performance, the German RiliBAK limits and the commercial EQAS programmes set different limits for the same analyte, derived in different ways. Any number quoted has to name its scheme
what counts as failing, under CLIA
for chemistry, haematology and immunology analytes, failure to attain at least 80% acceptable responses for an analyte in a testing event is unsatisfactory performance for that analyte; unsuccessful participation is failure in two consecutive testing events or two out of three consecutive events. So one isolated miss is a signal to monitor, and the second one is the one with consequences

Worked example

One analyte failed on every sample in the round. All five results were low, by amounts between 1.2 and 1.9 times the scheme's acceptance limit — the worst was 1.6 times on the sample being reviewed. Internal quality control for that analyte was in control before, during and after the event, and the sample was handled exactly as a patient specimen.
Handling was routine, so the analytical evidence can be read
The failure is confined to one analyte but spans every sample in the round → a property of that method, not a discrete event and not an instrument-wide fault
Every deviation is in the same direction → systematic error, not imprecision
Internal QC was clean throughout → the bias is invisible to the control chart, which happens when the control means were established while the bias was already present
So: compare your result with the method-group peer mean as well as the all-method target. Sitting with your peers and away from the consensus points at a method-group calibration or commutability issue rather than at your laboratory
Change one answer — QC showed a shift at the same time — and the verdict becomes a bias with a date attached, and the investigation starts at the calibration and reagent lot records instead

What the pattern points at

SpreadDirectionInternal QCMost likely
One analyte, one samplen/aCleanBlunder — transposition, transcription, dilution, or a chance excursion
One analyte, many samplesAll one wayShiftBias with a date — calibration, calibrator or reagent lot
One analyte, many samplesAll one wayCleanLong-standing calibration bias, or non-commutable PT material
One analyte, many samplesScatteredMore scatterImprecision — probe, fluidics, mixing, temperature
Many analytes, one instrumentAnyShiftInstrument-wide — water, temperature, optics, pipetting
Many analytes, one instrumentAnyCleanThe specimen — reconstitution, storage, matrix effect
Many analytes, several instrumentsAnyAnyNot analytical — clerical, units, submission mapping, shared utility failure
Read across, and note the two rows that differ only in the QC column. Identical proficiency results mean different things depending on what the control chart was doing, which is why the QC record has to be retrieved before the investigation starts rather than after it has gone down the wrong path.

Causes of proficiency failure in a large published series

CategoryFailuresShare
Technical2,94937.4%
Equipment1,85923.6%
Methodological1,27716.2%
Unexplained90311.5%
Clerical80210.2%
PT evaluation410.5%
Other520.7%
Total7,883100%
Li and colleagues classified 7,883 unacceptable routine chemistry results returned by 2,369 laboratories. The counts were re-added while this page was written and sum exactly to the stated total, and each share reproduces from its own count. The three commonest specific causes were missed scheduled instrument maintenance (19.5%), calibration problems (17.4%) and reagent management (10.0%) — note that the first two are preventable by schedule rather than by investigation, and that one failure in nine had no identifiable cause at all.

Acceptance limits differ by scheme — name the scheme with the number

AnalyteCLIA limit (effective 2024, implemented by PT providers January 2025)
PotassiumTarget value ± 0.3 mmol/L
SodiumTarget value ± 4 mmol/L
GlucoseTarget value ± 6 mg/dL or ± 8%, whichever is greater
Calcium, totalTarget value ± 1.0 mg/dL
CreatinineTarget value ± 0.2 mg/dL or ± 10%, whichever is greater
Cholesterol, totalTarget value ± 10%
ALTTarget value ± 15% or ± 6 U/L, whichever is greater
ASTTarget value ± 15% or ± 6 U/L, whichever is greater
TSHTarget value ± 20% or ± 0.2 mIU/L, whichever is greater
HaemoglobinTarget value ± 4%
These are the CLIA limits from the 2022 final rule, which took effect on 11 July 2024 and were implemented by proficiency testing providers from 1 January 2025. They are NOT the RCPA allowable limits of performance, the German RiliBAK limits or the limits used by the commercial EQAS programmes, which are derived differently and disagree with these and with each other. A result can be acceptable under one scheme and unacceptable under another, which is why the deviation on this page is entered as a multiple of your own scheme’s limit rather than in concentration units.

A scattered outlier with clean QC is a different problem from a one-way shift

When an external quality assessment result comes back outside the acceptance limit, the instinct is to look at the instrument. That is often the wrong place, and it is almost always the wrong place to look first. The proficiency report itself, read alongside the internal quality control record for the same period, usually identifies which of three error types you are dealing with before anybody opens a service manual — and bias, imprecision and blunders have almost nothing in common as investigations.

Systematic error announces itself through direction. If several samples at different concentrations all deviated the same way, the method was measuring consistently and consistently wrongly, and there will be a cause with a date: a calibration, a calibrator or reagent lot, a service event, a parameter change. Random error announces itself through scatter. Deviations falling on both sides of the target, especially alongside wider-than-usual control results, mean the method has become imprecise, and imprecision cannot be corrected by recalibration — it has to be repaired. Blunders announce themselves through isolation. A single badly wrong result in a round where everything else passed, with a clean control chart, is very unlikely to be an analytical property of anything, and the yield from hunting for one is poor.

The internal quality control record is what turns these from guesses into findings, and the most instructive case is the one where the control chart is clean while the proficiency result is not. Control limits are usually derived from a laboratory’s own historical performance, which means a bias that existed when those means were established has been absorbed into the chart and is invisible to it. Internal QC answers whether the method is where it has been; only external assessment answers whether it is where it should be. A consistent one-directional external deviation with untroubled internal control is the classic presentation of a long-standing calibration or standardisation bias — or of proficiency material that is not commutable on your method, which is a property of the material rather than a defect in the laboratory, and which is why many schemes grade against a method-group peer mean rather than an all-method consensus.

Two things bracket the whole exercise. First, the sample must have been handled exactly as a patient sample: same operators, same run, same number of measurements, one report, no consultation with another laboratory. Where that did not happen the analytical evidence cannot be interpreted, and the handling is itself a regulatory finding. Second, the number you failed by is meaningless without naming the scheme it was judged against, because CLIA, the RCPA allowable limits of performance, the German RiliBAK limits and the commercial programmes set different limits for the same analyte by different reasoning. A result can pass under one and fail under another, and the corrective action should be proportionate to the clinical consequence rather than to which scheme you happen to be enrolled in.

Frequently asked questions

My EQA result failed but internal QC was perfect. How is that possible?

Very easily, and it is the most informative combination in the whole exercise. Internal control limits are usually derived from the laboratory’s own historical data, so a bias that was already present when the control means were assigned is built into the chart and cannot be seen on it. Internal QC tells you whether the method is where it has been; external assessment tells you whether it is where it should be. A one-directional external deviation with clean internal QC points at a long-standing calibration or standardisation bias, or at proficiency material that is not commutable on your method.

How do I tell bias from imprecision from the EQA report alone?

By direction and spread. Deviations that are all the same way across several samples and concentrations are systematic error — bias — and there will be a cause with a date attached. Deviations that scatter either side of the target are random error — imprecision — and it cannot be recalibrated away. A single isolated failure in a round where everything else passed is usually neither, and should be investigated as a discrete event: transposition, transcription, dilution, or a genuine chance excursion.

Several analytes failed on one instrument but QC was fine. What does that mean?

Usually something about the proficiency specimen rather than the analyser, because internal control material and proficiency material are different matrices and a matrix effect is invisible to one and obvious to the other. Check the reconstitution first — one wrong diluent volume moves every analyte in that vial in the same direction — then storage, then commutability. Compare your result against the method-group peer mean as well as the overall target; sitting with your peers and away from the consensus is the signature of a method-group matrix effect.

What counts as failing, and when does it become serious?

It depends on the scheme, and the number has to be quoted with the scheme’s name. Under CLIA, for chemistry, haematology and immunology analytes, failure to reach at least 80% acceptable responses for an analyte in a testing event is unsatisfactory performance for that analyte, and unsuccessful participation is failure in two consecutive events or two out of three consecutive events. So one isolated miss is a reason to investigate and monitor; the second consecutive one has regulatory consequences.

Is it acceptable to repeat a proficiency sample and report the mean?

No. CLIA requires that proficiency samples be examined in the same manner as patient specimens — the routine method, the routine number of measurements, the routine operators, a single report — and prohibits sending them to another laboratory or discussing them with one. Repeating and averaging, running duplicates when patients are run singly, or assigning the sample to the most experienced operator all breach that and also make the result unrepresentative of what a patient would have received.

Related calculators

References

  1. 42 CFR Part 493 Subpart H — Participation in Proficiency Testing for Laboratories Performing Nonwaived Testing. §§ 493.801, 493.837, 493.841, 493.843, 493.845.
  2. Centers for Medicare & Medicaid Services. Clinical Laboratory Improvement Amendments of 1988 (CLIA) Proficiency Testing Regulations Related to Analytes and Acceptable Performance. Final rule. Fed Regist. 2022;87(131):40946-41014. Effective 11 July 2024.
  3. Li T, Zhao H, Zhang C, et al. Reasons for proficiency testing failures in routine chemistry analysis in China. Lab Med. 2019;50(1):103-110.
  4. Miller WG, Myers GL, Rej R. Why commutability matters. Clin Chem. 2006;52(4):553-554.
  5. Miller WG, Myers GL, Ashwood ER, et al. Specimen materials, target values and commutability for external quality assessment (proficiency testing) schemes. Clin Chim Acta. 2003;327(1-2):25-37.
  6. Jones GRD, Sikaris K, Gill J. ‘Allowable limits of performance’ for external quality assurance programs — an approach to application of the Stockholm criteria by the RCPA Quality Assurance Programs. Clin Biochem Rev. 2012;33(4):133-139.

Medical Disclaimer: The tools and content provided here are for educational and reference purposes only. They are not intended to substitute for professional medical advice, diagnosis, or treatment. Clinical decisions should always be based on the comprehensive assessment of a qualified healthcare professional.