Precision Verification Interpreter
Precision Verification Interpreter
Your five-day precision study came out higher than the manufacturer’s claim. That is not the same as failing it. CLSI EP15-A3 compares the observed SD against an upper verification limit built from the claim and the degrees of freedom, because a small study will exceed a true claim by chance about half the time. This page applies that comparison to both components at once, and tells you which one failed — which is the part that says where to look.
Does my precision verify the claim?
Observed SD vs claim → verified or notA five-day, five-replicate study on one control material. Observed repeatability SD 0.24 against a claim of 0.20; observed within-laboratory SD 0.38 against a claim of 0.30. Repeatability degrees of freedom 20; within-laboratory degrees of freedom looked up from EP15-A3 Table 6 as 5.
The upper verification limit, and where the degrees of freedom come from
Repeatability: dfR = N − k (total results minus number of runs). Five days x five replicates → 25 − 5 = 20
Within-laboratory: dfWL from EP15-A3 Table 6, entered with rho = claimed sWL / claimed sR and the number of runs — an effective, Satterthwaite-type degrees of freedom, always smaller than dfR
Verified if observed SD is at or below the claim, or, failing that, at or below the UVL
- why a limit above the claim at all
- an SD from a small study is a noisy estimate of the true SD. If the method's true precision is exactly the claim, the observed SD lands above the claim about half the time. A rule that failed every study whose SD exceeded the claim would fail half of all perfectly good methods. EP15-A3 notes that the UVL always exceeds its claim, generally by at least 30%
- dfR = N − k
- the residual degrees of freedom of the within-run mean square in the one-way ANOVA. For the standard 5 x 5 design that is 25 − 5 = 20, giving a factor of sqrt(31.4104 / 20) = 1.2532 — so a claimed SD of 0.20 has an upper verification limit of 0.251
- dfWL is the trap
- it is NOT N − k and it is NOT N − 1. Within-laboratory precision is a SUM of two variance components, each estimated with its own uncertainty, so its effective degrees of freedom is a Satterthwaite-type combination, tabulated in EP15-A3 Table 6 against the claims ratio and the number of runs. It is always smaller than dfR — for a five-run study, of the order of 4 to 6 — which makes the within-laboratory limit proportionally much wider than the repeatability limit. Using dfR for both is the single commonest misapplication of this protocol, and it makes the within-laboratory test far too strict
- the number of samples
- EP15-A3 obtains its factor from Table 7, indexed by degrees of freedom AND by the number of samples tested, so that running several materials does not inflate the chance of at least one false failure. That table is behind CLSI's paywall and is not reproduced here. The factors on this page are the single-sample chi-square values, which are SMALLER than EP15-A3's multi-sample factors — so this page's limits are slightly tighter than the standard's, and a verdict of "verified" here would also be verified under EP15-A3. A verdict of "not verified" on a multi-sample study should be checked against the published table before it is acted on
- what is not covered
- EP15-A3 also estimates bias from the same experiment, by comparing the grand mean against an assigned value inside a verification interval. That is a trueness question and is deliberately out of scope here
Worked example
A five-day, five-replicate study on one control material. Observed repeatability SD 0.24 against a claim of 0.20; observed within-laboratory SD 0.38 against a claim of 0.30. Repeatability degrees of freedom 20; within-laboratory degrees of freedom looked up from EP15-A3 Table 6 as 5.
Within-laboratory SD 0.38 is larger than repeatability SD 0.24, so the two estimates are internally consistent
Neither component is at or below its claim (0.24 is above 0.20; 0.38 is above 0.30), so the upper verification limits have to be computed
Repeatability: df 20 → factor sqrt(31.4104 / 20) = 1.2532 → UVL = 0.20 x 1.2532 = 0.251. Observed 0.24 is below it
Within-laboratory: df 5 → factor sqrt(11.0705 / 5) = 1.4880 → UVL = 0.30 x 1.4880 = 0.446. Observed 0.38 is below it
Both components are inside their limits, so the manufacturer's precision claim is verified — even though both observed SDs exceeded the claim itself
Use the repeatability degrees of freedom for both and the within-laboratory limit becomes 0.30 x 1.2532 = 0.376, and the same study now FAILS at 0.38. That single substitution is the commonest way this protocol is misapplied
Upper verification limit factors, sqrt(chi2(0.95, df) / df)
| Degrees of freedom | chi2(0.95, df) | Factor | UVL for a claimed SD of 0.20 |
|---|---|---|---|
| 4 | 9.488 | 1.540 | 0.308 |
| 5 | 11.071 | 1.488 | 0.298 |
| 6 | 12.592 | 1.449 | 0.290 |
| 8 | 15.507 | 1.392 | 0.278 |
| 10 | 18.307 | 1.353 | 0.271 |
| 15 | 24.996 | 1.291 | 0.258 |
| 20 | 31.410 | 1.253 | 0.251 |
| 24 | 36.415 | 1.232 | 0.246 |
| 28 | 41.337 | 1.215 | 0.243 |
| 40 | 55.759 | 1.181 | 0.236 |
What each failure pattern points at
| Repeatability | Within-laboratory | Where the variation is | First things to check |
|---|---|---|---|
| Verified | Verified | Nowhere — claim met | Record the limits, then move on to trueness |
| Verified | Fails | Between runs and between days | Calibration stability, reagent lots and on-board age, laboratory temperature across the week, maintenance events inside the study |
| Fails | Fails | Within runs, and therefore everywhere | Pipetting and probe, mixing, bubbles, cuvettes, lamp, and the condition of the control material itself |
| Fails | Verified | Probably in the inputs | That both claims came from the same row and level, that dfR is N − k, and whether the between-run variance came out at zero |
Exceeding a claim is not the same as failing it
A standard deviation calculated from twenty-five results is an estimate, and not a very stable one. If a method's true repeatability SD is exactly the figure on the package insert, a five-by-five study will produce an observed SD above that figure on about half of all occasions, purely because of where the twenty-five results happened to fall. A laboratory that treats any excess over the claim as a failure will therefore reject roughly half the methods that are performing exactly as promised, and will spend its time investigating instruments that have nothing wrong with them.
CLSI EP15-A3 handles this with a two-step rule. If the observed SD is at or below the claim, the claim is verified and no statistics are needed. If it is above the claim, the observed SD is compared against an upper verification limit — the claim multiplied by the square root of the ratio of the 95th percentile of the chi-square distribution to the degrees of freedom. That limit is the value an observed SD would exceed less than one time in twenty if the claim were true. EP15-A3 notes that the limit always exceeds the claim, generally by at least thirty per cent.
The step where this goes wrong in practice is the degrees of freedom, and it goes wrong specifically for the within-laboratory component. Repeatability is straightforward: the within-run mean square in a one-way analysis of variance has N minus k degrees of freedom, twenty for the standard five-day five-replicate design. Within-laboratory precision is not a single mean square. It is the sum of two variance components, each carrying its own uncertainty, so its effective degrees of freedom comes from a Satterthwaite-type combination that EP15-A3 tabulates against the ratio of the two claims and the number of runs. The number that comes out is always smaller than the repeatability degrees of freedom, and for a five-run study it is typically four to six rather than twenty. Because the factor grows quickly as degrees of freedom fall, using twenty for both components makes the within-laboratory limit far too tight, and fails methods that EP15-A3 would verify.
Running both components is worth the effort for a second reason that has nothing to do with passing or failing. When only one of them fails, the study has localised the problem. Repeatability failing means the variation appears inside a single run, which points at pipetting, mixing, probes, cuvettes and optics. Within-laboratory failing alone means the method is stable inside a run and unstable between runs, which points at calibration, reagent lots and the laboratory environment. Those are different investigations, and a single pooled precision figure would not have told you which one to start.
Frequently asked questions
My observed SD is above the manufacturer's claim. Have I failed?
Not necessarily, and probably not. An SD estimated from a small study scatters around the true value, so an observed SD will exceed a true claim about half the time. EP15-A3 compares it against an upper verification limit — the claim multiplied by sqrt(chi2(0.95, df) / df) — and the claim is verified if the observed SD sits below that limit. For twenty degrees of freedom the limit is about 25% above the claim; for five degrees of freedom it is about 49% above.
What degrees of freedom do I use for within-laboratory precision?
Not N − k, which is the repeatability figure. Within-laboratory precision is a sum of two variance components, so it takes an effective, Satterthwaite-type degrees of freedom that EP15-A3 tabulates in Table 6, entered with the ratio of the claimed within-laboratory SD to the claimed repeatability SD and the number of runs. It is always smaller than the repeatability degrees of freedom — for a five-run study, typically four to six. Using the repeatability figure for both is the commonest misapplication of this protocol and it makes the within-laboratory test far too strict.
Can my within-laboratory SD come out smaller than my repeatability SD?
Not as a real quantity, because within-laboratory precision contains repeatability. If the arithmetic produces it, the between-run mean square has come out below the within-run mean square and the estimated between-run variance is negative; EP15-A3 sets it to zero, which makes the two SDs equal rather than inverted. Equal values are legitimate and mean the study detected no between-run component.
Why does this page's limit not match my EP15-A3 spreadsheet exactly?
Because EP15-A3 takes its factor from Table 7, which is indexed by degrees of freedom and also by the number of samples tested, so that running several materials does not inflate the overall chance of a false failure. That table is not reproduced here. The factors on this page are single-sample chi-square values and are therefore slightly smaller, making the limits slightly tighter. A verdict of verified here would also be verified under the published table; a borderline failure on a multi-sample study should be checked against it.
Does verifying precision mean the method is fit for use?
No. It means the method is as precise as the manufacturer said. Whether that precision is good enough for your patients is a different question, answered by combining it with the bias and the allowable total error — which is what the sigma metric does, and what decides how many controls and which rules the method needs.
Related calculators
References
- CLSI EP15-A3. User Verification of Precision and Estimation of Bias; Approved Guideline. 3rd ed. Clinical and Laboratory Standards Institute; 2014.
- CLSI EP05-A3. Evaluation of Precision of Quantitative Measurement Procedures; Approved Guideline. 3rd ed. Clinical and Laboratory Standards Institute; 2014.
- Chesher D. Evaluating assay precision. Clin Biochem Rev. 2008;29 Suppl 1:S23-S26.
- Chi-square critical values computed from the chi-square distribution (95th percentile) rather than transcribed from a printed table; see the factor table on this page.
Medical Disclaimer: The tools and content provided here are for educational and reference purposes only. They are not intended to substitute for professional medical advice, diagnosis, or treatment. Clinical decisions should always be based on the comprehensive assessment of a qualified healthcare professional.
