Method Comparison Acceptance Interpreter
Method Comparison Acceptance Interpreter
This page does not fit a regression. It consumes one. Give it the slope and intercept you already have from the Deming regression calculator, the concentrations where a clinical decision actually gets made, and the error you are allowed — and it answers the question CLSI EP09c-Ed3 asks: is the predicted bias at each medical decision point acceptable? A slope of 1.05 is not a verdict; a 5% bias at a threshold that separates treating from not treating might be.
Is the bias at my decision points acceptable?
Slope + intercept → accept or rejectA Deming fit of a candidate method against the comparative method gives a slope of 1.04 and an intercept of 0.08 units. The analyte has decision points at 2.0 and 8.0 units. The laboratory works to an allowable total error of 10%, and the candidate method’s within-laboratory CV is 3.0%.
Predicted bias at a medical decision concentration
Predicted bias (%) = 100 x predicted bias / Xc
Full criterion: |predicted bias %| + 1.65 x CV% is at or below the allowable total error %
Weaker, bias-only criterion: |predicted bias %| is at or below the allowable total error %
- (slope − 1) x Xc
- the proportional part of the bias. It grows with concentration, so it decides the verdict at the high decision point and is nearly invisible at the low one. Its causes are calibration, calibrator value assignment and traceability, or a real difference in what the two assays measure
- intercept
- the constant part. It is the same absolute quantity at every concentration, so as a PERCENTAGE it is largest at the lowest decision point. Its causes are blanks, baselines, non-specific binding and carryover
- why both decision points
- because the two components fail at opposite ends. A method with a slope of 1.00 and an intercept of 0.1 is perfect at a decision point of 20 and 5% out at a decision point of 2. A method with a slope of 1.05 and a zero intercept is 5% out everywhere but only clinically relevant where the threshold is. Averaging over the range hides both
- 1.65
- the one-sided 95% normal deviate. |bias| + 1.65 x CV is the conventional total error model: the point beyond which about 5% of single results would fall. It is a model, not a measurement, and it assumes the errors are normally distributed and the bias is stable
- where the allowable total error comes from
- it has to be chosen and documented, because the candidate sources disagree. A CLIA proficiency limit, an EQA scheme’s acceptance limit, a specification derived from within-subject and between-subject biological variation, and a manufacturer’s claim will give different numbers for the same analyte, and a method can pass against one and fail against another. Name the source next to the number every time it is quoted
- what this page is not
- it is not a regression calculator. The Deming regression calculator produces the slope and intercept; this page consumes them. Nor does it replace a scatter plot and a difference plot — a regression summarises the relationship and cannot show you the three samples that disagree wildly
Worked example
A Deming fit of a candidate method against the comparative method gives a slope of 1.04 and an intercept of 0.08 units. The analyte has decision points at 2.0 and 8.0 units. The laboratory works to an allowable total error of 10%, and the candidate method's within-laboratory CV is 3.0%.
Predicted bias at 2.0 = (1.04 − 1) x 2.0 + 0.08 = 0.08 + 0.08 = 0.16 units, which is 0.16 / 2.0 = 8.0%
Predicted bias at 8.0 = (1.04 − 1) x 8.0 + 0.08 = 0.32 + 0.08 = 0.40 units, which is 0.40 / 8.0 = 5.0%
The worse of the two is the LOW decision point at 8.0% — the opposite of what the raw slope and intercept suggest, because the fixed 0.08 offset is 4% of 2.0 and only 1% of 8.0
Bias-only check: 8.0% is inside the 10% allowable total error, so a comparison against allowable total error alone would call this acceptable
Full check: 8.0% + 1.65 x 3.0% = 8.0 + 4.95 = 12.95%, which exceeds 10% → marginal
The practical reading: on average the new method agrees well enough, but around one result in twenty at the 2.0 threshold will be outside the allowable total error. Halve the imprecision to 1.2% and the same bias passes; halve the bias and it passes with the imprecision unchanged
Where each component of the bias does its damage
| Predicted bias at Xc = 2.0 | Predicted bias at Xc = 8.0 | Points at | |
|---|---|---|---|
| Slope 1.05, intercept 0 | 0.10 units (5.0%) | 0.40 units (5.0%) | Calibration, standardisation, assay specificity |
| Slope 1.00, intercept 0.10 | 0.10 units (5.0%) | 0.10 units (1.3%) | Blank, baseline, non-specific binding, carryover |
| Slope 1.04, intercept 0.08 | 0.16 units (8.0%) | 0.40 units (5.0%) | Both — and the low point is the one that fails |
| Slope 0.94, intercept 0.30 | 0.18 units (9.0%) | −0.18 units (−2.3%) | Opposing components that cancel mid-range |
The two criteria, and why they disagree
| Criterion | Form | What it assumes |
|---|---|---|
| Bias only | |bias %| at or below TEa% | That the whole error budget may be spent on bias — which leaves nothing for imprecision, so about half of single results would be expected to exceed the limit at the boundary |
| Bias plus imprecision | |bias %| + 1.65 x CV% at or below TEa% | That errors are normally distributed and the bias is stable, and that 95% of single results should be inside the allowable total error |
A slope and an intercept are not an answer
Method comparison studies usually end with a slope, an intercept, a correlation coefficient and a verdict that is really a matter of taste — a slope between 0.95 and 1.05 is often treated as acceptable, an intercept near zero as reassuring, and a correlation above 0.99 as proof of agreement. None of those three numbers answers a clinical question. A correlation coefficient measures whether the two methods rank samples in the same order, which they will if the samples span a wide range, however badly they agree. A slope and an intercept describe a line, and what matters is not the line but where it sits relative to identity at the particular concentrations where somebody makes a decision.
CLSI EP09c-Ed3 puts the question that way round. Fit the regression, then use it to predict the bias at each medical decision concentration, and compare that prediction against an acceptance criterion set before the study began. The arithmetic is trivial — the predicted bias at a concentration Xc is the slope minus one, times Xc, plus the intercept — but it produces answers that the raw regression parameters do not. The two components of bias fail at opposite ends of the range: the proportional part grows with concentration and the constant part, expressed as a percentage, shrinks with it. A method can therefore have a barely noticeable bias in the middle of its interval and an unacceptable one at a rule-out threshold near the bottom, and reporting a single average bias across the measuring interval hides exactly that.
There is a second reason this page asks for the method’s imprecision as well as its bias. Laboratories commonly compare a bias estimate against an allowable total error, and allowable total error is not a limit on bias. It is the budget for bias and imprecision combined. A method whose bias by itself consumes the entire allowance has nothing left over, and around half of its individual results at that concentration will lie outside the allowance even though the bias comparison said it passed. Adding 1.65 times the within-laboratory CV to the absolute bias, and requiring the sum to fit, is the version of the comparison that means what laboratories usually think the simpler one means.
None of this establishes that either method is correct. A method comparison measures agreement between two procedures, and if they share a calibration error it is invisible to the study. Nor does a good regression rule out individual discrepant samples, which is why a difference plot belongs beside the fit. And the conclusion is valid only across the concentration range the samples actually covered — a decision point reached by extrapolating beyond the highest sample has not been evaluated, whatever the arithmetic says.
Frequently asked questions
Why does this page ask for my CV in a bias calculation?
Because allowable total error is the budget for bias and imprecision together, not a limit on bias alone. If the bias by itself uses up the whole allowance, about half of individual results at that concentration will fall outside it, even though a bias-versus-TEa comparison said the method passed. The criterion applied here is |bias %| + 1.65 x CV% at or below the allowable total error, where 1.65 is the one-sided 95% normal deviate.
Why evaluate bias at decision points instead of averaging across the range?
Because the two components of bias fail at opposite ends. The proportional part, (slope − 1) x concentration, grows with concentration; the constant part, the intercept, is a fixed absolute amount and so is largest as a percentage at low concentrations. A method with a low slope and a high intercept has almost no bias somewhere in the middle of its range and can be unacceptable at both ends. Averaging conceals that, and CLSI EP09c-Ed3 asks for bias at the concentrations where clinical decisions are made.
Which regression should the slope and intercept come from?
Deming or Passing-Bablok, not ordinary least squares. Ordinary least squares assumes the comparative method is measured without error and systematically under-estimates the slope when it is not — which it never is. Use the Deming regression calculator for the fit and bring the slope and intercept here.
Where should the allowable total error come from?
From a documented source that you name alongside the number, because the candidates disagree. A CLIA proficiency limit, an EQA scheme’s acceptance limit, a specification derived from within-subject and between-subject biological variation, and a manufacturer’s claim will give different figures for the same analyte, and a method can pass against one and fail against another. Near the bottom of the measuring interval a percentage allowance also needs an absolute floor, or it shrinks to something smaller than the method’s own imprecision.
The regression looks excellent but a few samples disagree badly. Does that matter?
Yes, and a slope and intercept will never show it. A regression summarises the average relationship, and a handful of grossly discrepant samples barely move it while representing exactly the failure mode that harms patients. Plot the differences against the mean as well, and investigate the outliers individually rather than removing them — interference, an antibody effect or a carryover event will often be the explanation, and each of those is a finding about the method.
Related calculators
References
- CLSI EP09c-Ed3. Measurement Procedure Comparison and Bias Estimation Using Patient Samples. 3rd ed. Clinical and Laboratory Standards Institute; 2018.
- CLSI EP09-A3. Measurement Procedure Comparison and Bias Estimation Using Patient Samples; Approved Guideline. 3rd ed. Clinical and Laboratory Standards Institute; 2013.
- Bland JM, Altman DG. Statistical methods for assessing agreement between two methods of clinical measurement. Lancet. 1986;1(8476):307-310.
- Linnet K. Necessary sample size for method comparison studies based on regression analysis. Clin Chem. 1999;45(6 Pt 1):882-894.
Medical Disclaimer: The tools and content provided here are for educational and reference purposes only. They are not intended to substitute for professional medical advice, diagnosis, or treatment. Clinical decisions should always be based on the comprehensive assessment of a qualified healthcare professional.
