NGS Variant Confirmation Interpreter

NGS Variant Confirmation Interpreter

Whether a next-generation sequencing call needs orthogonal confirmation before it is reported, from its depth, allele fraction, strand support, genomic context and what it is being reported as. The two published threshold sets disagree with each other, and the page shows you where you are against both.

Does this call need Sanger?

Call metrics + context → confirm or not
This is the first thing the AMP and NSGC joint report asks, not the last. Its recommendation is to confirm variants with significant clinical implications unless the laboratory has criteria “rigorously demonstrated to ensure high positive predictive value”, and it counts pathogenic and likely pathogenic findings, secondary findings on the ACMG list, carrier variants in recessive genes and pharmacogenetic variants as significant — including where penetrance is low.
The AMP and NSGC report names the regions that should “always be subject to orthogonal confirmation”: “low-complexity repeats, such as homopolymers and short tandem repeats, mobile elements, and genes with highly similar paralogs or pseudogenes”. PMS2, SMN1, CYP21A2, GBA, NEB and STRC are the usual suspects for the third option, and in those genes the question is often not whether the variant is real but which copy it is in, which Sanger alone may not settle either.
Strand bias is one of the most reliable single indicators of an artefact, and its causes are known: oxidative damage during library preparation produces G to T changes seen predominantly on one strand, and formalin fixation produces C to T deamination artefacts the same way. There is no universal numeric cut-off for a strand bias score — the statistic differs between callers — so this selector asks for the judgement your pipeline’s own flag or your own view of the alignment supports.
Indels are harder in every direction: harder to align, harder to left-normalise, harder to represent unambiguously, and more likely to be described differently by two pipelines. The AMP and NSGC report keeps indel resolution and nomenclature ambiguity in its list of special circumstances for exactly that reason, and confirmation of a reportable indel also settles which representation is correct.
Total reads covering the position, not the number supporting the variant. The two published threshold sets differ by fivefold here — one study found no false positives at 20 reads or more given other quality filters, and another recommended more than 100 before confirmation could be dropped. Both are on this page, and the gap between them is where your own laboratory’s validation has to do the work.
Supporting reads as a percentage of total reads. For a germline heterozygote expect somewhere near 50%, and treat a fraction far from 50% or 100% as a question rather than a number — it may be a mosaic, a subclonal somatic variant, an allele-balance artefact, or a mismapped read pile. The arithmetic and the interpretation of an unusual fraction are on the variant allele frequency calculator.
Confirmation may be omitted — if your laboratory has validated omission at these metricsExample

A missense single-nucleotide variant in unique sequence, being reported as likely pathogenic. Read depth 180, variant allele fraction 47%, supporting reads balanced across both strands

The two published threshold sets, and the one document that governs both

floor nobody argues with: depth ≥ 20  AND  allele fraction ≥ 20%  AND  QUAL ≥ 100  AND  unique sequence
the conservative set: depth > 100  AND  allele fraction > 40%
depth ≥ 20, fraction ≥ 20%
from a study of 1,109 variants in 825 clinical exomes: 1,079 variants meeting FILTER=PASS, QUAL at least 100, depth at least 20 and fraction at least 20% gave 100% concordance with Sanger, with zero false-positive single-nucleotide variants (866) and zero false-positive indels (213). Homopolymer regions, pseudogene regions and variants with nomenclature uncertainty were excluded from that claim
depth over 100, fraction over 40%
from a study of 7,845 non-polymorphic variants in which 1.3% were NGS false positives. False positives had a mean allele fraction of 16% and a mean coverage of 223; true variants had 46% and 415. Dropping confirmation entirely would have reduced sensitivity from 100% to 97.8%, missing 176 confirmed variants
QUAL
not an input on this page, because the scale is not comparable between callers and a QUAL of 100 from one pipeline is not a QUAL of 100 from another. If your pipeline emits one, apply your own validated threshold alongside the two above
the governing rule
neither threshold set is a standard. The AMP and NSGC joint report states that “read depth, variant type, variant allele fraction, genomic context, and many current quality score calculations have been shown to be inadequate alone”, requires combinations of criteria validated on the laboratory’s own data using a dataset separate from the one they were derived on, and requires the policy to be written, approved by the laboratory director and stated on the report

Worked example

A missense single-nucleotide variant in unique sequence, being reported as likely pathogenic. Read depth 180, variant allele fraction 47%, supporting reads balanced across both strands
It is being reported as clinically actionable, so the question is live rather than moot
Unique sequence: no paralogue or pseudogene, no homopolymer or repeat. Neither of the two "always confirm" contexts applies
Strand support balanced: no artefact signature
A single-nucleotide variant, not an indel, so the nomenclature ambiguity argument does not apply either
Depth 180 is above 100 and the allele fraction of 47% is above 40% — this clears the more conservative of the two published threshold sets, not merely the floor
So the technical case for omitting confirmation is as strong as the literature gets. It is still not permission. Whether this laboratory may report it without Sanger depends on whether this laboratory has validated omission at these metrics, on this assay, and written it into a policy the director has approved
Now change one thing: make it an insertion. The answer flips to confirm, and not because the call looks any worse — because the exact representation of a reportable indel is itself something confirmation settles

What the two studies actually found

Exome study, 1,109 variantsPanel study, 7,845 variants
Threshold proposedFILTER=PASS, QUAL ≥ 100, depth ≥ 20, fraction ≥ 20%depth > 100 and fraction > 40%
False positives foundZero, among 1,079 high-quality calls — 866 single-nucleotide variants and 213 indels1.3% of variants were NGS false positives
What false positives looked likemean allele fraction 16%, mean coverage 223, against 46% and 415 for true variants
Cost of dropping confirmation entirelyNot assessed; 3 of 789 discrepancies were Sanger false negatives rather than NGS false positivesSensitivity would fall from 100% to 97.8%, missing 176 confirmed variants
Explicitly excludedhomopolymer regions, pseudogene regions, variants with nomenclature uncertainty, prenatal cases
Conclusion drawnSanger is no longer always necessarySanger confirmation is required to achieve optimal sensitivity and specificity
Two careful studies reached opposite headline conclusions, and the reason is not that one is wrong. They studied different assays with different chemistries, different capture designs and different variant callers, and the false positive behaviour of an NGS pipeline is a property of that pipeline. That is why the AMP and NSGC joint report declines to set a threshold and instead requires each laboratory to derive and validate its own — and to validate it on data separate from the data it was derived on, which is the step most easily skipped.

When confirmation is not optional

SituationWhy
Homopolymers, short tandem repeats, low-complexity sequenceNamed in the AMP and NSGC report as regions to “always” confirm. Polymerase slippage produces systematic, reproducible indel errors, so seeing the same call in other samples is not evidence it is real
Genes with a pseudogene or close paralogueThe failure is mispping, not miscalling, so every quality metric looks normal. Confirmation must be gene-specific — generic Sanger primers amplify both copies
Strongly strand-biased supportThe signature of oxidative or fixation damage introduced before sequencing. Depth does not rescue it
A variant allele fraction far below 50%Could be an artefact, a mosaic, clonal haematopoiesis or contamination. Note that Sanger’s own limit of detection is around 15 to 20%, so it may not be the right confirmatory method
Reportable insertions and deletionsConfirmation settles the representation as well as the existence, and the representation is what cascade testing and database submission depend on
Any doubt about sample identityThe AMP and NSGC report notes that a confirmation assay doubles as a check against sample and data mix-ups. A laboratory that stops confirming needs another way to catch them — a SNP identity assay or spiked-in controls
Nothing in this table is rescued by a high depth or a clean allele fraction, which is the common thread. These are failures of the alignment, the chemistry or the description rather than of the base calling, and the metrics a confirmation policy is usually built on cannot see any of them.

The argument for dropping Sanger, and what it does not cover

For the first decade of clinical next-generation sequencing, every reportable variant was confirmed by Sanger sequencing. That was a reasonable default when NGS was new and its error modes were poorly characterised, and it is expensive: confirmation adds days to turnaround and a substantial fraction of the cost of a panel or exome, for a step that, on most calls, changes nothing. The case for dropping it is therefore real and the literature behind it is serious.

The trouble is that the literature disagrees with itself, and understanding why is more useful than picking a side. One study of 1,109 variants in 825 clinical exomes found that 1,079 calls meeting a straightforward quality filter — the caller’s own PASS, a quality score of at least 100, at least 20 reads and a variant fraction of at least 20% — matched Sanger perfectly, with zero false positives among 866 single-nucleotide variants and 213 indels. It also found that three of the discrepancies it did see were Sanger false negatives rather than NGS false positives, which is a point worth sitting with: the confirmatory method is not automatically the truth. Another study, of 7,845 variants on targeted panels, found 1.3% of calls were NGS false positives, characterised them (mean allele fraction 16% and mean coverage 223, against 46% and 415 for the true ones), and concluded that confirmation should be retained for anything below 100 reads and a 40% fraction, because abandoning it entirely would have cost 176 confirmed variants and taken sensitivity from 100% to 97.8%.

Both are right about their own assay. False positive behaviour is a property of a particular capture design, chemistry, aligner and variant caller, and it does not transfer. That is exactly the position the 2023 AMP and NSGC joint report takes: it sets no threshold, states that read depth, variant type, allele fraction, genomic context and quality scores “have been shown to be inadequate alone”, and requires laboratories to derive combinations of criteria and validate them on their own data — using a separate dataset from the one the criteria were derived on, which is the step that prevents a policy from simply describing the data it was fitted to. It also requires the policy to be written down, approved by the laboratory director, made available to ordering clinicians and summarised on the report.

Three things the depth-and-fraction argument does not reach at all, and they are the reason this page asks about context before it asks about metrics. In a region with a close paralogue or a pseudogene, the error is in the mapping rather than the base calling, so a beautifully covered call with a perfect allele fraction can be in the wrong gene; PMS2, SMN1, CYP21A2, GBA, NEB and STRC are the standing examples, and generic Sanger primers do not resolve them either. In a homopolymer or short tandem repeat, polymerase slippage generates systematic errors that recur across samples, which is why the AMP and NSGC report warns specifically against treating repeated observation of a variant in other samples as confirmation. And an indel’s problem is often its description rather than its existence: two pipelines will write the same deletion at different coordinates, and confirmation settles which representation goes into the report and into every cascade test that follows.

One last thing that is easy to lose when confirmation is dropped. A Sanger confirmation is also, incidentally, a sample identity check — it catches a swap or a data mix-up that no amount of sequencing depth would reveal, because the wrong sample sequences perfectly well. A laboratory that stops confirming universally needs another mechanism for that, and the AMP and NSGC report names the options: an independent SNP identity assay, or spiked-in tracking oligonucleotides. Related pages: the variant allele frequency calculator for an allele fraction that does not look germline, the NGS run quality interpreter for whether the run the call came from passed QC at all, the NGS coverage depth calculator for the depth the assay was designed to deliver, and the ACMG variant classification interpreter for what the variant means once it is confirmed.

Frequently asked questions

Does every NGS variant need Sanger confirmation?

No, and no professional body requires it. The 2023 AMP and NSGC joint report allows confirmation to be omitted for variants that meet criteria “rigorously demonstrated to ensure high positive predictive value”, provided the laboratory has validated those criteria on its own data, written them into a director-approved policy and stated the policy on its reports. What it does not do is tell you what the criteria should be, because false positive behaviour differs between assays. Certain situations remain exceptions whatever the metrics: low-complexity and repeat regions, genes with close paralogues or pseudogenes, strand-biased calls, low allele fractions and reportable indels.

What read depth and variant allele fraction allow Sanger to be skipped?

The two published answers differ by roughly fivefold, and both are defensible for the assay they were derived on. A study of clinical exomes found zero false positives among 1,079 variants meeting FILTER=PASS, QUAL at least 100, depth at least 20 and variant fraction at least 20%, excluding homopolymer and pseudogene regions. A study of targeted panels found 1.3% false positives and recommended depth above 100 with a fraction above 40% before confirmation could be dropped. This page treats 20 reads and 20% as a floor nobody argues below, and 100 reads and 40% as the point above which even the conservative study agreed — with your own laboratory’s validation deciding everything in between.

Why do homopolymer regions always need confirmation?

Because the errors there are systematic rather than random. Polymerase slippage during library amplification inserts or deletes a base in a run of identical nucleotides at a rate that climbs with run length, and because the mechanism is reproducible, the same false call appears in sample after sample. That defeats the usual quality heuristics twice over: the call can have excellent depth and a convincing allele fraction, and seeing it in other samples looks like corroboration when it is really the same artefact recurring. The AMP and NSGC report names low-complexity repeats among the regions to confirm always, and the largest study supporting reduced confirmation excluded them from its own claim.

Can Sanger sequencing confirm a low-level mosaic variant?

Often not, and this is a real trap. Sanger’s limit of detection is around 15 to 20% variant allele fraction, so a genuine mosaic below that can come back negative and be wrongly dismissed as an NGS artefact. If the NGS allele fraction is low and mosaicism is plausible, the right response is a more sensitive orthogonal method — droplet digital PCR, deep amplicon sequencing, or a second independent NGS assay — and, where the clinical question allows it, testing a second tissue such as skin fibroblasts or buccal cells, since a mosaic may be absent from blood altogether.

If a variant was confirmed in another patient, does that count as confirmation?

No, and the AMP and NSGC report specifically discourages it. Systematic errors recur: an artefact generated by the chemistry or the alignment at a particular position will appear in every sample that covers that position, so seeing the same call confirmed previously can be evidence that the artefact is reproducible rather than that the variant is real. Repeated confirmation of the same variant across samples can support a laboratory’s validated policy as part of a formal analysis, but it is not a substitute for confirming the call in front of you.

What else does a confirmation assay do besides confirm the variant?

It checks that the sample is the sample. A specimen swap, a plate-position error or a data mix-up produces sequencing data of perfect quality belonging to the wrong person, and no depth or allele fraction threshold can detect that. Orthogonal confirmation catches it incidentally because the confirmatory assay is run on the original specimen. The AMP and NSGC report makes the point explicitly and says that a laboratory which does not confirm universally needs an alternative — an independent SNP identity assay, or spiked-in tracking oligonucleotides — rather than nothing.

Related calculators

References

  1. Crooks KR, et al. Recommendations for Next-Generation Sequencing Germline Variant Confirmation: A Joint Report of the Association for Molecular Pathology and National Society of Genetic Counselors. J Mol Diagn. 2023;25(7):411–427. doi:10.1016/j.jmoldx.2023.03.012. Confirmation policies must be written and director-approved; single metrics “have been shown to be inadequate alone”; “low-complexity repeats, such as homopolymers and short tandem repeats, mobile elements, and genes with highly similar paralogs or pseudogenes” should “always be subject to orthogonal confirmation”; criteria must be validated on a separate dataset; and confirmation doubles as a check for “sample or data mix-ups”.
  2. Mu W, et al. Sanger Confirmation Is Required to Achieve Optimal Sensitivity and Specificity in Next-Generation Sequencing Panel Testing. J Mol Diagn. 2016. 7,845 non-polymorphic variants, 1.3% NGS false positives; false positives mean heterozygous ratio 16% (range 10–40) and mean coverage 223 (range 10–689) against 46% and 415 for confirmed variants; recommended “minimal read depth coverage of >100 and a variant allele frequency (or heterozygous ratio) of >40%”; dropping confirmation would reduce sensitivity from 100% to 97.8%, missing 176 confirmed variants.
  3. Sanger sequencing is no longer always necessary based on a single-center validation of 1109 NGS variants in 825 clinical exomes. Sci Rep. 2021. doi:10.1038/s41598-021-85182-w. 1,079 variants meeting FILTER=PASS, QUAL at least 100, depth at least 20X and variant fraction at least 20% gave 100% concordance with Sanger and zero false positives among 866 single-nucleotide variants and 213 indels, excluding homopolymer regions, pseudogene regions and variants with nomenclature uncertainty; 3 of 789 discrepancies were Sanger false negatives.
  4. Lincoln SE, et al. A Rigorous Interlaboratory Examination of the Need to Confirm Next-Generation Sequencing-Detected Variants with an Orthogonal Method in Clinical Genetic Testing. J Mol Diagn. 2019. The interlaboratory study on which much of the case for selective confirmation rests; cited here for its existence and conclusions, as its full text could not be retrieved in this session and no numeric threshold on this page is attributed to it.
  5. Roy S, Coldren C, et al. Standards and Guidelines for Validating Next-Generation Sequencing Bioinformatics Pipelines: A Joint Recommendation of the Association for Molecular Pathology and the College of American Pathologists. J Mol Diagn. 2018. The framework within which a confirmation policy is validated.

Medical Disclaimer: The tools and content provided here are for educational and reference purposes only. They are not intended to substitute for professional medical advice, diagnosis, or treatment. Clinical decisions should always be based on the comprehensive assessment of a qualified healthcare professional.