NGS Coverage Depth Calculator

NGS Coverage Depth Calculator

Mean coverage from reads, read length and target size by the Lander-Waterman relation — and why the mean is the wrong number to report, because coverage is never evenly distributed across a target.

NGS Coverage Depth

Reads, length, target → mean ×
The number of reads, in millions. If the figure you have is read PAIRS — which it usually is for paired-end sequencing — enter the pair count here and select paired-end below.
The length of one read. For paired-end sequencing enter the length of a single mate, not the sum of the two: the selector below accounts for the second mate.
Paired-end sequencing yields two reads per fragment, and both contribute sequence to the target. Getting this wrong halves or doubles the answer, which is the commonest error on the page.
The number of bases you are trying to cover: about 3.1e9 for a haploid human genome, of the order of 3e7 to 6e7 for an exome, and anything from 1e4 to 1e6 for a gene panel. Scientific notation is accepted.
38.7× mean coverageExample

400 million read pairs of 150 bp, paired-end, over a 3.1 × 10⁹ base haploid genome

The Lander-Waterman relation

mean coverage = (number of reads × read length) ÷ target size in bases
paired-end: each pair contributes two reads
equivalently, coverage = total sequenced bases ÷ target bases
number of reads
how many reads were produced. For paired-end sequencing each fragment yields two reads and both contribute sequence, so a figure quoted in read pairs must be doubled. Instrument reports differ in which they quote, and the factor of two is the commonest error in this calculation
read length
the length of a single read. A run described as 2 × 150 bp produces two reads of 150 bases each, not one of 300
target size
the number of bases you are trying to cover — about 3.1 × 10⁹ for a haploid human genome, of the order of tens of megabases for an exome, kilobases to megabases for a panel
what the mean assumes
that reads land uniformly at random over the target. Lander and Waterman’s 1988 analysis makes that assumption explicit and derives the resulting Poisson distribution of depth. Real sequencing violates it systematically, which is why the mean flatters the result
what it ignores
duplicate reads, reads that fail quality filters, reads that do not map, and for targeted sequencing the off-target fraction — which is often a third or more of the data. The figure here is the theoretical ceiling; usable depth after filtering is lower, sometimes by half

Worked example

400 million read pairs of 150 bp, paired-end, over a 3.1 × 10⁹ base haploid genome
Reads = 400 × 10⁶ pairs × 2 mates = 8 × 10⁸ reads
Sequenced bases = 8 × 10⁸ × 150 = 1.2 × 10¹¹ bases, or 120 gigabases
Mean coverage = 1.2 × 10¹¹ ÷ 3.1 × 10⁹ = 38.7×
Had the 400 million been the TOTAL read count rather than read pairs, the answer would be 19.4× — a twofold error that changes the conclusion entirely, and the reason the read type selector exists
Doubling the output to 800 million pairs gives 77.4×, and the same 400 million pairs directed at a 35 Mb exome instead would give 428.6× before any allowance for off-target reads
None of these figures is the depth at any particular base. At a mean of 38.7×, some positions will be covered fifteen times and some ninety, and the number that matters for reporting is the percentage of the target at or above the minimum depth the laboratory requires

Orders of magnitude

ApplicationReads (millions)Read lengthTypeTargetMean coverage
Whole genome400150Paired-end3.1e938.7×
Whole genome, deeper800150Paired-end3.1e977.4×
Exome50150Paired-end3.5e7428.6×
Gene panel2150Paired-end5e51,200.0×
Single-end run30100Single-end3.5e785.7×
Illustrative configurations, not recommendations. Note the exome and panel figures are theoretical: a substantial fraction of reads in any capture or amplicon experiment falls outside the target, so the achieved on-target depth is lower, often by a third or more.

Why the mean is the wrong number to report

ProblemEffect on coverageWhat to report instead
GC-rich regionsSystematically under-covered — amplification and capture both favour moderate GC content, so first exons and promoter-adjacent sequence are among the worst-covered parts of an exomePercentage of the target at or above a stated minimum depth
Repetitive and homologous sequenceReads map ambiguously or not at all, so depth after filtering can be near zero in regions where the raw data looks adequatePer-region depth, with known problem regions listed explicitly
Capture or amplicon imbalanceSome baits and some amplicons work far better than others, producing a long low tail that a mean conceals entirelyThe coverage distribution, or its low percentiles
Duplicate readsInflate the mean without adding independent observations, since duplicates are copies of the same original moleculeDepth after duplicate removal, and the duplicate rate alongside it
A mean of 100× is compatible with 2% of the target being uncovered, and it is precisely the clinically interesting 2% — GC-rich first exons, segmental duplications — that fails. This is why coverage metrics in clinical reporting are stated as a percentage of the target at or above a threshold.

The mean is cheap; the distribution is what matters

The arithmetic of sequencing depth is as simple as it looks. Multiply the number of reads by the read length to get the total bases sequenced, divide by the size of the target, and the result is the mean number of times each base of the target has been read. Lander and Waterman set this out in 1988 for physical mapping, and the relation carries over unchanged to short-read sequencing. The only arithmetic trap is the factor of two: paired-end sequencing produces two reads per fragment, both of which contribute sequence, and instrument reports variously quote read pairs or total reads. Getting that wrong is a twofold error in the answer.

The more important point is what the mean does not tell you. The Lander-Waterman calculation assumes reads fall uniformly at random across the target, which real sequencing does not do. GC-rich sequence is systematically under-represented, because both PCR amplification and hybridisation capture disfavour extremes of base composition — and GC-rich sequence is not randomly distributed either, being concentrated in first exons and promoter-adjacent regions. Repetitive and homologous sequence attracts reads that cannot be placed unambiguously, so depth after mapping quality filtering collapses in regions where the raw data looked fine. Capture baits and amplicons differ in efficiency by orders of magnitude, giving a long tail of poorly covered target that a mean conceals completely.

The consequence is that a mean coverage figure is a statement about how much sequencing was bought, not about what was achieved. A target with a mean of 100× can still have two per cent of its bases essentially uncovered, and the failing two per cent is not a random sample: it is the GC-rich and repetitive fraction, which is disproportionately the clinically interesting fraction. This is why coverage in a clinical context is reported as the percentage of the target at or above a stated minimum depth — the proportion of target bases at 20× or more, say — and why a report that quotes only a mean has not answered the question that matters. The threshold itself is set by what the assay must detect, and a laboratory’s validation establishes it.

Typical depths are best treated as orders of magnitude rather than standards. Whole-genome germline sequencing is commonly run around 30×, which is enough for heterozygous single-nucleotide variants in well-behaved sequence and marginal for much else. Exomes and panels are run considerably deeper, partly because depth is cheap over a small target and partly because capture uniformity is poorer, so a higher mean is needed to lift the low tail above a usable threshold. Somatic work is deeper again by an order of magnitude or more, because the variant may be present in a small minority of molecules — and at those depths the limiting factor stops being depth and becomes the error rate of the chemistry itself.

Frequently asked questions

How do you calculate sequencing coverage depth?

Mean coverage = (number of reads × read length) ÷ target size in bases. For 400 million paired-end read pairs of 150 bp over a 3.1 Gb genome: 800 million reads × 150 = 120 Gb of sequence, divided by 3.1 Gb, gives 38.7×.

Do paired-end reads count once or twice?

Twice. Each pair produces two reads of the stated length and both contribute sequence to the target, so a figure quoted in read pairs must be doubled. A run described as 2 × 150 bp gives two 150-base reads per fragment, not one of 300 bases.

Why is mean coverage a poor summary?

Because coverage is not evenly distributed. GC-rich and repetitive regions are systematically under-covered, and capture or amplicon imbalance produces a long low tail that a mean conceals. A target with a mean of 100× can have a few per cent of bases effectively uncovered, and those are often the clinically important ones.

What should be reported instead of mean coverage?

The percentage of the target at or above a stated minimum depth — the proportion of bases at 20× or more, for example — together with the depth achieved after duplicate removal and the regions known to fail. That answers whether a negative result can be trusted; a mean does not.

What coverage depth is typical?

As orders of magnitude rather than standards: whole-genome germline sequencing is commonly run around 30×, exomes and panels considerably deeper because capture uniformity is poorer, and somatic variant detection deeper again because the variant may be in a small minority of molecules. The right depth follows from what the assay must detect.

Related calculators

References

  1. Lander ES, Waterman MS. Genomic mapping by fingerprinting random clones: a mathematical analysis. Genomics. 1988;2(3):231–239.
  2. Sims D, Sudbery I, Ilott NE, Heger A, Ponting CP. Sequencing depth and coverage: key considerations in genomic analyses. Nat Rev Genet. 2014;15(2):121–132.
  3. Rehm HL, Bale SJ, Bayrak-Toydemir P, et al. ACMG clinical laboratory standards for next-generation sequencing. Genet Med. 2013;15(9):733–747.
  4. Meynert AM, Ansari M, FitzPatrick DR, Taylor MS. Variant detection sensitivity and biases in whole genome and exome sequencing. BMC Bioinformatics. 2014;15:247.

Medical Disclaimer: The tools and content provided here are for educational and reference purposes only. They are not intended to substitute for professional medical advice, diagnosis, or treatment. Clinical decisions should always be based on the comprehensive assessment of a qualified healthcare professional.