Copy Number from Log2 Ratio Calculator

Copy Number from Log2 Ratio Calculator

Copy number as 2 × 2^log2 ratio for an autosomal locus — and the purity correction that has to be applied before that reading means anything in a tumour sample.

Copy Number from Log2 Ratio

Log2 ratio, purity → copies
The segment-level log2 copy ratio from the array or the sequencing pipeline. It is a ratio against a reference and is normally median-centred across the autosomes, so it describes the segment relative to the sample's own average rather than against an absolute two copies.
The percentage of nuclei in the sequenced material that are tumour. Take it from a pathologist's estimate on the same block, or from an allele-fraction-based estimate in the sequencing data — the two often disagree, and the disagreement propagates straight into this answer. Enter 100 for a germline sample or for data already rescaled for purity.
3.13copiesExample

An observed log2 ratio of +0.42 in a specimen estimated at 60% tumour, at an autosomal locus

Formula

pure sample: copy number = 2 × 2^(log2 ratio)
with purity p (as a fraction), autosomal locus: copy number = (2 × 2^(log2 ratio) − 2 × (1 − p)) ÷ p
and forwards: observed log2 ratio = log₂( (p × CN + 2 × (1 − p)) ÷ 2 )
the pure-sample form
a log2 ratio is log₂(copies ÷ 2) for an autosomal locus in a diploid genome, so copies = 2 × 2^log2. That gives 0 for two copies, −1 for one, +0.585 for three and +1 for four — the values a clean germline sample produces
the purity correction
derived, not asserted. The observed ratio is a mixture: p of the DNA carries CN copies and (1 − p) carries the normal two, so 2^log2 = (p × CN + 2(1 − p)) ÷ 2. Rearranging for CN gives the expression above, and at p = 1 it collapses back to 2 × 2^log2
why it matters
THE POINT OF THE PAGE: normal DNA pulls every observed ratio towards zero. A single-copy loss in a 50% pure sample reads as three-quarters of the neutral coverage, a log2 of −0.415 rather than −1. Read without the correction, the segment looks like a shallow dip instead of a deletion
ploidy
the second thing that breaks the simple form. Log2 ratios are median-centred, so they measure a segment against the sample's own average. In a whole-genome-doubled tumour the baseline is four copies, a truly diploid region reads as a loss, and a gain to six copies reads as a modest rise. Ploidy must be established independently
why a homozygous deletion is not −∞
because the normal cells are still there. At zero tumour copies the observed ratio bottoms out at log₂(1 − p) — about −0.74 at 40% purity, −1.32 at 60% and −2.32 at 80%. A finite floor, set by contamination before mapping noise or background is considered at all

Worked example

An observed log2 ratio of +0.42 in a specimen estimated at 60% tumour, at an autosomal locus
Copy ratio = 2^0.42 = 1.338, so the naive reading is 2 × 1.338 = 2.68 copies — an awkward, equivocal number that is not obviously a gain
Now correct it. The sample is 60% tumour, so 40% of the DNA contributes the normal two copies: numerator = 2 × 1.338 − 2 × 0.40 = 2.676 − 0.800 = 1.876
Divide by the purity: 1.876 ÷ 0.60 = 3.13 copies in the tumour cells. A single-copy gain, which the uncorrected figure obscured
Run it backwards as a check. Three copies at 60% purity would give an observed ratio of log₂((0.6 × 3 + 2 × 0.4) ÷ 2) = log₂(1.30) = +0.379, and the observed +0.42 sits just above it. Consistent
The same arithmetic at lower purity is more dramatic. An observed +0.30 in a sample that is only 40% tumour reads naively as 2.46 copies and corrects to 3.16 — the entire gain was being hidden by the stroma
And in the other direction: a segment with no copies at all in a 60% pure sample cannot read below log₂(0.40) = −1.32. A homozygous deletion in an impure sample is a shallow trough, not a hole

Expected log2 ratios in a clean diploid sample

Copy numberCopy ratiolog2 ratio
0 — homozygous deletion0−∞ in theory; in practice bounded by contamination, mapping noise and background
1 — single-copy loss1/2−1.000
2 — copy neutral2/20.000
3 — single-copy gain3/2+0.585
4 — two-copy gain4/2+1.000
66/2+1.585
88/2+2.000
These are the only values a perfectly pure, perfectly diploid, perfectly clonal sample would produce. Real profiles do not land on them, and the distance from them is not measurement error: it is the combined effect of normal-cell contamination, subclonality and whatever the sample's actual ploidy is.

The same copy number seen through different purities

Tumour purityObserved log2 at CN 0at CN 1at CN 3at CN 4at CN 8
100%−∞−1.000+0.585+1.000+2.000
80%−2.322−0.737+0.485+0.848+1.766
60%−1.322−0.515+0.379+0.678+1.485
40%−0.737−0.322+0.263+0.485+1.138
20%−0.322−0.152+0.138+0.263+0.678
Every column compresses towards zero as purity falls, and the compression is not uniform — deletions flatten faster than amplifications. At 20% tumour, a homozygous deletion and a single-copy loss differ by 0.17 log2 units, which is inside the noise of most assays. This is why low-purity specimens produce copy-number profiles that look quiet rather than wrong.

Why an array log2 and a sequencing log2 are not the same number

IssueConsequence
Different referencesA log2 ratio is always against something: a co-hybridised or pooled reference on an array, a panel of normal samples for sequencing. Two platforms with different references produce different values for the same segment
Median centringSequencing log2 values are centred on the median depth across the autosomes, so they describe a segment relative to the sample's own average. A genome-doubled tumour therefore reads as flat, and its truly diploid regions read as losses
Different sources of biasArrays carry probe-level hybridisation and GC effects; sequencing carries amplification bias, capture efficiency and mapping quality. Neither is a scaled version of the other
Measured disagreementComparing a targeted-sequencing caller against array CGH at the same genes, the 95% prediction intervals for the deviation ran from 0.287 to 0.464 log2 units — and 0.3 log2 units is a factor of 1.23 in copy ratio
What followsThresholds are platform-specific. A log2 cut-off validated on one assay cannot be carried across to another, and serial samples should be compared on the same platform or not at all
A deviation of three tenths of a log2 unit is not a rounding difference: at a true copy number of three it is the difference between calling a gain and calling the segment neutral. Copy-number thresholds belong to the assay that was validated, not to the quantity.

Two copies, one exponent, and everything the exponent hides

A log2 ratio is a comparison, and the comparison is what makes it convenient. For an autosomal locus in a diploid genome, the ratio of observed to expected signal is the copy number divided by two, and taking the base-two logarithm puts loss and gain symmetrically either side of zero: one copy is −1, two copies is 0, three copies is +0.585 and four is +1. Inverting it is a single line — copy number equals two times two raised to the log2 ratio — and for a germline sample, a cell line or anything else that is genetically uniform, that line is the whole story.

A tumour specimen is not genetically uniform, and that is where the simple form fails. What was sequenced was a mixture: tumour nuclei carrying whatever copy number the segment has, and stroma, immune cells and vasculature carrying the normal two. The observed ratio is the weighted average, so every gain is pulled down towards zero and every loss is pulled up towards it. The algebra is short. If a fraction p of the DNA is tumour with copy number CN, then 2^log2 = (p × CN + 2(1 − p)) ÷ 2, and rearranging gives copy number = (2 × 2^log2 − 2(1 − p)) ÷ p, which collapses back to the simple form at p = 1. The worked example shows what it is worth: an observed +0.42 in a 60% pure sample reads naively as 2.68 copies, an equivocal number that nobody would call a gain, and corrects to 3.13 — a clean single-copy gain that the stroma was concealing. The effect scales with impurity. A documented illustration: a single-copy loss in a 50% pure sample has three-quarters of the coverage of a neutral site, a log2 of −0.415 rather than −1.

The same relation explains why a homozygous deletion is never the minus infinity the algebra promises. Run the mixture equation forwards with CN set to zero and the observed ratio is log₂(1 − p): −0.74 at 40% purity, −1.32 at 60%, −2.32 at 80%. The normal cells put a floor under the signal, and that floor is reached before mapping noise, background from repetitive sequence or uneven coverage are considered at all. It also explains why low-purity profiles look quiet rather than obviously wrong: at 20% tumour, a homozygous deletion and a single-copy loss are separated by 0.17 log2 units, which is inside the noise of most assays, so the profile flattens instead of failing.

Two further cautions belong with the arithmetic. The first is ploidy. Log2 ratios are centred on the median across the autosomes, which means they describe a segment relative to the sample's own average rather than against an absolute two copies. In a tumour that has undergone whole-genome doubling the flat baseline is four copies, a genuinely diploid region reads as a loss, and a gain to six copies reads as a modest rise — so purity and ploidy have to be estimated together, which is what the dedicated allele-specific callers exist to do, and neither can be recovered from a single segment's log2 value. The second is that log2 ratios are not portable between platforms. They are ratios against a reference, and the reference differs; the biases differ too, hybridisation and probe effects on an array against amplification and capture and mapping on a sequencer. The size of the disagreement has been measured: comparing a targeted-sequencing caller with array comparative genomic hybridisation at the same genes, the 95% prediction intervals for the deviation ran from 0.287 to 0.464 log2 units, and three tenths of a log2 unit is a factor of 1.23 in copy ratio — at a true copy number of three, the difference between calling a gain and calling the segment neutral. Copy-number thresholds therefore belong to the assay that validated them. And in every case the copy number is only half the picture: what a deletion means depends on what sits on the surviving allele, which is a question for the variant allele fraction rather than for this page.

Frequently asked questions

How do you convert a log2 ratio to copy number?

For an autosomal locus in a diploid genome, copy number = 2 × 2^(log2 ratio). That gives 0 for two copies, −1 for one and +0.585 for three. In a tumour sample the result must then be corrected for purity, because normal DNA in the specimen pulls every observed ratio towards zero.

How does tumour purity change the calculation?

Copy number = (2 × 2^(log2 ratio) − 2 × (1 − p)) ÷ p, where p is the purity as a fraction. It follows from treating the sample as a mixture of tumour at copy number CN and normal cells at two copies. An observed +0.42 in a 60% pure sample is 2.68 copies uncorrected and 3.13 corrected — a gain that the uncorrected figure hides.

What are the expected log2 values for a diploid sample?

Zero for two copies, −1 for one copy, +0.585 for three, +1 for four and +2 for eight. Those are the values a perfectly pure, perfectly clonal, perfectly diploid sample would produce. Real profiles sit between them, and the gap is contamination, subclonality and ploidy rather than measurement error.

Why is a homozygous deletion not log2 of minus infinity?

Because the normal cells in the specimen still contribute two copies. With no tumour copies at all, the observed ratio bottoms out at log₂(1 − purity): about −0.74 at 40% purity, −1.32 at 60% and −2.32 at 80%. Mapping noise and background raise that floor further.

Does whole-genome doubling break the log2 calculation?

Yes, because log2 ratios are median-centred on the sample's own average. In a genome-doubled tumour the flat baseline is four copies, so a truly diploid region reads as a loss and a gain to six copies looks modest. Purity and ploidy have to be estimated together by an allele-aware caller; neither is recoverable from one segment's log2 value.

Can array and sequencing log2 ratios be compared?

Not directly. They are ratios against different references and carry different biases. Comparing a targeted-sequencing caller with array CGH at the same genes gave 95% prediction intervals for the deviation of 0.287 to 0.464 log2 units, and 0.3 log2 units is a factor of 1.23 in copy ratio. Thresholds belong to the validated assay, and serial samples should stay on one platform.

Related calculators

References

  1. Talevich E, Shain AH, Botton T, Bastian BC. CNVkit: genome-wide copy number detection and visualization from targeted DNA sequencing. PLoS Comput Biol. 2016;12(4):e1004873 — and the CNVkit documentation and source, from which the purity-adjusted relation and its derivation are taken.
  2. Van Loo P, Nordgard SH, Lingjærde OC, et al. Allele-specific copy number analysis of tumors. Proc Natl Acad Sci USA. 2010;107(39):16910–16915 — on estimating purity and ploidy jointly rather than assuming either.
  3. Carter SL, Cibulskis K, Helman E, et al. Absolute quantification of somatic DNA alterations in human cancer. Nat Biotechnol. 2012;30(5):413–421.
  4. Zack TI, Schumacher SE, Carter SL, et al. Pan-cancer patterns of somatic copy number alteration. Nat Genet. 2013;45(10):1134–1140 — on the frequency of whole-genome doubling and arm-level events.
  5. Li MM, Datto M, Duncavage EJ, et al. Standards and guidelines for the interpretation and reporting of sequence variants in cancer. J Mol Diagn. 2017;19(1):4–23.

Medical Disclaimer: The tools and content provided here are for educational and reference purposes only. They are not intended to substitute for professional medical advice, diagnosis, or treatment. Clinical decisions should always be based on the comprehensive assessment of a qualified healthcare professional.