DNA Copy Number Calculator

DNA Copy Number Calculator

Convert a mass of double-stranded DNA into a number of molecules, using 660 g/mol per base pair and Avogadro’s number — the calculation behind every qPCR standard curve made from a plasmid or an amplicon.

DNA Copy Number

Mass, length → copies
The mass in the tube, in nanograms. For a standard curve this is the mass you are about to dilute, measured fluorometrically rather than by absorbance wherever possible.
The full length of the molecule, not the length of the amplicon within it. For a linearised plasmid carrying a 200 bp insert, the length is the whole plasmid. For human genomic DNA use about 3.1e9 for the haploid genome.
304,141,414copiesExample

1 ng of a linearised 3,000 bp plasmid

Formula

copies = (mass in ng × 6.022e23) ÷ (length in bp × 1e9 × 660)
6.022e23
Avogadro's number — molecules per mole. The calculator's parser reads scientific notation directly, so 6.022e23 and 1e9 can be typed as they stand
660 g/mol
the average molecular weight of one base pair of double-stranded DNA. It is an average across base composition: an AT pair and a GC pair differ, so the result is a good estimate rather than an exact count. For single-stranded DNA the usual average is about 330 g/mol per base
1e9
converts nanograms to grams. It is folded into the denominator so the formula can be typed exactly as written above
length
the length of the whole molecule. For a plasmid standard, that is the entire plasmid including the vector backbone, not the amplicon it carries — using the amplicon length is the commonest error on this page and inflates the copy number by whatever factor the backbone contributes
what you get
copies in the mass you entered. Divide by the volume that mass is dissolved in to get copies per microlitre, which is what a dilution series is actually built from

Worked example

1 ng of a linearised 3,000 bp plasmid
Numerator: 1 × 6.022e23 = 6.022e23
Denominator: 3,000 × 1e9 × 660 = 1.98e15
6.022e23 ÷ 1.98e15 = 3.04 × 10⁸ copies, about 304 million
Dissolved in 100 µL, that stock holds roughly 3.04 × 10⁶ copies per microlitre — the top of a standard curve, from which a tenfold series runs down to single figures in six steps
The same 1 ng of human genomic DNA, taking the haploid genome as 3.1e9 bp, gives only about 294 haploid genome equivalents — so a single-copy locus is present at about 294 copies per nanogram of haploid genome, or roughly 150 diploid cells' worth

Copies per nanogram at different lengths

MoleculeLength (bp)Copies per ng
Short amplicon1009.1 × 10⁹
Long amplicon1,0009.1 × 10⁸
Small plasmid standard3,0003.0 × 10⁸
Typical cloning plasmid5,0001.8 × 10⁸
Large construct10,0009.1 × 10⁷
Human haploid genome3.1 × 10⁹294
Copy number is inversely proportional to length, which is why the length used has to be the whole molecule. Quantifying a plasmid but entering only the insert length overstates the copy number by the ratio of the two.

Where the estimate loses precision

AssumptionEffect
660 g/mol is an average per base pairA very GC-rich or AT-rich sequence departs from it by a few per cent. Immaterial for a standard curve, material if you are claiming an absolute molecule count
The mass is correctCopy number inherits every error in the quantification. A260 overstates double-stranded DNA whenever RNA or single-stranded material is present, so a fluorometric measurement is preferable
The whole mass is the molecule you think it isResidual RNA, genomic carryover in a plasmid prep or a partial digest all contribute mass without contributing copies of the target
The molecules survive dilutionAt high dilution, DNA adsorbs to plastic. Carrier nucleic acid or a low-bind tube matters far more at the bottom of a standard curve than the arithmetic does
The formula is exact; the inputs are not. Treat the output as a well-founded estimate of molecule number, which is what a standard curve needs, rather than as a count.

From nanograms to molecules

A standard curve made from a plasmid or a purified amplicon needs to be expressed in copies, because that is what the y-intercept of the curve then means and what an unknown sample is reported against. The conversion from mass to molecules is elementary chemistry: divide the mass by the molecular weight to get moles, then multiply by Avogadro's number. The only piece of biology in it is the molecular weight, and for double-stranded DNA that is taken as 660 g/mol per base pair.

That figure is an average. An A–T pair and a G–C pair have different masses, so a sequence with unusual base composition departs from 660 by a few per cent. Over a plasmid of several thousand base pairs the composition averages out and the error is negligible against the uncertainty in the mass measurement. Over a fifty-base oligonucleotide it does not average out at all, and the supplier's calculated molecular weight should be used instead. For single-stranded DNA the corresponding average is about 330 g/mol per base, which is simply the same figure per strand.

The commonest mistake is not in the arithmetic but in the length. The length required is that of the whole molecule being weighed, which for a plasmid standard means the entire construct including the vector backbone. Entering the length of the amplicon inside it, which is the number most people have to hand, overstates the copy number by the ratio of plasmid to insert — often a factor of ten or more, and invisible afterwards because the curve still looks perfectly linear. Linearising the plasmid before making the dilution series is standard practice for a separate reason: supercoiled template amplifies less efficiently than linear template, which shifts the curve without changing its shape.

Two practical consequences follow from the size of the numbers. One nanogram of a 3,000 bp plasmid contains about 3 × 10⁸ molecules, so the stock has to be diluted many thousandfold before it belongs anywhere near a reaction — and an aerosol from that tube carries more template than any patient sample will, which is why plasmid standards are handled in a separate room from extraction and set-up. At the other end, when the calculation says a well should contain fewer than a few hundred copies, sampling becomes stochastic: the Poisson variation in how many molecules are actually pipetted, not the assay, sets the reproducibility. That is the real floor of a standard curve, and no amount of replicate averaging moves it.

Frequently asked questions

How do you calculate DNA copy number from nanograms?

Copies = (mass in ng × 6.022e23) ÷ (length in bp × 1e9 × 660). The 660 is the average molecular weight of a base pair of double-stranded DNA and the 1e9 converts nanograms to grams. One nanogram of a 3,000 bp plasmid is about 3.04 × 10⁸ copies.

Why 660 g/mol per base pair?

It is the average mass of one base pair of double-stranded DNA including the sugar-phosphate backbone and the counter-ion, averaged across base composition. Because it is an average, the result is an estimate rather than an exact molecule count, though the error is small for anything longer than a few hundred base pairs.

Should I use the plasmid length or the amplicon length?

The plasmid length. The mass you measured is the mass of whole plasmid molecules, so the molecular weight has to be that of the whole plasmid. Using the amplicon length inflates the copy number by the ratio of the two, and nothing downstream will reveal the error.

How many copies of a gene are in a nanogram of human DNA?

About 294 haploid genome equivalents per nanogram, taking the haploid genome as 3.1 × 10⁹ base pairs — so roughly 300 copies of a single-copy locus per nanogram of DNA, or about 150 diploid cells' worth.

Why do replicates scatter at the bottom of a standard curve?

Because at a few hundred copies or fewer, the number of molecules actually transferred into each well varies by chance. This Poisson sampling variation is a property of counting small numbers of molecules, not a fault in the assay, and it sets the practical limit of quantification.

Related calculators

References

  1. Bustin SA, Benes V, Garson JA, et al. The MIQE guidelines: minimum information for publication of quantitative real-time PCR experiments. Clin Chem. 2009;55(4):611–622.
  2. Sambrook J, Russell DW. Molecular Cloning: A Laboratory Manual. 3rd ed. Cold Spring Harbor Laboratory Press; 2001.
  3. Dhanasekaran S, Doherty TM, Kenneth J. Comparison of different standards for real-time PCR-based absolute quantification. J Immunol Methods. 2010;354(1–2):34–39.

Medical Disclaimer: The tools and content provided here are for educational and reference purposes only. They are not intended to substitute for professional medical advice, diagnosis, or treatment. Clinical decisions should always be based on the comprehensive assessment of a qualified healthcare professional.