DNA Copy Number Calculator
DNA Copy Number Calculator
Convert a mass of double-stranded DNA into a number of molecules, using 660 g/mol per base pair and Avogadro’s number — the calculation behind every qPCR standard curve made from a plasmid or an amplicon.
DNA Copy Number
Mass, length → copies1 ng of a linearised 3,000 bp plasmid
Formula
- 6.022e23
- Avogadro's number — molecules per mole. The calculator's parser reads scientific notation directly, so 6.022e23 and 1e9 can be typed as they stand
- 660 g/mol
- the average molecular weight of one base pair of double-stranded DNA. It is an average across base composition: an AT pair and a GC pair differ, so the result is a good estimate rather than an exact count. For single-stranded DNA the usual average is about 330 g/mol per base
- 1e9
- converts nanograms to grams. It is folded into the denominator so the formula can be typed exactly as written above
- length
- the length of the whole molecule. For a plasmid standard, that is the entire plasmid including the vector backbone, not the amplicon it carries — using the amplicon length is the commonest error on this page and inflates the copy number by whatever factor the backbone contributes
- what you get
- copies in the mass you entered. Divide by the volume that mass is dissolved in to get copies per microlitre, which is what a dilution series is actually built from
Worked example
1 ng of a linearised 3,000 bp plasmid
Numerator: 1 × 6.022e23 = 6.022e23
Denominator: 3,000 × 1e9 × 660 = 1.98e15
6.022e23 ÷ 1.98e15 = 3.04 × 10⁸ copies, about 304 million
Dissolved in 100 µL, that stock holds roughly 3.04 × 10⁶ copies per microlitre — the top of a standard curve, from which a tenfold series runs down to single figures in six steps
The same 1 ng of human genomic DNA, taking the haploid genome as 3.1e9 bp, gives only about 294 haploid genome equivalents — so a single-copy locus is present at about 294 copies per nanogram of haploid genome, or roughly 150 diploid cells' worth
Copies per nanogram at different lengths
| Molecule | Length (bp) | Copies per ng |
|---|---|---|
| Short amplicon | 100 | 9.1 × 10⁹ |
| Long amplicon | 1,000 | 9.1 × 10⁸ |
| Small plasmid standard | 3,000 | 3.0 × 10⁸ |
| Typical cloning plasmid | 5,000 | 1.8 × 10⁸ |
| Large construct | 10,000 | 9.1 × 10⁷ |
| Human haploid genome | 3.1 × 10⁹ | 294 |
Where the estimate loses precision
| Assumption | Effect |
|---|---|
| 660 g/mol is an average per base pair | A very GC-rich or AT-rich sequence departs from it by a few per cent. Immaterial for a standard curve, material if you are claiming an absolute molecule count |
| The mass is correct | Copy number inherits every error in the quantification. A260 overstates double-stranded DNA whenever RNA or single-stranded material is present, so a fluorometric measurement is preferable |
| The whole mass is the molecule you think it is | Residual RNA, genomic carryover in a plasmid prep or a partial digest all contribute mass without contributing copies of the target |
| The molecules survive dilution | At high dilution, DNA adsorbs to plastic. Carrier nucleic acid or a low-bind tube matters far more at the bottom of a standard curve than the arithmetic does |
From nanograms to molecules
A standard curve made from a plasmid or a purified amplicon needs to be expressed in copies, because that is what the y-intercept of the curve then means and what an unknown sample is reported against. The conversion from mass to molecules is elementary chemistry: divide the mass by the molecular weight to get moles, then multiply by Avogadro's number. The only piece of biology in it is the molecular weight, and for double-stranded DNA that is taken as 660 g/mol per base pair.
That figure is an average. An A–T pair and a G–C pair have different masses, so a sequence with unusual base composition departs from 660 by a few per cent. Over a plasmid of several thousand base pairs the composition averages out and the error is negligible against the uncertainty in the mass measurement. Over a fifty-base oligonucleotide it does not average out at all, and the supplier's calculated molecular weight should be used instead. For single-stranded DNA the corresponding average is about 330 g/mol per base, which is simply the same figure per strand.
The commonest mistake is not in the arithmetic but in the length. The length required is that of the whole molecule being weighed, which for a plasmid standard means the entire construct including the vector backbone. Entering the length of the amplicon inside it, which is the number most people have to hand, overstates the copy number by the ratio of plasmid to insert — often a factor of ten or more, and invisible afterwards because the curve still looks perfectly linear. Linearising the plasmid before making the dilution series is standard practice for a separate reason: supercoiled template amplifies less efficiently than linear template, which shifts the curve without changing its shape.
Two practical consequences follow from the size of the numbers. One nanogram of a 3,000 bp plasmid contains about 3 × 10⁸ molecules, so the stock has to be diluted many thousandfold before it belongs anywhere near a reaction — and an aerosol from that tube carries more template than any patient sample will, which is why plasmid standards are handled in a separate room from extraction and set-up. At the other end, when the calculation says a well should contain fewer than a few hundred copies, sampling becomes stochastic: the Poisson variation in how many molecules are actually pipetted, not the assay, sets the reproducibility. That is the real floor of a standard curve, and no amount of replicate averaging moves it.
Frequently asked questions
How do you calculate DNA copy number from nanograms?
Copies = (mass in ng × 6.022e23) ÷ (length in bp × 1e9 × 660). The 660 is the average molecular weight of a base pair of double-stranded DNA and the 1e9 converts nanograms to grams. One nanogram of a 3,000 bp plasmid is about 3.04 × 10⁸ copies.
Why 660 g/mol per base pair?
It is the average mass of one base pair of double-stranded DNA including the sugar-phosphate backbone and the counter-ion, averaged across base composition. Because it is an average, the result is an estimate rather than an exact molecule count, though the error is small for anything longer than a few hundred base pairs.
Should I use the plasmid length or the amplicon length?
The plasmid length. The mass you measured is the mass of whole plasmid molecules, so the molecular weight has to be that of the whole plasmid. Using the amplicon length inflates the copy number by the ratio of the two, and nothing downstream will reveal the error.
How many copies of a gene are in a nanogram of human DNA?
About 294 haploid genome equivalents per nanogram, taking the haploid genome as 3.1 × 10⁹ base pairs — so roughly 300 copies of a single-copy locus per nanogram of DNA, or about 150 diploid cells' worth.
Why do replicates scatter at the bottom of a standard curve?
Because at a few hundred copies or fewer, the number of molecules actually transferred into each well varies by chance. This Poisson sampling variation is a property of counting small numbers of molecules, not a fault in the assay, and it sets the practical limit of quantification.
Related calculators
References
- Bustin SA, Benes V, Garson JA, et al. The MIQE guidelines: minimum information for publication of quantitative real-time PCR experiments. Clin Chem. 2009;55(4):611–622.
- Sambrook J, Russell DW. Molecular Cloning: A Laboratory Manual. 3rd ed. Cold Spring Harbor Laboratory Press; 2001.
- Dhanasekaran S, Doherty TM, Kenneth J. Comparison of different standards for real-time PCR-based absolute quantification. J Immunol Methods. 2010;354(1–2):34–39.
Medical Disclaimer: The tools and content provided here are for educational and reference purposes only. They are not intended to substitute for professional medical advice, diagnosis, or treatment. Clinical decisions should always be based on the comprehensive assessment of a qualified healthcare professional.
