PHQ-2 Depression Screen Calculator

PHQ-2 Depression Screen Calculator

Score the two PHQ-2 items out of 6 against the cut-off of 3. The validation gives two different sensitivity and specificity pairs for two different target conditions, and its own abstract and table disagree.

PHQ-2 total

2 items, 0 to 6
Anhedonia. One of the two cardinal symptoms, and the reason a two-item instrument works at all: the PHQ-2 keeps the two items that carry the diagnosis and discards the seven that carry severity.
Depressed mood. Endorsing either item at any frequency scores at least 1, so a total of 1 or 2 is common and sits below the cut-off of 3.
3pointsExample

Anhedonia: more than half the days (2). Depressed mood: several days (1)

Advertisement

Scoring

PHQ-2 = anhedonia (0 to 3) + depressed mood (0 to 3), each scored not at all, several days, more than half the days or nearly every day over the last two weeks
Range 0 to 6 · cut-off 3 or more
two target conditions, two different answers
Table 2 of the 2003 paper gives 82.9 per cent sensitivity and 90.0 per cent specificity for MAJOR depressive disorder (likelihood ratio 2.9) and 62.3 per cent and 95.4 per cent for ANY depressive disorder (likelihood ratio 5.4). They are not interchangeable, and a page quoting one figure without saying which condition it belongs to is quoting half a result
the source disagrees with itself
the paper’s abstract prints a specificity of 92 per cent at “PHQ-2 score >3” where its own Table 2 prints 90.0 per cent at 3 or more, and its discussion cites 585 patients in the interview sample where the methods and the tables use 580. This page uses the table’s figures and prints the discrepancy rather than reconciling it silently
the cohort’s own prevalences
7 per cent for major depressive disorder and 18 per cent for any depressive disorder, in the 580 interviewed patients. Those are the prevalences the predictive values on this page are computed at, so the first number a reader meets is the instrument’s own rather than an invented one
the ninth item is not here
the PHQ-2 is items 1 and 2 of the PHQ-9 and contains no question about self-harm. A reader using it as a first stage should know that the item many people associate with the PHQ-9 is absent from the short form

Worked example

Anhedonia: more than half the days (2). Depressed mood: several days (1)
2 + 1 = 3 points, which reaches the cut-off of 3 or more
Apply the table's 82.9 per cent sensitivity and 90.0 per cent specificity at the cohort's own 7 per cent major-depression prevalence: of 1,000 people, 58 true positives against 93 false ones — a positive predictive value of 38.4 per cent
At a prevalence of 20 per cent the same pair gives 67.5 per cent
Now switch target condition. For ANY depressive disorder the pair is 62.3 and 95.4, and at the cohort's 18 per cent prevalence the positive predictive value is 74.8 per cent — higher than for major depression, from a LOWER sensitivity, because the specificity and the prevalence both rose
Which is why the condition a screening figure belongs to has to be printed next to it
Advertisement

The same cut-off against two target conditions

Target conditionPrevalence in the cohortSensitivitySpecificityLikelihood ratio
Major depressive disorder7 per cent82.9 per cent90.0 per cent2.9
Any depressive disorder18 per cent62.3 per cent95.4 per cent5.4
Both rows are a PHQ-2 of 3 or more in the same 580 patients. Widening the target condition lowers the sensitivity by twenty points and raises the likelihood ratio, because the broader condition includes milder states that two items do not catch.

Positive predictive value at four prevalences

Target and prevalencePositive predictive valueNegative predictive value
Major depression, 20 in 10067.5 per cent95.5 per cent
Major depression, 7 in 100 (the cohort’s own)38.4 per cent98.6 per cent
Any depressive disorder, 18 in 100 (the cohort’s own)74.8 per cent92.0 per cent
Any depressive disorder, 5 in 10041.6 per cent98.0 per cent
Computed from a 2 × 2 on the figures above. At the cohort’s own major-depression prevalence, fewer than two in five positives have major depression — which is what a two-item screen is for: a first stage, not an answer.

A first stage, and the figure that depends on which question you asked

The PHQ-2 is the first two items of the PHQ-9 — anhedonia and depressed mood — scored 0 to 3 each over two weeks, for a total of 0 to 6, with a cut-off of 3 or more. Kroenke, Spitzer and Williams validated it in 2003 in 6,000 patients: 3,000 in 8 primary care clinics and 3,000 in 7 obstetrics-gynaecology clinics across 8 states and the District of Columbia, with a blinded structured telephone interview by a PhD clinical psychologist or a senior psychiatric social worker in 580 of them as the criterion standard.

Its central result is not one sensitivity and specificity pair but two, and they belong to different questions. For major depressive disorder, a total of 3 or more carried 82.9 per cent sensitivity and 90.0 per cent specificity, with a likelihood ratio of 2.9. For any depressive disorder — a broader target that includes milder states — the same cut-off in the same patients carried 62.3 per cent sensitivity and 95.4 per cent specificity, with a likelihood ratio of 5.4. Widening the target cost twenty points of sensitivity and bought five of specificity. A figure quoted without its target condition is half a result, and both halves are on this page.

The paper also disagrees with itself in two places, and both are printed here rather than quietly reconciled. Its abstract gives a specificity of 92 per cent at a “PHQ-2 score >3” where Table 2 gives 90.0 per cent at 3 or more; and its discussion cites 585 interviewed patients where the methods and the tables use 580. This page uses the table.

What follows from all of that is how a two-item screen should be used. At the cohort’s own 7 per cent major-depression prevalence, a total of 3 or more carries a positive predictive value of 38.4 per cent: fewer than two positives in five have major depression. That is not a failure of the instrument, it is what a first stage is. A positive PHQ-2 is a reason to ask the remaining seven PHQ-9 items, and one of those seven is the self-harm item that is not in the short form at all.

Attribution, as the instrument’s own notice requires. Developed by Drs Robert L. Spitzer, Janet B.W. Williams, Kurt Kroenke and colleagues, with an educational grant from Pfizer Inc. The published form carries “No permission required to reproduce, translate, display or distribute”, and that notice is the reason this page exists while no page here reproduces the AUDIT.

A total on this page is a number, not a diagnosis. These instruments quantify what a person reports, or what an observer records at one moment; a diagnosis rests on a clinical assessment that no questionnaire total stands in for. A low total does not exclude what the instrument screens for. No sensitivity quoted here is 1, every figure was measured in a published cohort rather than in the person in front of you, and someone who endorses nothing may still have the disorder. This page renders no dose, no medication, no treatment regimen and no disposition. It computes the published total and prints the published thresholds with the body that published each one; what follows from the number is a clinical decision this page does not make. Every cut-off, sensitivity, specificity and severity band here is printed with the cohort it was measured in, or with a statement that no source read for this page attaches one — because an instrument’s accuracy is a property of the population it was measured in and not of the instrument. If you are reading this about yourself and you are in distress or thinking about harming yourself, please contact your local emergency number or a crisis line now rather than treating a number as an answer: 999 or Samaritans on 116 123 in the UK and Ireland, 988 in the United States and Canada, or your local emergency service elsewhere.

Frequently asked questions

What is a positive PHQ-2 score?

Three or more out of six. In the 2003 validation that cut-off gave 82.9 per cent sensitivity and 90.0 per cent specificity for major depressive disorder in 580 patients interviewed by a mental health professional, drawn from a cohort of 6,000.

Why are two different sensitivities quoted for the PHQ-2?

Because the paper reports two target conditions. For major depressive disorder the pair is 82.9 and 90.0; for any depressive disorder, in the same patients at the same cut-off, it is 62.3 and 95.4. The broader condition includes milder states that two items miss, so the sensitivity falls and the specificity rises.

Does a PHQ-2 below 3 mean someone is not depressed?

A low total does not exclude what the instrument screens for. No sensitivity quoted here is 1, every figure was measured in a published cohort rather than in the person in front of you, and someone who endorses nothing may still have the disorder. At 62.3 per cent sensitivity for any depressive disorder, more than a third of those cases score below the cut-off. A total on this page is a number, not a diagnosis. These instruments quantify what a person reports, or what an observer records at one moment; a diagnosis rests on a clinical assessment that no questionnaire total stands in for.

Does the PHQ-2 ask about self-harm?

No. It is items 1 and 2 of the PHQ-9 only — anhedonia and depressed mood — and contains no question about thoughts of self-harm or of being better off dead. If you are reading this about yourself and you are in distress or thinking about harming yourself, please contact your local emergency number or a crisis line now rather than treating a number as an answer: 999 or Samaritans on 116 123 in the UK and Ireland, 988 in the United States and Canada, or your local emergency service elsewhere.

How much is a positive PHQ-2 worth in general practice?

At the validation cohort’s own major-depression prevalence of 7 per cent, a total of 3 or more carries a positive predictive value of 38.4 per cent. At 20 per cent it carries 67.5 per cent. The instrument is identical; the population is what changed.

Related calculators

References

  1. Kroenke K, Spitzer RL, Williams JB. The Patient Health Questionnaire-2: validity of a two-item depression screener. Med Care. 2003;41(11):1284–1292. Full text read, including Tables 1 to 4: 6,000 patients, 3,000 in 8 primary care clinics and 3,000 in 7 obstetrics-gynaecology clinics across 8 states and the District of Columbia; the criterion standard a blinded structured telephone interview by a PhD clinical psychologist or senior psychiatric social worker in 580 patients; at a cut-off of 3 or more, sensitivity 82.9 per cent and specificity 90.0 per cent for major depressive disorder (likelihood ratio 2.9) and sensitivity 62.3 per cent and specificity 95.4 per cent for any depressive disorder (likelihood ratio 5.4). Note for a future author: the paper’s own abstract prints a specificity of 92 per cent at “PHQ-2 score >3” where Table 2 prints 90.0 per cent at 3 or more, and the discussion cites 585 patients in the interview sample where the methods and tables use 580.
  2. Patient Health Questionnaire-2 (PHQ-2), National HIV Curriculum, University of Washington. Source of the two items, the response values, the 0-to-6 range, the cut-off of 3 or more, and the cohort PREVALENCES against which the operating characteristics were measured — 7 per cent for major depressive disorder and 18 per cent for any depressive disorder, in the 580 patients who had an independent mental health professional interview.
  3. Patient Health Questionnaire (PHQ-9) and Generalized Anxiety Disorder 7-item scale (GAD-7), the combined published form as distributed by the North Dakota Department of Health and Human Services. Source of every item of both instruments in its exact wording, the four response options and their point values, the shared functional-impairment question, and the developer notice: “Developed by Drs. Robert L. Spitzer, Janet B.W. Williams, Kurt Kroenke and colleagues, with an educational grant from Pfizer Inc. No permission required to reproduce, translate, display or distribute, 1999.”

Not medical advice. For healthcare professionals and education. Reference intervals vary by laboratory and assay — always use your own laboratory's. Never base a dose or a treatment decision on this page alone. Full disclaimer at calcengines.com/disclaimer/