4AT Delirium Assessment Calculator
4AT Delirium Assessment Calculator
Four items, 0 to 12, with two of them worth 4 points each so either can reach the threshold alone. Pooled sensitivity and specificity are both 0.88 across 17 studies — and at low prevalence most positives are still false.
4AT rapid delirium assessment
4 items, 0 to 12Alertness normal (0), one mistake on the AMT4 (1), starts months backwards but manages fewer than seven (1), and the nursing staff describe a change over the last three days that is still fluctuating today (4)
Scoring
0 · 1–3 · 4 or above
- two items worth 4
- alertness and acute change each score 0 or 4, with nothing in between, so either one on its own reaches the 4-or-above threshold. That is a design decision rather than an accident: an abnormal level of arousal, or a documented acute change in mental function, is enough to warrant assessment whatever the cognitive items show
- the acute-change item needs an informant
- evidence of significant change or fluctuation in mental function arising over the last two weeks and still evident in the last 24 hours. It cannot be scored by looking at the patient, and the sources of the information are staff, carers, family, the notes or the clinician’s own prior knowledge. Left at 0 because nobody asked, it costs 4 points
- designed for repeated use
- Delirium fluctuates, which is why the published instruments are designed for repeated use: a single negative does not exclude it, and the acute-and-fluctuating criterion cannot be scored from the bedside alone without an informant history or the notes.
- untestable scores, rather than skips
- a patient too drowsy or too dysphasic to answer the AMT4 scores 2 on it, and one untestable on months backwards scores 2 there. A cognitive test would record “not done”; a delirium screen scores the inability, because the inability is the finding
- pooled accuracy
- sensitivity 0.88 (95% CI 0.80–0.93) and specificity 0.88 (0.82–0.92), from 17 studies and 3,702 observations across 11 countries — acute medicine, surgery, the emergency department, geriatric and stroke wards, post-operative units, a care home and nursing homes. Excluding the three stroke studies: sensitivity 0.86 (0.77–0.92), specificity 0.89 (0.83–0.93) in 14 studies and 3,440 observations
- what a positive screen is worth
- the table below applies the published figures at four prevalences. At the validation cohort’s own 12.3 per cent the positive predictive value is 44 per cent; at 2 per cent it is 10 per cent, with the negative predictive value above 99 per cent either way. The instrument does not change between those rows. This is a rule-out tool before it is a rule-in one, and the 2 × 2 arithmetic is what shows it
- inter-rater reliability
- NOT MEASURED, and the validation paper says so in its own limitations: inter-rater reliability and kappa were not assessed. The meta-analysis of 17 studies reports no kappa for the instrument either. What is reported is internal consistency, Cronbach’s alpha 0.80, which is a different property — it says the four items hang together, not that two raters agree. Two of the four items involve judgement and one depends on an informant, so this is a real gap rather than a formality
- licence and attribution
- CC BY 4.0, free to use with no permission, payment or registration required, and free to incorporate into electronic record systems. The full required attribution is printed at the foot of the article below, in the wording the instrument’s own attribution page specifies
Worked example
Alertness normal (0), one mistake on the AMT4 (1), starts months backwards but manages fewer than seven (1), and the nursing staff describe a change over the last three days that is still fluctuating today (4)
0 + 1 + 1 + 4 = 6 points
6 is in the 4-or-above band: possible delirium, with or without cognitive impairment, which the instrument says should prompt clinical assessment
Note where four of the six points came from. Without the informant history the total would be 2 — possible cognitive impairment, a different band and a different pathway. The acute-change item is the one that cannot be scored by looking at the patient, and it is worth 4
Now the arithmetic that matters. At the validation cohort's prevalence of 12.3 per cent, with 89.7 per cent sensitivity and 84.1 per cent specificity: of 10,000 patients, 1,230 have delirium and 1,103 of them screen positive, while 8,770 do not and 1,394 of them screen positive. Positive predictive value 1,103 / 2,497 = 44 per cent
The same two figures where delirium affects 2 per cent: a positive predictive value of 10 per cent, with the negative predictive value above 99 per cent
So a positive 4AT identifies a patient who needs assessing, and in most populations most positives will not have delirium. A negative is the stronger result — and even then, Delirium fluctuates, which is why the published instruments are designed for repeated use: a single negative does not exclude it, and the acute-and-fluctuating criterion cannot be scored from the bedside alone without an informant history or the notes.
The four items and their available points
| Item | Points | Needs |
|---|---|---|
| Alertness | 0 or 4 | Observation before stimulation; brief drowsiness under 10 seconds on waking is normal |
| AMT4 — age, date of birth, place, year | 0, 1 or 2 | Four questions; untestable scores 2 |
| Attention — months of the year backwards | 0, 1 or 2 | Seven or more correct back to June scores 0; one prompt allowed; untestable scores 2 |
| Acute change or fluctuating course | 0 or 4 | An informant, the notes, or prior knowledge of the patient — not obtainable from the bedside alone |
Accuracy, and what it is worth at four prevalences
| Figures used | Prevalence | Positive predictive value | Negative predictive value |
|---|---|---|---|
| Sensitivity 89.7%, specificity 84.1% (234-patient validation cohort) | 12.3% (that cohort’s own) | 44% | 98% |
| The same two figures | 2% | 10% | 99.8% |
| Sensitivity 88%, specificity 88% (pooled, 17 studies, 3,702 observations) | 20% | 65% | 97% |
| The same two figures | 2% | 13% | 99.7% |
Four items, two of them worth four points, and one that needs an informant
The 4AT was built to be usable: four items, no training requirement, under two minutes, and scoreable in a patient who cannot cooperate. Alertness and acute change are worth four points each with nothing in between, so either reaches the threshold of four on its own; the AMT4 and months-backwards items are worth up to two each, so together they cannot. A total of one, two or three therefore carries information beyond its size — it means normal arousal and no documented acute change.
The item that decides most scores is the one that cannot be scored by looking at the patient. Evidence of significant change or fluctuation in mental function over the last two weeks, still evident in the last twenty-four hours, comes from staff, family, carers, the notes or prior knowledge. In the worked example above it supplies four of six points, and without it the same patient lands in a different band. Delirium fluctuates, which is why the published instruments are designed for repeated use: a single negative does not exclude it, and the acute-and-fluctuating criterion cannot be scored from the bedside alone without an informant history or the notes.
The accuracy is good and it is not the whole story. Pooled across seventeen studies and 3,702 observations in eleven countries, sensitivity and specificity are both 0.88. In the original two-centre validation of 234 hospitalised older people, mean age 83.9, with delirium in 12.3 per cent and dementia in 31.2 per cent, the threshold of four was 89.7 per cent sensitive and 84.1 per cent specific with an area under the curve of 0.93. The table above applies those figures at four prevalences: the positive predictive value runs from 65 per cent down to 10 per cent while the negative predictive value stays above 97 per cent. This is a rule-out instrument first, and a positive result is a reason to assess rather than a finding.
Two honest gaps. Inter-rater reliability has not been measured: the validation paper lists it among its own limitations and the meta-analysis reports no kappa either, and what is published is internal consistency, Cronbach’s alpha 0.80, which says the items hang together rather than that two raters agree. And that paper is internally inconsistent in two places, giving the cohort as 236 in the abstract and 234 in the results, and delirium-with-dementia as 7.2 per cent against 23 per cent; the figures used here are the ones that match the patient counts. For comparison, the Confusion Assessment Method has sensitivities from 0.09 to 1.0 across 23 accuracy studies in 2,629 patients, with one large implementation study at 0.28 where it was scored without the preceding interview its manual requires. It is not reproduced here, because its reuse terms could not be established as free.
4AT © 2011–2014 Alasdair MacLullich, Tracy Ryan and Helen Cash. Licensed under CC BY 4.0 (creativecommons.org/licenses/by/4.0/). Official and current version: www.the4at.com. The items, response options and scoring thresholds here are those of the official 4AT and have not been changed; the on-screen wording is abbreviated for a web form, and the full official wording and administration instructions are on the4at.com. No warranty is given as to accuracy or fitness for purpose.
A screening score is not a diagnosis: a published sensitivity is a property of the instrument in the population it was validated in, not a statement about this patient. This page reports published figures and recommends no action. Every weight, cut-off and outcome figure here comes from a named derivation cohort, and cohorts differ in case mix, era, coding and outcome definition; where your own institution’s protocol or analysis plan differs, it takes precedence.
Frequently asked questions
What does a 4AT score of 4 or more mean?
The instrument’s own interpretation is possible delirium, with or without cognitive impairment, and that this should prompt clinical assessment. It is not a diagnosis: in the 234-patient validation cohort, with delirium in 12.3 per cent, the positive predictive value at that threshold was 44 per cent. A screening score is not a diagnosis: a published sensitivity is a property of the instrument in the population it was validated in, not a statement about this patient.
Does a 4AT of 0 exclude delirium?
No. The instrument’s own wording is that delirium or severe cognitive impairment is unlikely but that delirium remains possible if the information is incomplete — and the incomplete information is usually the acute-change item, which needs an informant. Delirium fluctuates, which is why the published instruments are designed for repeated use: a single negative does not exclude it, and the acute-and-fluctuating criterion cannot be scored from the bedside alone without an informant history or the notes.
Why are two items worth 4 points?
So that either can reach the threshold alone. A clearly abnormal level of arousal, or a documented acute change or fluctuation in mental function, is enough to warrant assessment whatever the two cognitive items show. The corollary is that the AMT4 and attention items together cannot reach 4, so a total of 1 to 3 always means normal arousal and no documented acute change.
Is the 4AT free to use?
Yes, and explicitly so. Its own site states that the 4AT is completely free to use, that no permission, payment or registration is required, that it is made freely available under the CC BY 4.0 licence, and that it is free to incorporate into electronic record systems without licensing fees or specific permission. Attribution is required and is printed on this page. That is unusual: most delirium and cognitive instruments are not reproducible on these terms, which is why several are named on this site and none of them is built.
How reliable is the 4AT between different raters?
Unknown. The validation study did not assess inter-rater reliability and lists that as a limitation; the seventeen-study meta-analysis reports no kappa for the instrument. The published reliability figure is internal consistency, Cronbach’s alpha 0.80, which is a different thing. Two of the four items involve judgement and one depends on an informant, so the absence of a kappa is a real gap and not a formality.
Related calculators
References
- the4AT. 4AT: frequently asked questions. the4at.com/4at-faq (accessed 9 October 2026).
- the4AT. Attribution and reuse. the4at.com/attribution (accessed 9 October 2026).
- Bellelli G, Morandi A, Davis DHJ, Mazzola P, Turco R, Gentile S, Ryan T, Cash H, Guerini F, Torpilliesi T, Del Santo F, Trabucchi M, Annoni G, MacLullich AMJ. Validation of the 4AT, a new instrument for rapid delirium screening: a study in 234 hospitalised older people. Age Ageing. 2014;43(4):496–502.
- Tieges Z, Maclullich AMJ, Anand A, Brookes C, Cassarino M, O’Connor M, Ryan D, Saller T, Arora RC, Chang Y, Agarwal K, Taffet G, Quinn T, Shenkin SD, Galvin R. Diagnostic accuracy of the 4AT for delirium detection in older adults: systematic review and meta-analysis. Age Ageing. 2021;50(3):733–43.
- Regenstrief Institute. LOINC term 52495-9, Confusion Assessment Method, copyright and terms-of-use block. loinc.org/52495-9 (accessed 9 October 2026).
Not medical advice. For healthcare professionals and education. Reference intervals vary by laboratory and assay — always use your own laboratory's. Never base a dose or a treatment decision on this page alone. Full disclaimer at calcengines.com/disclaimer/
