Hydronephrosis Grade Interpreter

Hydronephrosis Grade Interpreter

An SFU grade and a UTD category read side by side on the same kidney, with the anteroposterior renal pelvic diameter thresholds for the gestational or postnatal timing entered. The two systems do not map onto one another, and this page is built to show where they part company.

SFU grade and UTD category, side by side

Grade + dilatation → both systems, and their disagreements
The UTD thresholds differ for each. A postnatal scan in the first 48 hours is excluded by the consensus itself, because transient dehydration understates the dilatation — so an early reassuring scan is not a reassuring scan.
Measured on a transverse image at the maximum intrarenal diameter. Only 64 per cent of physicians agreed on how to measure it in the surveys cited by the grading review read for this page, which is worth knowing before a 9 mm is treated as different from an 11 mm.
These descriptors are the structure the two text sources read for this page agree on, attributed to Fernbach, Maizels and Conway’s 1993 introduction of the system. THE FULL CRITERION WORDING IS NOT REPRODUCED HERE: the review that carries it in full carries it as an image, and a grade descriptor invented from a figure caption would be worse than none. Note also that one reliability study had to split the system into two readings it called SFU-A and SFU-B, which agreed with the UTD classification at kappa 0.50 and 0.75 respectively.
In the UTD classification a parenchymal abnormality places the kidney in the highest risk group whatever the diameter is. It is also what defines SFU grade 4, which is why entering grade 4 with a normal parenchyma raises an inconsistency on this page rather than an answer.
Central and peripheral dilatation are assessed separately in the UTD system and carry different weight: central dilatation is compatible with the low-risk group, peripheral dilatation is not.
The consensus treats transient visualisation of the ureter after birth as normal; a persistently dilated ureter is not, and escalates the category on its own.
UTD P1 postnatally, or UTD A1 antenatally, with a mild SFU gradeExample

A postnatal ultrasound at three weeks of age: anteroposterior renal pelvic diameter 12 mm, SFU grade 2, normal parenchyma, central calyceal dilatation, ureter not dilated

Advertisement

The UTD thresholds, and what the systems measure

Antenatal, 16–27 weeks — normal under 4 mm · A1 4 to under 7 mm · A2-3 7 mm and above
Antenatal, 28 weeks and over — normal under 7 mm · A1 7 to under 10 mm · A2-3 10 mm and above
Postnatal, after 48 hours — normal under 10 mm · P1 10 to under 15 mm · P2 15 mm and above · P3 any parenchymal abnormality or abnormal bladder
Peripheral calyceal dilatation, or a dilated ureter, escalates the category irrespective of the diameter
the systems do not map onto each other
The grading systems do not map onto one another and were derived in infants. Record which system a grade came from whenever you write one down. The SFU grade describes the anatomy of one kidney’s collecting system; the UTD classification assigns a risk group across the whole urinary tract including the ureter and the bladder; the anteroposterior diameter is a single measurement. The review read for this page states that the diameter and the SFU grades are not parallel for many patients because of the configuration of the renal pelvis, and that no threshold diameter separates obstructive from non-obstructive dilatation
7 mm means two different things
it is the A2-3 floor before 28 weeks and the A1 floor from 28 weeks. The same measurement on the same fetus two weeks apart can move down a category without anything changing, which is why the timing is asked first
derived in infants
both systems were developed for antenatally detected and neonatal dilatation. The SFU system was, in its authors’ own framing, designed to grade neonatal and infant pelvicalyectasis. Nothing here is a grading system for the acute hydronephrosis of an adult obstructing ureteric stone, which is conventionally described rather than graded
what this page deliberately does not reproduce
the full SFU criterion wording. The review that carries all four grades and the subdivisions in full carries them in figures, and the fetched text held captions only; two text sources were used instead and the page reproduces only the structure they agree on. It also does not ask about the bladder or about oligohydramnios, both of which place a case in a higher UTD category on their own
the reliability
see the second table. Short version: in 180 infant kidneys read twice by four radiologists, intra-observer agreement was 0.64 to 0.88 for the SFU system and 0.48 to 0.92 for the UTD system, with SFU grades 2 and 3 only fair to moderate between observers; and in a multicentre study with eight reviewers from five institutions reading 30 renal units, Light’s kappa for the UTD grade was 0.43. This page reports a figure and what the published sources attach to it. It does not make a clinical decision and cannot.

Worked example

A postnatal ultrasound at three weeks of age: anteroposterior renal pelvic diameter 12 mm, SFU grade 2, normal parenchyma, central calyceal dilatation, ureter not dilated
Parenchyma normal and SFU grade is not 4, so neither the inconsistency rule nor the high-risk rule applies
Calyceal dilatation is central rather than peripheral and the ureter is not dilated, and 12 mm is below the postnatal P2 floor of 15 mm — so the intermediate-risk rules do not fire
12 mm is at or above the postnatal P1 floor of 10 mm → UTD P1, the low-risk group, and an SFU grade of 2 sits in the same part of its range, so the two systems agree here
Change one entry: make the calyceal dilatation peripheral and the UTD category becomes P2 at the same 12 mm, while the SFU grade of 2 does not move — the page would then report a disagreement rather than an agreement
Change the timing instead: the same 12 mm at 20 weeks' gestation is UTD A2-3, because the A2-3 floor before 28 weeks is 7 mm. One measurement, three categories, depending on when and what else was seen
Advertisement

The UTD thresholds by timing, from the 2014 consensus

TimingNormalLow risk (A1 / P1)Increased risk (A2-3 / P2)
Antenatal, 16 to 27 weeksUnder 4 mm4 to under 7 mm7 mm and above
Antenatal, 28 weeks and overUnder 7 mm7 to under 10 mm10 mm and above
Postnatal, more than 48 hours after birthUnder 10 mm10 to under 15 mm15 mm and above
Overriding findings, any timingNo calyceal dilatation, normal parenchyma, ureter not seen, normal bladderCentral calyceal dilatation is permitted herePeripheral calyceal dilatation, a dilated ureter, an abnormal parenchyma or an abnormal bladder — each on its own, whatever the diameter
Read as text out of the consensus paper itself, not off a figure. Note that a parenchymal abnormality — thinning, raised echogenicity, loss of corticomedullary differentiation — or an abnormal bladder puts a case in P3 irrespective of everything else, and that this page does not ask about the bladder. Note also that the first 48 hours after birth are excluded: transient dehydration understates the diameter, so an early scan cannot be reassuring.

How reproducible these grades actually are

FigureValueCohort
SFU intra-observer agreement0.64 to 0.88180 infant kidneys, four radiologists each reading twice
UTD intra-observer agreement0.48 to 0.92The same 180 kidneys and raters
SFU grades 2 and 3, between observersFair to moderate only, and overall inter-observer agreement was significantly higher for the UTD system than for the SFU systemThe same cohort
UTD against SFU, agreement between the two systemsKappa 0.75 against one reading of the SFU system and 0.50 against the other (the authors had to define two, SFU-A and SFU-B)The same cohort
UTD grade, Light’s kappa0.43, described as weak to moderate; the commonest disagreement was between P2 and P330 renal units from 15 infant sonograms, eight reviewers from five institutions
Anteroposterior diameter as a measurementExcellent reliability as a continuous measure, but only 64 per cent of physicians agreed on how to measure itThe multicentre study above, and the surveys cited by the 2020 grading review
Parenchymal thickness as a binary, above or below 3 mm70 per cent agreementThe same multicentre study
These are the numbers that decide how much weight a single recorded grade can carry. The measurement is more reliable than the grade, and the grade that matters most — the distinction between the two intermediate categories — is the least reliable part of either system. A grade written in a record without the system it came from is close to uninterpretable, which is why the two are reported together here rather than converted.

Two systems, one kidney, and no conversion between them

Hydronephrosis has more than one grading system in active use and they measure different things. The Society for Fetal Urology grades, introduced in 1993, describe the anatomy of one kidney’s collecting system from 0 to 4: pelvis only, pelvis plus major calyces within the renal border, pelvis beyond the border with major and minor calyces dilated and the parenchyma preserved, and the same with the parenchyma thinned. The anteroposterior renal pelvic diameter is a single measurement in millimetres. The urinary tract dilation classification, published as a multidisciplinary consensus in 2014, uses that diameter together with the calyces, the parenchyma, the ureter and the bladder to assign a risk group across the whole tract. The grading systems do not map onto one another and were derived in infants. Record which system a grade came from whenever you write one down.

They do not convert. The 2020 review read for this page states that the diameter and the SFU grades are not parallel for many patients because of the configuration of the renal pelvis — a capacious extrarenal pelvis gives a large diameter with little calyceal dilatation, a small intrarenal pelvis the reverse — and that no threshold diameter separates obstructive from non-obstructive dilatation. The study that tried to convert between them had to define two readings of the SFU system, reaching kappa 0.75 against one and 0.50 against the other in 180 infant kidneys. This page therefore reports both readings and names the disagreement when there is one, rather than returning a single grade.

The reliability figures are the other reason not to over-read one recorded grade. In those 180 kidneys, read twice by four radiologists, intra-observer agreement was 0.64 to 0.88 for the SFU system and 0.48 to 0.92 for the UTD system — a range wide enough that some observers barely agreed with themselves — and SFU grades 2 and 3 showed only fair to moderate agreement between observers. A multicentre study with eight reviewers from five institutions reading 30 renal units found Light’s kappa for the UTD grade to be 0.43, with the commonest disagreement between P2 and P3. The continuous diameter was the most reliable thing measured, and even there only 64 per cent of physicians agreed on how to measure it.

Three practical points follow. Write the system beside the grade, always — ‘2’ without ‘SFU’ or ‘UTD’ beside it is close to uninterpretable. Do not grade a postnatal scan taken inside the first 48 hours: the consensus excludes it because transient dehydration understates the diameter, so an early reassuring scan is not reassuring. And note what this page cannot see: an abnormal bladder, a ureterocele, a dilated posterior urethra or oligohydramnios each place a case in the highest UTD category on their own and are not among the inputs. The downstream kidney function this is all a proxy for is estimated with the CKD-EPI 2021 eGFR calculator, staged acutely with the KDIGO AKI stage calculator, and the prerenal question is addressed by the fractional excretion of sodium calculator.

Frequently asked questions

Can I convert an SFU grade into a UTD category?

Not reliably, and the attempt is instructive. The study that measured the correspondence in 180 infant kidneys had to define two different readings of the SFU system, and reached kappa 0.75 against one and 0.50 against the other. The systems also measure different things: an SFU grade describes one kidney’s collecting system, a UTD category assigns risk across the whole urinary tract.

Which system should I use?

This page does not choose, and the literature has not. In the comparison study, overall inter-observer agreement was significantly higher for the UTD system than for the SFU system; the 2020 review read for this page calls UTD “not an evidence-based grading system” and concludes that neither the diameter nor the radiology, SFU or UTD systems is a gold standard for severity. That conclusion is one author’s position, not a consensus finding, and is reported as such. Use whichever your unit uses, and record which.

Why does the page refuse when I enter SFU grade 4 with a normal parenchyma?

Because grade 4 is defined by parenchymal thinning — it is grade 3 plus a thinned parenchyma. The two entries contradict each other, so one of them is wrong, and answering either reading would be inventing information. The page says which entries are in conflict instead.

Why does 7 mm mean different things at different gestations?

Because the consensus sets different thresholds before and after 28 weeks. Before 28 weeks, 7 mm is the floor of the increased-risk group A2-3; from 28 weeks, 7 mm is the floor of the low-risk group A1. The same measurement two weeks later is a lower category with nothing having changed, which is why the timing is the first field.

Do these systems apply to an adult with an obstructing stone?

No. Both were developed for antenatally detected and neonatal dilatation — the SFU system was designed to grade neonatal and infant pelvicalyectasis — and the published reliability data come from infant cohorts. Acute hydronephrosis in an adult with a ureteric stone is conventionally described rather than graded on these scales; the stone itself is covered by the stone volume page and the spontaneous passage page.

What this page will not tell me

Every threshold, range and performance figure on this page is the published figure from the source named beside it, and each one depends on the population, the method and the equipment it was derived in. Where your own report, laboratory or local guideline gives a different figure, that figure governs. A score, an index, a measured volume or an attenuation value is not a diagnosis, and a proportion measured in a cohort is not a probability for one patient. This page reports a figure and what the published sources attach to it. It does not make a clinical decision and cannot.

Related calculators

References

  1. Nguyen HT, et al. Multidisciplinary consensus on the classification of prenatal and postnatal urinary tract dilation (UTD classification system). J Pediatr Urol. 2014;10(6):982–998.
  2. Fernbach SK, Maizels M, Conway JJ. Ultrasound grading of hydronephrosis: introduction to the system used by the Society for Fetal Urology. Pediatr Radiol. 1993;23(6):478–480 — cited as the origin of the SFU grades in every secondary source read; the grade criteria themselves are reproduced here only as far as two independent text sources agree.
  3. Han M, Kim HG, Lee JD, Park SY, Sur YK. Conversion and reliability of two urological grading systems in infants: the Society for Fetal Urology and the urinary tract dilatation classification system. Pediatr Radiol. 2017;47(1):65–73 — 180 kidneys, four radiologists, each reading twice.
  4. Interrater reliability of objective and subjective measures of urinary tract dilation by pediatric urologists: a multicenter analysis. J Pediatr Urol. 2025;21(6):1764–1770 — eight reviewers from five institutions, 30 renal units.
  5. Onen A. Grading of hydronephrosis: an ongoing challenge. Front Pediatr. 2020;8:458 — read for the discussion of why the systems do not map onto one another; its full grade criteria are held in figures and are not reproduced here.

Not medical advice. For healthcare professionals and education. Reference intervals vary by laboratory and assay — always use your own laboratory's. Never base a dose or a treatment decision on this page alone. Full disclaimer at calcengines.com/disclaimer/