CFD Solver Memory & Runtime Calculator

CFD Solver Memory, Runtime and Storage Calculator

How much RAM a case needs, whether it fits the node you have, how many cells per core that leaves you, what Amdahl’s law says about the speed-up and why it is the wrong question, how long the run takes and how much disk the transient output will eat. Every per-cell memory figure is attributed to published vendor guidance and given as a range.

Memory, cores, runtime and storage

cells, solver, precision and cores → RAM, cells per core, wall clock and disk
Fluid cells, in millions. Solid zones in a conjugate problem count too but cost less, because they carry only the energy equation. Boundary faces add roughly what their own count suggests and are ignored here, which matters only on a mesh that is mostly surface.
Memory per cell tracks the number of faces per cell, because the matrix and the face fluxes are stored per face. A tetrahedron has four faces, a hexahedron six, and a polyhedron from a dualised tet mesh commonly twelve to fifteen. The factors used here come from published vendor guidance and are attributed in the references.
A segregated solver holds one scalar matrix at a time and reuses it, so its matrix storage is small. A coupled solver assembles a block matrix whose every entry is a 4×4 or 5×5 sub-block, and the algebraic multigrid hierarchy on top of it is proportionally larger. That is the whole reason coupled runs need roughly twice the memory.
Double precision costs about 60% more memory, not 100%, because the mesh connectivity and the index arrays are integers and do not change size. Use it for large ranges of cell size or pressure, high aspect ratios, multiphase, and anything where the pressure is a small difference between large numbers.
Each additional Eulerian phase brings its own volume fraction and momentum set while sharing the pressure field and the mesh, so this page adds 80% of the single-phase field and matrix storage per extra phase. That is a stated modelling assumption, not a vendor figure. A Lagrangian or discrete-phase model is not this — its cost goes with the number of parcels, not the number of cells, and is not modelled here.
Species, user scalars, extra turbulence equations beyond two, a transition model’s two, a Reynolds-stress model’s extra four. The base figure already includes a two-equation turbulence model and energy. Each extra scalar is charged 5% of the base here, which is a derived floor rather than a measurement: one field, one old-time level, three gradient components and a face flux is about 48 bytes against a base of roughly a kilobyte per cell. Measured increments run higher, commonly 5 to 15% each, and far higher with reacting flow where every species also carries reaction rates and transport properties.
If you have measured your own solver on your own model set, put it here and the model above is bypassed entirely. This is the input to prefer: per-cell memory is vendor-specific and model-specific, and one measurement from your own case beats any published factor.
Total cores across all nodes. Used for memory per core, cells per core, and the parallel speed-up.
Physical cores in one compute node. Together with the memory per node this decides whether the case fits, because memory is allocated per node and a node that runs out does not care that another node has spare.
Usable RAM per node, not the sticker figure. Leave 5 to 10% for the operating system, the file cache and the MPI buffers; a run that fits in 99% of the node’s memory does not fit.
For a steady run, outer iterations. For a transient run, time steps multiplied by the inner iterations each one takes — a 10,000-step run at 5 inner iterations is 50,000 here. That multiplication is where transient runtime estimates most often go wrong by a factor of five.
Microseconds of single-core time to advance one cell by one iteration. Roughly 1 to 3 µs for a segregated incompressible solver, 3 to 10 µs coupled, more again with radiation, combustion chemistry or a Reynolds-stress model. Measure it: run 50 iterations on the real mesh, take the wall clock, multiply by the cores, divide by cells times iterations.
The fraction of the work that does not parallelise: reading the case, partitioning, writing files, global reductions. Measured values for production CFD sit between 10⁻⁴ and 10⁻², and the maximum possible speed-up is its reciprocal — 0.001 caps you at 1,000× however many cores you buy. Amdahl is a strong-scaling model and it is the wrong one for the way CFD is usually run; see the note under the result.
How many times the full field is written to disk. Zero for a steady run that only saves at the end. The commonest cause of a filled filesystem is a write frequency set in iterations rather than in time steps.
Scalars saved per cell per write. Pressure, three velocity components, two turbulence quantities and temperature is eight; add species, derived quantities, wall distance and any user fields. Writing only what the post-processing needs is the single cheapest storage saving available.
Post-processing almost never needs double-precision fields; writing them doubles the storage for no benefit a contour plot can show. Restart files are a different matter and should stay in the solver’s own precision.
106.56GBExample

20 million polyhedral cells, pressure-based coupled solver, double precision, single phase, no extra equations, 256 cores at 128 per node with 512 GB per node, 5,000 iterations at 3 µs per cell per iteration, a serial fraction of 0.001, and 200 transient writes of 10 single-precision fields

Advertisement

Memory, cells per core, Amdahl’s law and the write volume

RAM = N × m × fmesh fsolver fprecision fphases fequations  ·  cells/core = N/P  ·  S(P) = 1/(s + (1 − s)/P)  ·  Smax = 1/s  ·  η = S(P)/P  ·  T(P) = T1 (s + (1 − s)/P)  ·  storage = N nfields b nwrites
m
base memory per million cells: 1.0 GB for tetrahedra, single precision, segregated, with a two-equation turbulence model and energy. Ansys’s published Fluent guidance; attributed rather than invented, and used as the anchor for every factor that follows
f_mesh
1.0 tetrahedral, 1.20 hexahedral, 1.80 polyhedral. It tracks faces per cell, because the matrix coefficients and the face fluxes are stored per face
f_solver
1.0 segregated, 1.85 pressure-based coupled. The density-based figure of 2.30 is derived from the coupled one by the ratio of block sizes, 25/16 against 16/16, rather than taken from a source
f_precision
1.60 for double precision. Not 2.0, because the connectivity and index arrays are integers and do not change size
s
serial fraction. The maximum speed-up is its reciprocal whatever the core count, and efficiency reaches one half at P = (1 − s)/s
T_1
single-core time, cells × iterations × cost per cell per iteration. This is the quantity to measure rather than estimate: it is the only linear factor in the runtime that nothing else on this page constrains
b
bytes per written value, 4 or 8. Post-processing fields rarely need 8

Worked example

20 million polyhedral cells, pressure-based coupled solver, double precision, single phase, no extra equations, 256 cores at 128 per node with 512 GB per node, 5,000 iterations at 3 µs per cell per iteration, a serial fraction of 0.001, and 200 transient writes of 10 single-precision fields
Memory first. The base figure is 1.0 GB per million cells for tetrahedra, single precision, segregated, with a two-equation turbulence model and energy. Polyhedral multiplies it by 1.80, the pressure-based coupled solver by 1.85 and double precision by 1.60: 1.0 × 1.80 × 1.85 × 1.60 = 5.328 GB per million cells. Twenty million cells is 106.6 GB as a central estimate, and the range carried through the page is 74.6 to 159.8 GB
SANITY CHECK AGAINST THE SOURCE, because a factor model is only worth what its ingredients are. The vendor's published table gives 5.6 GB per million cells for polyhedral, double precision, coupled. The model says 5.328, which is 4.9% low — the largest disagreement anywhere in the twelve entries, and the reason the range is quoted rather than the point value
DOES IT FIT. 256 cores at 128 per node is 2 nodes, so 53.3 GB per node on the central estimate against 512 GB available: 10.4% of the node memory, and 15.6% at the high end. Comfortable. Note that the test is per node and not in aggregate — a job needing 800 GB does not fit on two 512 GB nodes, because partitioning is never balanced and the system wants its share
CELLS PER CORE: 20 million over 256 cores is 78,125. Well above the 10,000 Ansys call excellent and a long way from the 2,500 below which communication takes over; also well below the 200,000 at which CFD Direct still held 90% scaling. This is a sensible working point and there is room to add cores if wall clock matters more than core-hours
PARALLEL SPEED-UP, from Amdahl's law with s = 0.001: S = 1/(0.001 + 0.999/256) = 204× on 256 cores, an efficiency of 79.7%. The ceiling this serial fraction imposes at any core count is 1/s = 1,000×, and efficiency would be down to one half at (1 − s)/s = 999 cores
WALL CLOCK. Single-core time is 20×10⁶ cells × 5,000 iterations × 3 µs = 3.0×10⁵ core-seconds, which is 83.3 core-hours. Spread over 256 cores with perfect scaling that would be 0.33 hours; with the serial fraction it is 0.41 hours, about 25 minutes, and the run consumes 104.5 core-hours rather than the 83.3 the work needs. The difference, 21 core-hours, is what the serial fraction costs
AND BE SUSPICIOUS OF THAT NUMBER, because Amdahl's law is the wrong model here. It contains no term for halo exchange, which is what actually limits CFD scaling, and it describes strong scaling when CFD is usually run in weak scaling. The cells-per-core figure is the better guide and the page warns on it separately
TRANSIENT OUTPUT: 20×10⁶ cells × 10 fields × 4 bytes = 0.8 GB per write, times 200 writes is 160 GB, at an average of 392 GB per hour of run time. Doubling the fields or switching the output to double precision doubles that; setting the write frequency in iterations when it was meant to be in time steps is how it becomes 8 TB

Where the memory factors come from

ChangeFactorSource
Baseline: tetrahedra, single precision, segregated, two-equation turbulence and energy1.0 GB per million cellsAnsys published Fluent guidance
Hexahedral instead of tetrahedral×1.20fits the same vendor table to 2%
Polyhedral instead of tetrahedral×1.80same vendor table; note the vendor’s own GPU memory page says 20 to 40% over hexahedral, while its sizing table implies 47 to 56%
Pressure-based coupled instead of segregated×1.85same vendor table, consistent across all three mesh types
Density-based implicit instead of segregated×2.30derived here from the 4×4 block figure by the block-size ratio 25/16, not sourced
Double precision instead of single×1.60same vendor table; the vendor’s GPU memory page says 50%, its sizing table implies 56 to 60%
Each extra Eulerian phase×1.80 for two phases, ×2.60 for threethis page’s assumption, stated as such
Each extra transported scalar equation×1.05derived floor from field, old level, gradient and flux storage; measured values run higher
Every multiplicative factor here was checked against the vendor table it came from before use: a three-factor model of mesh type, solver and precision reproduces all twelve published entries to within 4.9%, and the cross-ratios are consistent across all three mesh types, which is what makes it safe to use them independently. Two figures disagree between the vendor’s own two documents, and both disagreements are in the same direction — the sizing table implies a larger penalty than the memory page states. That is precisely why this page reports a range and not a byte count.

Cells per core: the number practitioners argue about

Cells per coreWhat happensWho says so
Below about 2,500the halo exchange costs more than the arithmetic it supports; adding cores makes the wall clock worseAnsys, from Fluent scaling on thousands of cores at Argonne
2,500 to 10,000good scaling, and the usual place to run when wall clock matters more than core-hoursAnsys, same source
Above 10,000excellent efficiency; the network is not the constraintAnsys, same source
100,000a reasonable weak-scaling working point, though departures from linear scaling appear by a few hundred coresCFD Direct, OpenFOAM on AWS
200,00090% scaling held at 504 cores on a 97-million-cell external aerodynamics caseCFD Direct, same source
There is no single right answer and anyone who gives you one is quoting their own machine. The trade-off is real in both directions: too few cells per core and communication dominates, too many and you are waiting for arithmetic that could have been spread out, with memory per rank as the hard limit at the top end. The figures differ between solvers because the communication pattern differs — a segregated solver exchanges one halo per equation per iteration while a coupled one exchanges a block — and between machines because the interconnect differs. Measure two core counts on your own case and interpolate; it takes twenty minutes and it beats every table including this one.

An attributed range beats an invented byte count, and cells per core beats Amdahl’s law

Per-cell memory is vendor-specific, model-specific and mesh-specific, and a page that quotes a precise byte count it cannot source is worse than one that says 1 to 2 GB per million cells and explains what moves it. So this page starts from a figure that is actually published — Ansys’s guidance of about 1 GB per million cells for tetrahedra, single precision, a segregated solver, a two-equation turbulence model and energy — and applies factors taken from the same vendor’s sizing table for the things that change it. Every factor is attributed in the table above, the two that are not sourced are labelled as derived, and every number comes with a range.

Why a range rather than a number, in the vendor’s own words. Ansys publishes two documents that bear on this and they do not quite agree. The memory page for the GPU solver says a polyhedral mesh needs 20 to 40% more memory than a hexahedral one and that double precision needs 50% more. The hardware sizing table in the buying guidance implies 47 to 56% for polyhedral over hexahedral, and 56 to 60% for double over single. Both disagreements point the same way: the table asks for more than the prose. Add to that the difference between CPU and GPU solvers, between one vendor and another, and between model sets, and a spread of −30% to +50% around a central estimate is honest where a single figure would not be. If you have measured your own solver, the page takes your number instead and bypasses the model entirely — that is the input to prefer over anything published.

What drives per-cell memory, and it is not the number of equations. It is faces per cell and the size of the matrix block. Matrix coefficients and face fluxes are stored per face, so a polyhedron with fourteen faces costs more than a tetrahedron with four, and that is where the mesh factor comes from. The solver factor is larger and comes from the same place: a segregated solver holds one scalar matrix at a time and reuses it for each equation in turn, while a coupled solver assembles a block matrix in which every entry is a 4×4 sub-block for pressure and the three velocities — sixteen times the storage per coefficient — with an algebraic multigrid hierarchy on top of it that is proportionally larger. Extra transported scalars, by contrast, are cheap in a segregated solver: one field, one old-time level, three gradient components and a face flux is about 48 bytes against a base of roughly a kilobyte per cell, which is where the 5% per equation used here comes from. That figure is a floor, and reacting flow will exceed it substantially.

Memory fits node by node, not in aggregate. This is the arithmetic that most often catches people out. A job needing 800 GB does not fit on two 512 GB nodes just because 1,024 exceeds 800: partitioning is never perfectly balanced, and the operating system, the file cache and the MPI buffers all want a share. Eighty-five per cent of the node total is the practical ceiling, and it should be checked against the high end of the memory range rather than the central estimate. The partitioning step itself deserves a separate thought, because in several solvers it is done on a single rank and can peak higher than the solve it is preparing — a case that runs comfortably can still fail to start.

Amdahl’s law is on this page, and it is the wrong model for CFD. It is here because it is what people ask for and because the one thing it captures is real: a fixed serial fraction caps the speed-up at its reciprocal however many cores you buy, so 1% serial means never better than 100× and efficiency is already down to one half at 99 cores. But CFD does not fail to scale because of a serial fraction. It fails to scale because halo exchange grows relative to the arithmetic as the partitions get smaller, which is a term Amdahl’s law does not contain at all. That is why the calculator prints the cells-per-core figure next to the speed-up and warns on it independently: below roughly 2,500 cells per core the communication overtakes the work and adding cores makes the wall clock worse, while Amdahl’s law will cheerfully promise a speed-up the network cannot deliver. The other reason to be careful with it is that CFD is usually run in weak scaling — more cores because the mesh got bigger — and Amdahl’s law describes strong scaling, a fixed problem spread thinner. Measured the right way, from two runs at different core counts, the serial fraction for production CFD comes out between 10⁻⁴ and 10⁻².

The cells-per-core question has no single answer and the honest thing is to give the range. Ansys, from Fluent scaling runs at Argonne on thousands of cores, report good scaling down to about 2,500 cells per core and excellent efficiency above 10,000. CFD Direct, running OpenFOAM on AWS, held 90% scaling at 504 cores with about 200,000 cells per core on a 97-million-cell external aerodynamics case, and saw departures from linear scaling by about 100 cores in a weak-scaling test at 100,000 cells per core. Those figures are not in conflict; they are different solvers, different communication patterns and different interconnects. The trade-off is genuine in both directions — too few cells per core and communication dominates, too many and memory per rank becomes the hard limit and you are waiting for arithmetic that could have been spread out. Two runs at different core counts on your own case settles it in twenty minutes.

Transient output is where the filesystem gets filled, and the arithmetic is trivial enough that nobody does it. Cells times fields times bytes times writes. Twenty million cells, ten fields, single precision, two hundred writes is 160 GB, which is fine. The same case at five hundred writes in double precision is 800 GB, which is often not. Two further things worth checking before a long transient run: the write rate, because a write that takes longer than the interval between writes will throttle the solver rather than the disk; and whether a write frequency has been set in iterations when it was meant to be in time steps, which is the commonest way a run produces fifty times the data intended. Writing surfaces or a subset of fields instead of the whole volume is usually the cheapest saving available, and the run you are sizing here is a good place to decide what the post-processing actually needs. If the run in question is a scale-resolving one, the LES and DES affordability page sizes the cell count and the step count that feed into this one.

Advertisement

Frequently asked questions

How much RAM per million cells should I assume?

For a single-phase segregated run in single precision, 1 to 2 GB per million cells covers most solvers and mesh types. Double precision adds about 60%, a coupled solver about 85%, and a polyhedral mesh about 80% over tetrahedra — so a polyhedral, coupled, double-precision case lands between 4 and 6 GB per million cells. Those are the figures this page uses and they are attributed in the table above. What moves them most, in order: the solver, the mesh type, the precision, and then the model set. What almost nobody guesses right is the direction of the mesh-type effect, which is about faces per cell rather than about cells.

Why is double precision only 60% more rather than twice?

Because a large share of what a solver stores is not floating-point at all. Mesh connectivity, face-to-cell maps, owner and neighbour arrays, partition boundaries and the multigrid coarse-level topology are integers and do not change size when the floating-point precision does. The vendor’s own two documents bracket the effect at 50% and 56 to 60%, and this page uses 1.60.

Does my case fit on the node?

Check the fraction-of-node-memory figure against the high end of the range, not the central estimate, and treat 85% as the ceiling rather than 100%. Three reasons: partitioning is never perfectly balanced, so the worst rank holds more than the average; the operating system, the file cache and the MPI buffers want a share; and the partitioning step itself is done on a single rank in several solvers and can peak higher than the solve. A case that fits at 95% on paper will sometimes fail to start and sometimes fail after four hours, which is worse.

What is the right number of cells per core?

There is no right number, which is the point. Ansys report good scaling down to about 2,500 cells per core and excellent efficiency above 10,000 from Fluent runs on thousands of cores. CFD Direct held 90% scaling at 504 cores with 200,000 cells per core on an OpenFOAM external aerodynamics case. Both are correct for their solver and machine. The trade-off in words: too few cells per core and the halo exchange costs more than the arithmetic it supports; too many and memory per rank becomes the constraint and you are waiting for work that could have been spread out. If wall clock is what you are buying, run nearer the bottom of the range; if core-hours are what you are paying for, run nearer the top.

Is Amdahl’s law a good model for CFD scaling?

No, and the page says so beside the number. Amdahl’s law captures one real effect — a fixed serial fraction caps the speed-up at its reciprocal, so 1% serial means never better than 100× — and misses the one that actually limits CFD, which is that halo exchange grows relative to the arithmetic as partitions shrink. It also describes strong scaling, a fixed problem spread across more cores, while CFD is usually run in weak scaling. Use it for the ceiling and use the cells-per-core figure for the practical limit. If you want a real number, measure the wall clock at two core counts and solve for s; for production CFD you will get something between 10⁻⁴ and 10⁻².

Why does my run get slower when I add cores?

Almost always because the cells per core has fallen below the point where communication dominates. Each rank’s halo is proportional to its surface area while its work is proportional to its volume, so as partitions shrink the ratio of communication to arithmetic grows as the inverse cube root of the cells per core — and at some point the curve turns over. Secondary causes worth checking: the partitioning has become badly balanced, several ranks are now sharing one memory channel, or the case has crossed from one node to two and is paying for the interconnect for the first time. None of these appear in Amdahl’s law.

How do I measure the cost per cell per iteration?

Run 50 to 100 iterations on the real mesh with the real models on the real machine, take the wall clock for the iterations alone — not including reading and partitioning — multiply by the number of cores, and divide by cells times iterations. That gives microseconds of core time per cell per iteration directly. Do it before committing an allocation: it is the only linear factor in the runtime that nothing else on this page constrains, and the defaults here are a guess about your solver.

How much disk will a transient run need?

Cells times fields times bytes times writes, which the page computes. The two mistakes worth naming: writing double-precision fields for post-processing, which doubles the volume for nothing a contour plot can show; and setting the write frequency in iterations when it was meant to be in time steps, which on a run with five inner iterations per step produces five times the data and on an autotimestepping run can produce far more. Also check the average write rate against what the file system can sustain — a write that takes longer than the interval between writes throttles the solver, and the symptom looks like a solver problem rather than a storage one.

Should I use a segregated or a coupled solver if memory is tight?

Segregated, on memory grounds alone: it holds one scalar matrix at a time rather than a block matrix with sixteen or twenty-five entries per coefficient, and this page charges coupled 85% more. But the choice is not usually a memory choice. Coupled solvers converge in far fewer iterations on strongly coupled problems — high-speed compressible flow, natural convection, rotating machinery, anything where pressure and velocity are tightly linked — so a coupled run that needs a third of the iterations can be cheaper in wall clock even at twice the memory. Decide on convergence behaviour and then size the memory, not the other way round.

Why is the estimate a range rather than a number?

Because the underlying figures are. The vendor whose published numbers anchor this page states the polyhedral penalty as 20 to 40% in one document and implies 47 to 56% in another, and states the double-precision penalty as 50% while implying 56 to 60%. Other vendors differ again, CPU and GPU solvers differ, and model sets differ. A −30% to +50% band around the central estimate is honest; a single byte count would not be. If you have measured your own solver on your own model set, put that figure into the override box and the whole factor model is bypassed.

Related calculators

References

  1. Ansys, Fluent GPU Solver hardware guidance (Ansys Learning / Innovation Space knowledge article) and Fluent User’s Guide, GPU memory usage section. Source of every sourced memory factor on this page: approximately 1.0 GB of memory per million fluid cells for a tetrahedral mesh in single precision with a segregated solver, a two-equation turbulence model and the energy equation; approximately 1.2 GB for hexahedral and 1.8 GB for polyhedral on the same basis; a coupled solver about 1.85 times a segregated one; and double precision about 1.6 times single. Attributed to the named manufacturer and used as factors; the published table itself is not reproduced.
  2. Ansys, Ansys HPC Seminar 2021 — CFD presentation material (public copy hosted by CMC Microsystems), reporting Fluent scaling measurements made at Argonne National Laboratory across thousands of cores. Source of the cells-per-core guidance quoted here: excellent efficiency above about 10,000 cells per core, and good scaling down to about 2,500. Attributed, not presented as a general law.
  3. CFD Direct, OpenFOAM HPC with AWS. Source of the second set of cells-per-core figures: 90% scaling maintained at 504 cores with about 200,000 cells per core on a 97-million-cell external aerodynamics case, and departure from linear scaling from about 100 cores in a weak-scaling test held at 100,000 cells per core. Quoted beside the Ansys figures rather than in place of them, because they are different solvers on different hardware and both are right.
  4. G. M. Amdahl, Validity of the single processor approach to achieving large scale computing capabilities, AFIPS Conference Proceedings 30 (1967), 483–485. The speed-up law used here, S(P) = 1/(s + (1 − s)/P), and the ceiling Smax = 1/s.
  5. J. L. Gustafson, Reevaluating Amdahl’s Law, Communications of the ACM 31 (1988), 532–533. The reason Amdahl’s law is the wrong model for the way CFD is usually run: the problem size grows with the core count rather than staying fixed. Cited here because the page reports an Amdahl figure and needs to say what it does not mean.
  6. Verification performed for this page, not taken from a source (1): the twelve published per-million-cell figures were checked for internal consistency before any of them was used, because a shifted column in a retrieved table would not survive it. The coupled-to-segregated ratio is 1.80, 1.83 and 1.89 across the three mesh types; the double-to-single ratio is 1.60, 1.58 and 1.56; the hexahedral-to-tetrahedral ratio is 1.19 to 1.22 down all four columns. A three-factor multiplicative model of mesh type, solver and precision then reproduces all twelve entries to within 4.9%, which is what makes it legitimate to apply the factors independently.
  7. Verification performed for this page, not taken from a source (2): the two unsourced factors are labelled as such. The density-based figure of 2.30 is extrapolated from the coupled figure of 1.85 by the ratio of block sizes, 25/16 against 16/16, on the reasoning that a five-equation block matrix stores 25 sub-entries per coefficient where a four-equation one stores 16. The per-equation increment of 5% is a derived floor: one field, one old-time level, three gradient components and one face flux is about 48 bytes per cell against a base of roughly one kilobyte per cell implied by 1 GB per million cells.

Setup guidance, not validation. Correlations have ranges of validity and cell-count estimates are order-of-magnitude. A converged simulation is not a correct one. Full disclaimer at calcengines.com/disclaimer/