Memory and performance#
HTTomo divides each section between MPI processes and then processes
each chunk in memory-sized blocks. Two controls help
when planning a run: memory-check estimates host-memory demand before the
run, while --max-memory limits memory use by each process at runtime.
Estimate CPU memory before a run#
Run the estimate with the same input, pipeline and process count that will be used for processing:
$ python -m httomo memory-check INPUT.nxs PIPELINE.yaml NPROCS
The result is an estimated peak across all NPROCS processes. Divide it by
NPROCS for an approximate per-process value. The calculation includes the
input data type and preview, section padding, shape changes and memory required
for a re-slice.
Leave additional capacity for Python, MPI, HDF5, the operating system and other jobs. The command estimates HTTomo’s main section storage; it is not a guarantee that the complete process will remain below the reported value.
Choose a process count#
More processes divide the section data into smaller chunks, but every process has runtime overhead and may allocate method-specific buffers. For GPU runs, start with one process per GPU. For CPU runs, increase the process count only while the machine has sufficient memory and I/O bandwidth.
Use a runtime memory ceiling#
--max-memory is a per-process ceiling, unlike the total reported by
memory-check:
$ python -m httomo run INPUT.nxs PIPELINE.yaml OUTPUT \
--max-memory 32G
When a section’s estimated host-memory requirement reaches the ceiling, HTTomo
uses a temporary HDF5-backed store instead of keeping the section data in RAM.
For GPU sections, the same value caps the memory budget used to choose block
sizes; available device memory still provides an upper bound. A value of 0
disables the user ceiling.
Disk-backed sections protect memory at the cost of extra I/O. The warning
Chunk does not fit in memory - using a file-based store indicates that this
path was selected. Put --reslice-dir on fast storage that every
participating process can access.
Reduce resource use#
If an estimate or run is too large:
crop unused detector regions with preview;
process a smaller angular range while developing a pipeline;
reduce the number of simultaneous processes if aggregate host memory is the limit;
lower
--max-cpu-slicesfor CPU-only sections;lower
--max-memoryto reduce GPU blocks or select disk-backed storage;avoid
--save-allunless every intermediate result is needed; anduse
--compress-intermediatewhen storage capacity matters more than compression overhead.
For the complete option definitions, see Command-line interface. Developers implementing or correcting method estimates should read Memory estimation and block sizing.