Command-line interface#
HTTomo provides a command-line interface (CLI) for validating and running processing pipelines.
Before using the CLI:
outside Diamond, activate the environment in which HTTomo is installed;
at Diamond, run
module load httomo.
Outside Diamond, invoke the CLI using python -m httomo. Diamond users can
use httomo as a shortcut.
To list the available commands, run:
$ python -m httomo --help
The main commands are:
checkValidate a YAML pipeline, optionally against an input HDF5 file.
memory-checkEstimate the peak CPU memory required to run a pipeline.
runRun a pipeline on an input dataset.
Use --help after any command to see its current arguments and options:
$ python -m httomo run --help
The check command#
Validate a YAML pipeline before running it:
$ python -m httomo check PIPELINE [IN_DATA_FILE]
PIPELINEPath to the YAML pipeline to validate.
IN_DATA_FILEOptional path to the input HDF5 file. When supplied, HTTomo also checks that the dataset paths referenced by the pipeline loader exist in the file.
For details of the validation performed, see Validate the pipeline.
Note
The check command accepts pipeline files only. Checking a pipeline
supplied as a string is not currently supported.
The memory-check command#
Estimate the peak CPU memory required to process a dataset:
$ python -m httomo memory-check IN_DATA_FILE PIPELINE NPROCS
IN_DATA_FILEPath to the input HDF5 file.
PIPELINEPath to the pipeline that will process the data.
NPROCSNumber of processes that will run the pipeline. The value must be at least one.
The reported value is the estimated peak memory across all processes. It is
calculated from the estimated peak memory for one process multiplied by
NPROCS.
The estimate accounts for the input data type and dimensions, loader previews, padding, re-slicing and changes in data shape between pipeline sections. See Memory and performance for planning a process count and comparing this total with the per-process runtime ceiling.
The run command#
Run a processing pipeline:
$ python -m httomo run [OPTIONS] IN_DATA_FILE PIPELINE OUT_DIR
Arguments#
IN_DATA_FILEPath to an existing HDF5 input file.
PIPELINEPath to a YAML pipeline. HTTomo can also accept a JSON pipeline supplied as a string when
--pipeline-format jsonis used.OUT_DIRParent directory in which HTTomo creates the run output directory.
By default, the output directory is named using the start time of the run:
DD-MM-YYYY_HH_MM_SS_output
For example, a run started at 15:30:45 on 1 May 2023 with OUT_DIR set to
/home/myuser would write to:
/home/myuser/01-05-2023_15_30_45_output/
Options#
Output and intermediate data#
--output-folder-name DIRECTORYUse the given output-directory name instead of the timestamp-based default. For example,
--output-folder-name test-1createsOUT_DIR/test-1.--save-allSave intermediate datasets for every task in the pipeline. Without this option, datasets are saved only for tasks whose
save_resultsetting is enabled, either explicitly in the pipeline or by the method’s default configuration.--save-snapshotsSave image snapshots at selected points in the pipeline. Snapshots are useful for inspecting intermediate processing without saving every complete intermediate dataset.
--intermediate-format hdf5Store intermediate datasets in HDF5 format. This is currently the only supported intermediate format and is selected by default.
--compress-intermediateStore intermediate datasets in chunked HDF5 files with BLOSC compression.
--frames-per-chunk INTEGERSet the number of frames per HDF5 chunk for intermediate data. The value must be at least
-1:-1selects the chunk size automatically and is the default;0uses contiguous storage;a positive value sets the number of frames per chunk.
Compression requires chunked storage. If
--compress-intermediateis combined with--frames-per-chunk 0, HTTomo changes the chunk setting to-1and selects it automatically.--recon-filename-stem NAMESet the filename stem used for reconstruction output. HTTomo adds the
.h5extension. For example,--recon-filename-stem my-reconproducesmy-recon.h5.
Execution and resource use#
--gpu-id INTEGERSelect the GPU device to use. The default is
-1, which does not explicitly select a different CUDA device.--max-memory SIZESet a per-process memory ceiling. Values may be supplied as bytes or with a
K,MorGsuffix, for example--max-memory 32G.When the estimated memory for a pipeline section reaches this limit, HTTomo uses disk-backed intermediate storage. For GPU sections, the same value also caps the memory budget used to calculate the block size; HTTomo uses the smaller of this ceiling and the available GPU memory. The default is
0, which disables the user-supplied ceiling. GPU block sizing still respects the memory reported by the device.See Memory and performance for practical sizing guidance.
--max-cpu-slices INTEGERSet the maximum number of slices in a block for CPU-only pipeline sections. The value must be at least one and defaults to
64.Adjusting this value may affect the performance of CPU-only processing. See Core concepts for information about blocks, chunks and sections.
--reslice-dir DIRECTORYChoose the directory used for temporary re-slicing files. The directory must already exist and be writable. The run output directory is used by default.
When the output is on network-mounted storage, using a local temporary directory can substantially improve file-based re-slicing performance. For a multi-node run, the directory must be accessible to every participating process.
--continuous-scan-subset START STOPSelect a subset of projections along the angular dimension. This option overrides the
continuous_scan_subsetvalue in the pipeline loader configuration. See Continuous-scan subsets.--mpi-abort-hookAbort all MPI processes when any process encounters an unhandled exception. This prevents the remaining processes from waiting indefinitely for a failed process.
This option is mainly intended for debugging. Because termination occurs at the MPI level, the exception traceback may be incomplete.
Pipeline format and parameter sweeps#
--pipeline-format {yaml,json}Select the pipeline format. The value is case-insensitive and defaults to YAML.
YAML pipelines must be provided as files. JSON pipelines must be provided as strings.
--bits-sweep-images INTEGERSet the bit depth of TIFF images produced by a Parameter Sweeping run. Use
8,16or32. The default is32.The CLI currently accepts any integer, although the supported output bit depths are 8, 16 and 32.
Monitoring#
--monitor NAMEEnable a performance monitor. The available monitors are
summaryandbench. This option can be supplied more than once.summaryReport aggregate timings and a per-method breakdown.
benchReport detailed timings for every process, including CPU and GPU execution, data transfers and file operations.
--monitor-output FILENAMEWrite monitoring results to a file. By default, results are written to standard output.
The
summarymonitor produces human-readable text, while thebenchmonitor produces CSV data.
System logging#
--syslog-host HOSTSet the hostname of the syslog server. The default is
localhost.--syslog-port PORTSet the syslog server port. The default is
514.