From d2b03ec2ff782bc29dc1ba8ce8e047621af30fb0 Mon Sep 17 00:00:00 2001 From: lmoresi Date: Wed, 26 Aug 2026 16:44:08 +1000 Subject: [PATCH] Document the two HPC findings that were sitting in issues (#645, #635) Both were carefully measured reports that had nowhere to live except the issue tracker, so the next person to hit either would rediscover it. #645 -- a long output path segfaults parallel HDF5 output. The failure tracks the length of the full GENERATED filename, including the suffixes UW3 appends (output.mesh.U.00000.h5), and it fires well below the advertised limits: every component under NAME_MAX, PATH_MAX=4096, PETSC_MAX_PATH_LEN=4096, and a 286-character name still crashed. Written into the troubleshooting guide with the reported thresholds marked as observations on one stack rather than a portable limit, since the first unsafe length was never established. #635 -- Gadi rank placement for a MatMult-bound Stokes solve. ~97% of solve time inside MatMult with the solver path held fixed (four SNES, 24 KSP, same mesh hash), so runtime tracks throughput rather than convergence. The table goes into the performance guide. The binding result is deliberately NOT written as a recommendation. Explicit core binding was slower at 8 ranks, but the runs were concurrent on different nodes and a repeat of the same configuration varied by 8.6% on its own -- the node-to-node noise is the size of the effect. The doc says that outright rather than letting a reader take the table as advice. Neither issue is closed by this: #635 asks for an investigation we cannot run from here, and both keep their open questions listed in the text. Underworld development team with AI support from Claude Code --- docs/advanced/performance.md | 65 ++++++++++++++++++++++++++++++++ docs/advanced/troubleshooting.md | 39 +++++++++++++++++++ 2 files changed, 104 insertions(+) diff --git a/docs/advanced/performance.md b/docs/advanced/performance.md index 8bf2113b6..6f55498ad 100644 --- a/docs/advanced/performance.md +++ b/docs/advanced/performance.md @@ -22,6 +22,71 @@ Tools for identifying bottlenecks: - Batch operations for better performance - Parallel scaling +## Choosing MPI ranks for a MatMult-bound Stokes solve (#635) + +A large nonlinear Stokes solve is usually **memory-bandwidth bound, not +compute bound**. In a controlled Gadi campaign on a 2-D power-law Barr–Houseman +fault benchmark (P2/P1, cell size 0.01, `n=3`, viscosity contrast 1e5, tolerance +1e-8), roughly **97% of solve time was inside PETSc `MatMult`**, and every +configuration ran the identical numerical workload — four SNES solves, 24 KSP +solves, the same mesh SHA-256. + +That is the useful diagnostic: when the solver path is fixed and the runtime still +moves, you are looking at throughput, not convergence. + +| queue / launch | ranks | Stokes solve (s) | peak memory | `MatMult` (Mflop/s) | +|---|---:|---:|---:|---:| +| `normal`, default placement | 8 | 4932 | 8.99 GB | 6707 | +| `normal`, explicit core binding | 8 | 5751 | 9.13 GB | 5738 | +| `normal`, binding + no-ML attempt | 8 | 5708 | 8.97 GB | 5782 | +| `normal`, binding + no-ML attempt | 4 | 8964 | 7.31 GB | 3842 | +| `normal`, binding + no-ML attempt | 12 | **3288** | 10.1 GB | **9995** | +| `normalsr` (Sapphire Rapids) | 8 | 3992 | 9.25 GB | 8308 | + +What this supports: + +- **Runtime tracks `MatMult` throughput**, not nonlinear behaviour. Twelve ranks gave + the best wall time in this sample (56:02); four ranks cut peak memory to 7.3 GB but + took 2:30:48. +- **Fewer ranks buy memory, not speed.** If a job is near a memory ceiling, dropping + ranks is a legitimate trade — but expect the wall time to move roughly inversely. +- **Newer nodes help.** The `normalsr` 8-rank run beat the `normal` 8-rank run, though + it was still slower than 12 ranks on `normal`. + +```{warning} +**Explicit core binding did not help here**, and this campaign cannot say whether it +ever does. `--map-by core --bind-to core` was *slower* than default placement at 8 +ranks (5751 s vs 4932 s), but the configurations ran concurrently on different nodes, +and a repeat of the same default-placement case varied by 8.6% on its own (4540 s vs +4932 s). Node-to-node variation is the same size as the effect. A controlled same-node +comparison, or a socket-distributed mapping, is needed before drawing any affinity +conclusion — do not read this table as a recommendation against binding. +``` + +**Harmless Open MPI noise.** Messages of the form + +```text +[LOG_CAT_ML] component basesmuma is not available but requested in hierarchy +[LOG_CAT_ML] ml_discover_hierarchy exited with error +``` + +are **nonfatal**. Setting `OMPI_MCA_coll=^ml` neither suppressed them nor changed the +timings materially. + +**Serialise your BLAS.** Every run above set one thread per rank, which is what you +want when MPI already owns the cores: + +```bash +export OPENBLAS_NUM_THREADS=1 +export OMP_NUM_THREADS=1 +export MKL_NUM_THREADS=1 +export NUMEXPR_NUM_THREADS=1 +``` + +**Still open.** A recommended Gadi rank placement for bandwidth-bound solves, whether +UW3's examples should ship explicit mapping/binding options, and a small repeatable +PETSc timing benchmark to accompany the setup notes — see #635. + ## Related Documentation - [Developer: Performance Guidelines](../developer/guidelines/performance-optimization.md) diff --git a/docs/advanced/troubleshooting.md b/docs/advanced/troubleshooting.md index 588cbd99e..bb3458cef 100644 --- a/docs/advanced/troubleshooting.md +++ b/docs/advanced/troubleshooting.md @@ -50,3 +50,42 @@ everywhere. See the `Stokes.penalty` property docstring for the full description, and the `CONSTRAINED_FREESLIP_MULTIPLIER` design note for the derivation. + +## A long output path can segfault parallel HDF5 output (#645) + +**Symptom.** A native segmentation fault inside `mesh.write()` or +`mesh.write_timestep()` on an HPC filesystem, with no Python traceback. The mesh, +labels, fields, solve, MPI size and output calls are all valid, and the same script +succeeds when only the output directory is renamed. + +**Cause.** The failure follows the *length of the full generated filename*, not the +validity of the path. UW3 appends its own suffixes to whatever you pass — a mesh write +becomes `output.mesh.00000.h5`, and a P2 velocity variable becomes +`output.mesh.U.00000.h5` — so a descriptive case directory can push the complete path +past what the native PETSc/HDF5/MPI-I/O stack tolerates. + +```{warning} +This happens **well below** the advertised limits. In the reported case every path +component was under `NAME_MAX` (255), the filesystem allowed `PATH_MAX=4096`, and PETSc +was configured with `PETSC_MAX_PATH_LEN=4096`. A 286-character filename still crashed. +``` + +**Reported thresholds.** On Gadi (PETSc 3.25.4, HDF5 1.12.2p, Open MPI 4.1.7, MPI-enabled +h5py 3.16.0, GPFS) a staged test failed at a **259-character** filename, and 252 +characters is the shortest length observed to fail on that stack. The exact first unsafe +length was not established, so treat these as observations on one stack rather than a +portable threshold. A macOS stack (PETSc 3.25.0, HDF5 1.14.6, Open MPI 5.0.10) did not +reproduce the original failure, though the full threshold matrix was not repeated there. + +**What to do.** + +1. Keep the output root and the generated case identifier **compact**, especially on HPC. +2. Put fixed configuration in HDF5 metadata or a metrics file rather than encoding every + parameter in the directory name. A directory named for the two or three parameters + that actually vary across a campaign is enough to tell runs apart. +3. If output segfaults natively, **print the full generated filename first** — including + the UW3 suffixes — before investigating the mesh, the labels or the solver. + +Shortening the longest filename from 286 to 189 characters made an otherwise unchanged +eight-rank benchmark pass end to end, and a subsequent 192-rank run at 1/64 resolution +completed every mesh, Stokes, HDF5, XDMF and postprocessing stage.