Slurm: tungsten performance on LUMI¶
Scripts and TOML inputs for running tungsten on LUMI-C (CPU / FFTW) and LUMI-G (HIP / rocFFT).
0.2 GPU scaling campaign (start here)¶
Issue #87 first slice: size tungsten_hip on one GCD, then strong-scale
1/2/4/8 GCDs with I/O off. Recipe, account, and scratch paths:
LUMI GPU scaling campaign.
export TUNGSTEN_HIP_BIN=/path/to/tungsten_hip # 0.2 HIP build
./docs/lumi_slurm/submit_tungsten_hip_scaling.sh size
TUNGSTEN_LX=512 ./docs/lumi_slurm/submit_tungsten_hip_scaling.sh strong
PARTITION=standard-g TUNGSTEN_LX=768 ./docs/lumi_slurm/submit_tungsten_hip_scaling.sh multinode
3D FD HIP twin (heat3d_fd_hip, device halo + stencil):
export HEAT3D_HIP_BIN=/path/to/heat3d_fd_hip
./docs/lumi_slurm/submit_heat3d_fd_hip_scaling.sh size
HEAT3D_N=256 ./docs/lumi_slurm/submit_heat3d_fd_hip_scaling.sh strong
PARTITION=standard-g HEAT3D_N=256 ./docs/lumi_slurm/submit_heat3d_fd_hip_scaling.sh multinode
3D spectral HIP twin (heat3d_spectral_hip, implicit Euler, 2 FFTs/step):
export HEAT3D_SPECTRAL_HIP_BIN=/path/to/heat3d_spectral_hip
./docs/lumi_slurm/submit_heat3d_spectral_hip_scaling.sh size
HEAT3D_N=768 ./docs/lumi_slurm/submit_heat3d_spectral_hip_scaling.sh strong
PARTITION=standard-g HEAT3D_N=768 ./docs/lumi_slurm/submit_heat3d_spectral_hip_scaling.sh multinode
Files: tungsten_hip_scaling.sbatch, tungsten_hip_scaling.toml,
submit_tungsten_hip_scaling.sh, heat3d_fd_hip_scaling.sbatch,
submit_heat3d_fd_hip_scaling.sh, heat3d_spectral_hip_scaling.sbatch,
submit_heat3d_spectral_hip_scaling.sh. Account project_462001519. Do not point
these jobs at the 0.1.4 binaries or project_462001245 scratch used below.
Layout (legacy 1024³ helpers)¶
This directory (under the git repo):
*.sbatch,submit_tungsten_performance.sh,verify_gpu_aware_mpi.sh, andtungsten_performance_*.toml.Domain size:
tungsten_performance_*.tomluse 1024³ grid points (heavy run; sbatch time limit 12 h).Scratch (runtime logs and per-job working dirs):
/scratch/project_462001245/$USER/tungsten_perf_jobs/{logs,runs}/.Binaries (default):
/projappl/project_462001245/openpfc/0.1.4-hip/bin/tungstenandtungsten_hip. Override withTUNGSTEN_CPU_BIN/TUNGSTEN_HIP_BINinside the job environment if needed.
Submit all six jobs (1, 2, 4 nodes × CPU + GPU)¶
./docs/lumi_slurm/submit_tungsten_performance.sh
Or single job (CLI --nodes overrides the #SBATCH --nodes default):
sbatch --nodes=2 --job-name=tperf-cpu-2n docs/lumi_slurm/tungsten_cpu.sbatch
sbatch --nodes=2 --job-name=tperf-gpu-2n docs/lumi_slurm/tungsten_gpu.sbatch
LUMI-specific behaviour¶
CPU (tungsten_cpu.sbatch)
partition/C+small: 128 MPI ranks per node (one rank per physical core), full node (--exclusive).srun --cpu-bind=coresas recommended in Distribution and binding.TOML uses
use_gpu_aware = falsefor the FFTW path.
GPU (tungsten_gpu.sbatch)
partition/G+small-g: 8 MPI ranks per node (one per GCD),--gpus-per-node=8,--exclusive.MPICH_GPU_SUPPORT_ENABLED=1for GPU-aware MPI (required for device pointers; see INSTALL.LUMI.md).Optional smoke check:
sbatch docs/lumi_slurm/verify_gpu_aware_mpi.sh(or runverify_gpu_aware_mpifrom the OpenPFCbin/after settingVERIFY_GPU_MPI_BINif needed).tungsten_hip/heat3d_fd_hip/heat3d_spectral_hipcallbind_local_device()(local_rank % n_devices). Scaling sbatch leaves every GCD visible so GPU-aware MPI can use intra-node HIP IPC; it does not wrap ranks withROCR_VISIBLE_DEVICES=$SLURM_LOCALID.Full-node 8-GCD jobs bind 7 cores per GCD with the LUMI CCD
mask_cpumap (not a singlemap_cputhread). That cut 16-GCD 768³wall_stepfrom 138 ms to 136 ms.Optional
OPENPFC_FFT_SLAB_AXIS=x|y|zforces the 1D HeFFTe slab split.
Outputs¶
Each job creates a directory under runs/ named with the Slurm job name and ID. Profiling files use the [profiling] output = "timing_profile" stem in that directory (see docs/performance_profiling.md in the repo).
Account and partition¶
The scripts use #SBATCH --account=project_462001519. Change if your billing project differs. Use standard / standard-g instead of small / small-g if you need the default full-node binding policy without --exclusive.