Slurm: tungsten performance on LUMI

Scripts and TOML inputs for running tungsten on LUMI-C (CPU / FFTW) and LUMI-G (HIP / rocFFT).

0.2 GPU scaling campaign (start here)

Issue #87 first slice: size tungsten_hip on one GCD, then strong-scale 1/2/4/8 GCDs with I/O off. Recipe, account, and scratch paths: LUMI GPU scaling campaign.

export TUNGSTEN_HIP_BIN=/path/to/tungsten_hip   # 0.2 HIP build
./docs/lumi_slurm/submit_tungsten_hip_scaling.sh size
TUNGSTEN_LX=512 ./docs/lumi_slurm/submit_tungsten_hip_scaling.sh strong
PARTITION=standard-g TUNGSTEN_LX=768 ./docs/lumi_slurm/submit_tungsten_hip_scaling.sh multinode

3D FD HIP twin (heat3d_fd_hip, device halo + stencil):

export HEAT3D_HIP_BIN=/path/to/heat3d_fd_hip
./docs/lumi_slurm/submit_heat3d_fd_hip_scaling.sh size
HEAT3D_N=256 ./docs/lumi_slurm/submit_heat3d_fd_hip_scaling.sh strong
PARTITION=standard-g HEAT3D_N=256 ./docs/lumi_slurm/submit_heat3d_fd_hip_scaling.sh multinode

3D spectral HIP twin (heat3d_spectral_hip, implicit Euler, 2 FFTs/step):

export HEAT3D_SPECTRAL_HIP_BIN=/path/to/heat3d_spectral_hip
./docs/lumi_slurm/submit_heat3d_spectral_hip_scaling.sh size
HEAT3D_N=768 ./docs/lumi_slurm/submit_heat3d_spectral_hip_scaling.sh strong
PARTITION=standard-g HEAT3D_N=768 ./docs/lumi_slurm/submit_heat3d_spectral_hip_scaling.sh multinode

Files: tungsten_hip_scaling.sbatch, tungsten_hip_scaling.toml, submit_tungsten_hip_scaling.sh, heat3d_fd_hip_scaling.sbatch, submit_heat3d_fd_hip_scaling.sh, heat3d_spectral_hip_scaling.sbatch, submit_heat3d_spectral_hip_scaling.sh. Account project_462001519. Do not point these jobs at the 0.1.4 binaries or project_462001245 scratch used below.

Layout (legacy 1024³ helpers)

  • This directory (under the git repo): *.sbatch, submit_tungsten_performance.sh, verify_gpu_aware_mpi.sh, and tungsten_performance_*.toml.

  • Domain size: tungsten_performance_*.toml use 1024³ grid points (heavy run; sbatch time limit 12 h).

  • Scratch (runtime logs and per-job working dirs): /scratch/project_462001245/$USER/tungsten_perf_jobs/{logs,runs}/.

  • Binaries (default): /projappl/project_462001245/openpfc/0.1.4-hip/bin/tungsten and tungsten_hip. Override with TUNGSTEN_CPU_BIN / TUNGSTEN_HIP_BIN inside the job environment if needed.

Submit all six jobs (1, 2, 4 nodes × CPU + GPU)

./docs/lumi_slurm/submit_tungsten_performance.sh

Or single job (CLI --nodes overrides the #SBATCH --nodes default):

sbatch --nodes=2 --job-name=tperf-cpu-2n docs/lumi_slurm/tungsten_cpu.sbatch
sbatch --nodes=2 --job-name=tperf-gpu-2n docs/lumi_slurm/tungsten_gpu.sbatch

LUMI-specific behaviour

CPU (tungsten_cpu.sbatch)

  • partition/C + small: 128 MPI ranks per node (one rank per physical core), full node (--exclusive).

  • srun --cpu-bind=cores as recommended in Distribution and binding.

  • TOML uses use_gpu_aware = false for the FFTW path.

GPU (tungsten_gpu.sbatch)

  • partition/G + small-g: 8 MPI ranks per node (one per GCD), --gpus-per-node=8, --exclusive.

  • MPICH_GPU_SUPPORT_ENABLED=1 for GPU-aware MPI (required for device pointers; see INSTALL.LUMI.md).

  • Optional smoke check: sbatch docs/lumi_slurm/verify_gpu_aware_mpi.sh (or run verify_gpu_aware_mpi from the OpenPFC bin/ after setting VERIFY_GPU_MPI_BIN if needed).

  • tungsten_hip / heat3d_fd_hip / heat3d_spectral_hip call bind_local_device() (local_rank % n_devices). Scaling sbatch leaves every GCD visible so GPU-aware MPI can use intra-node HIP IPC; it does not wrap ranks with ROCR_VISIBLE_DEVICES=$SLURM_LOCALID.

  • Full-node 8-GCD jobs bind 7 cores per GCD with the LUMI CCD mask_cpu map (not a single map_cpu thread). That cut 16-GCD 768³ wall_step from 138 ms to 136 ms.

  • Optional OPENPFC_FFT_SLAB_AXIS=x|y|z forces the 1D HeFFTe slab split.

Outputs

Each job creates a directory under runs/ named with the Slurm job name and ID. Profiling files use the [profiling] output = "timing_profile" stem in that directory (see docs/performance_profiling.md in the repo).

Account and partition

The scripts use #SBATCH --account=project_462001519. Change if your billing project differs. Use standard / standard-g instead of small / small-g if you need the default full-node binding policy without --exclusive.