Skip to main content
Version: 3.0.0 (experimental)

Hello World

A minimal DMR application that shows how DMR reconfigurations work without transferring application data.

Source: the hello-world example repository.

What it does​

  1. Initialises MPI and DMR.
  2. Registers restart, checkpoint, and finalize hooks through DMR_AUTO.
  3. Registers the built-in round policy and requests reconfiguration with USE_POLICY.
  4. Prints which rank is running, checkpointing, restarting, or exiting.
  5. Stops after MAX_ITERS reconfigurations.

The same source can be compiled and launched for DMR@Jobs or MiniDMR.

Prerequisites​

Clone or enter the example repository, and check out the v3 branch:

cd hello-world
git switch v3

The Makefile expects DMR_PATH to point to the DMR installation, and compiles with -DDMR_WITH_TEST_POLICIES so the built-in round policy (dmr_get_policy_round) is declared; see Policy Headers.

Choose an execution mode​

DMR@Jobs uses the system Slurm instance. On MN5, use the pre-built DMR module:

module load dmr

Check that the module exported DMR_PATH:

echo "$DMR_PATH"

Compile:

make clean
make

Configure the MN5 batch script (start_dmratjobs.sh):

#SBATCH --time=00:30:00
#SBATCH --exclusive
#SBATCH -N1
#SBATCH --qos=gp_bsccs
#SBATCH -A bsc85

export DMR_PROCS_PER_NODE=1
export DMR_DEFAULT_POLICY_MIN=1
export DMR_DEFAULT_POLICY_MAX=2

Run:

sbatch start_dmratjobs.sh

start_dmratjobs.sh builds the PRRTE host list from the Slurm allocation and launches:

$DMR_PATH/bin/dmr_wrapper mpirun --host $NODELIST_WITH_COUNTS ./hello-world

This mode drives the reconfiguration through the round policy: it expands until DMR_DEFAULT_POLICY_MAX nodes, then shrinks back to DMR_DEFAULT_POLICY_MIN, up to MAX_ITERS reconfigurations.

The exact node names and rank ordering depend on the allocation. The output should show rank 0 reporting the current reconfiguration count, ranks checkpointing/finalizing before they leave, restarted ranks joining the new allocation, and a final Goodbye world line once MAX_ITERS reconfigurations are reached. For example:

[1/1] Hello world from mc-slurmd-1. DMR's reconfiguration count is 0. Suggestion to DMR is: USE_POLICY (round policy).
mc-slurmd-1 rank 0 checkpointed. In a real program, the current process would save some data..
mc-slurmd-1 rank 0 restarted. In a real program, the current process would read some data.
[1/2] Hello world from mc-slurmd-1. DMR's reconfiguration count is 1. Suggestion to DMR is: USE_POLICY (round policy).
...
Goodbye world from rank 0 on mc-slurmd-1. DMR's reconfiguration count is 4.

Key points​

  • Reconfiguration bounds and stride for the round policy come from DMR_DEFAULT_POLICY_MIN, DMR_DEFAULT_POLICY_MAX and DMR_DEFAULT_POLICY_STRIDE, read when dmr_set_policy(dmr_get_policy_round()) registers the policy; they are not set in the source. See Built-in Policies.
  • restart prints that the rank restarted; a real program would reload or rebuild its data.
  • checkpoint prints that the rank checkpointed; a real program would save or transfer data before reconfiguration.
  • finalize prints that the rank is about to exit; a real program would release resources.
  • The example uses an infinite wait loop because expansion requests are non-blocking by default.
  • During a reconfiguration, ranks call the checkpoint/finalize hooks before exiting, and restarted ranks call the restart hook after DMR relaunches the program.