MI300A high-order finite-element kernels and HiFiMagnet

Forward-performance benchmark: solve the nonlinear thermo-electric high-field-magnet model at fixed accuracy and expose state fields, nonlinear/linear solver behaviour, device residency and CPU/APU performance; no inverse or control objective is introduced in this application.

Quick Facts

Type

Extended Mini App

Status

Planned

Work Packages

WP1, WP3, WP7

Frameworks

Feel++, PETSc, HPDDM

1. Overview

Forward-performance benchmark: solve the nonlinear thermo-electric high-field-magnet model at fixed accuracy and expose state fields, nonlinear/linear solver behaviour, device residency and CPU/APU performance; no inverse or control objective is introduced in this application.

2. Technical Stack

2.1. Frameworks Used

  • Feel++

  • PETSc

  • HPDDM

2.2. Parallel Frameworks

  • MPI

  • Kokkos

  • GPU - HIP

  • Multithread - OpenMP

3. Methods & Algorithms

3.1. WP1: Discretization

  • high-order cG/H1 and HDG

  • full/partial/matrix-free assembly

  • homogeneous worksets

  • facets

  • traces

  • coupled blocks

  • skeleton terms

3.2. WP3: Solvers

  • PETSc/HPDDM Krylov solvers

  • preconditioning

  • nonlinear and block solvers

  • device-resident sparse algebra

4. Data Flow

4.1. Inputs

  • Gmsh

  • JSON config

  • Spack environment

  • Apptainer image

4.2. Outputs

  • Electric potential φ

  • temperature T

  • unknown potential drop ΔV

  • current density J

  • Joule losses

  • JSON/CSV performance reports

  • profiles

  • HDF5 fields

  • reproducible native and container environments

5. Benchmarking

5.1. Metrics

  • Forward Residual

  • Discretization Error

  • Fixed Accuracy

  • Time To Solution

  • Time To Precision

  • Nonlinear And Linear Iterations

  • Assembly Cost

  • Solve Cost

  • Memory Residency

  • Host Device Transfers

  • Strong Scalability

  • Weak Scalability

  • Multi Apu Scalability

5.2. Benchmark Scope

  • Single Node

  • Multi Node

  • Multi Gpu

  • Full System

  • Method Verification

  • Solver Scaling

6. Timeline

Milestone Date

Specification Due

N/A

Prototype Due

N/A

7. Team

Partners: Unistra, Sorbonne U, LNCMI/CNRS, CINES

Responsible: C. Prud’homme; V. Chabannes; P. Jolivet

WP7 Engineer: Javier Cladellas (UNISTRA)

8. Notes

Common forward operator used by all three applications: state y=(φ,T,ΔV), J=-σ(T;p)∇φ, ∇·J=0, -∇·(k(T;p)∇T)=J·E, V=0 on Γ-, V=ΔV on Γ+, ∫Γ+ J·n ds=Iset, with cooling boundary conditions. This row benchmarks only the forward map and its cG/H1 and HDG implementations. Grand Challenge AMD MI300A at CINES/Adastra: 319 MI300A in September-October 2026, then 128 MI300A from November 2026 for six months. Compare native Spack and Apptainer execution, CPU multicore and hybrid CPU+GPU/APU.