FPGA accelerator for an incomplete-Cholesky preconditioned conjugate-gradient
(ICCG) solver, written in Vitis HLS 2024.2 C++ (hls::task data-driven task
networks) for the Alveo U55C. It is built from two TAPA HLS reference
projects, Callipepla (Jacobi-preconditioned CG) and LevelST (sparse
triangular solver).
callipepla/ Callipepla kernel, host program, HLS / link configuration
callipepla/floorplan/ floorplan spec of the kernel for the U55C
levelst/ LevelST triangular solver kernel (stand-alone and as IC preconditioner),
host programs, HLS / link configuration
include/ shared host headers (mmio.h, sparse_helper.h, host_utils.h,
levelst_prep.h = LevelST matrix scheduling / packing)
tools/hls_floorplan/ floorplanner for hls::task kernels (relay stations + pblocks)
docs/ development notes (see below)
Documentation:
- docs/callipepla-vitis-port.md: how the TAPA constructs were mapped to Vitis HLS, deviations from the TAPA source, simulation and synthesis status.
- docs/floorplan.md: floorplanning the kernel (relay
stations, pblock constraints, HBM binding) with
tools/hls_floorplan. - docs/callipepla-tapa-comparison.md: implementation and on-board comparison with the TAPA/AutoBridge build and the authors' prebuilt bitstream.
- docs/levelst-vitis-port.md: the LevelST port, its software-simulation status, and the design of the incomplete-Cholesky preconditioner built from it (how it plugs into Callipepla).
- Rewrite Callipepla in Vitis HLS C++ (
hls::task) - Verify the rewrite with software simulation (residuals match a CPU reference and, on the board, the TAPA bitstream bit for bit)
- C synthesis and C/RTL co-simulation of the rewrite
- Floorplan the kernel for Vivado (slot assignment + relay stations
derived from the AutoBridge result of the TAPA build) —
tools/hls_floorplan; implementation compared with the TAPA build in docs/callipepla-tapa-comparison.md - Implementation:
callipepla/build_hw.sh→Callipepla.xclbin(U55C, kernel clock 231.9 MHz) - Run the rewritten version on-board and compare with the TAPA version (41 cases bit-exact, 1% slower per iteration; docs/callipepla-tapa-comparison.md)
- Rewrite LevelST in Vitis HLS C++ (
hls::task), verified by software simulation against CPU forward substitution (levelst/) - Modularize: the solver core is a reusable task network
(
levelst/src/levelst_core.h); it runs stand-alone (TrigSolver) and as the preconditionerz = L^-T L^-1 r(TrigPrecond, two solves per application, verified in an ICCG loop in software simulation) - C synthesis of
TrigSolver(levelst/hls_work/TrigSolver.xo, all loops II=1, 3.49 ns estimated; see docs/levelst-vitis-port.md) - C/RTL co-simulation of
TrigSolver, C synthesis ofTrigPrecond - Plug the preconditioner into Callipepla with a Jacobi / incomplete Cholesky compile-time switch (design in docs/levelst-vitis-port.md; waits for the Callipepla bitstream work to finish)
- Build, run and verify the whole ICCG solver
- Vitis / Vivado 2024.2 at
/tools/Xilinx(source env.sh), platformxilinx_u55c_gen3x16_xdma_3_202210_1, XRT at/opt/xilinx/xrt. - Python 3 (standard library only) for
tools/hls_floorplan. - Test matrices (Matrix Market
.mtx) come from the SuiteSparse Matrix Collection;scripts/get_matrices.shdownloadsLFAT5(Oberwolfach) andmhd3200b(Bai) into the git-ignoredmatrices/directory (matrices/<name>/<name>.mtx); pass<group>/<name>arguments for others.
scripts/get_matrices.sh # once: matrices/LFAT5/LFAT5.mtx, matrices/mhd3200b/mhd3200b.mtx
cd callipepla
make # g++ with the Vitis HLS headers (override VITIS_DIR=...)
./build/callipepla ../matrices/mhd3200b/mhd3200b.mtx 100source env.sh
cd callipepla
v++ -c --mode hls --config hls_config.cfg --work_dir hls_work # -> hls_work/Callipepla.xo (~50 min)
vitis-run --mode hls --cosim --config hls_config.cfg --work_dir hls_workThe co-simulation needs a workaround for an xelab crash, see
docs/callipepla-vitis-port.md. Matrix and
iteration count for csim/cosim are set in hls_config.cfg (LFAT5 from
matrices/, 10 iterations).
cd levelst
make # build/trigsolver, build/trigprecond
./build/trigsolver ../matrices/mhd3200b/mhd3200b.mtx # L = lower triangle of A, solve L x = f
./build/trigsolver synth:300000:4:7 # synthetic 300k-row system (3 rounds)
./build/trigprecond ../matrices/mhd3200b/mhd3200b.mtx 8 # IC(0): one apply + 8 ICCG iterationstrigsolver compares the kernel's x with CPU forward substitutions in
float and double; trigprecond compares z = L^-T L^-1 r with a double
CPU solve and then runs a CPU conjugate-gradient loop that calls the kernel
for every preconditioner application. C synthesis:
source env.sh
cd levelst
v++ -c --mode hls --config hls_config.cfg --work_dir hls_work # TrigSolver -> hls_work/TrigSolver.xo
v++ -c --mode hls --config hls_precond.cfg --work_dir hls_work_precond # TrigPrecondsource env.sh
cd callipepla
./build_hw.sh # hw (default) or hw_emu; output in build_hw/<target>/build_hw.sh runs tools/hls_floorplan on hls_work/Callipepla.xo with
floorplan/u55c_halfslr.json (patched xo + floorplan.tcl +
floorplan_report.txt) and then v++ --link with link_config.ini.
To inspect the kernel's tasks and streams, or to edit the spec:
python3 ../tools/hls_floorplan list --xo hls_work/Callipepla.xosource /opt/xilinx/xrt/setup.sh
cd callipepla
make xrt # build/callipepla-xrt (host + XRT native API)
./build/callipepla-xrt ../matrices/mhd3200b/mhd3200b.mtx 100 build_hw/hw/Callipepla.xclbinThe same program without the xclbin argument runs the software model, so
the two residual sequences can be diffed directly. make board builds
build/callipepla-board, which writes the result-directory format of the
TAPA campaign for file-by-file comparisons (see the comparison document). The card must carry the
xilinx_u55c_gen3x16_xdma_3_202210_1 shell (xrt-smi examine).