Mess 2.1 Released: Consumer Devices, GPUs and a New GUI
Mess 2.1 brings memory bandwidth-latency benchmarking to consumer devices, adds NVIDIA GPU measurements, and introduces a desktop interface for configuring runs and exploring results.
We are releasing Mess 2.1, expanding memory-system benchmarking beyond servers with support for consumer devices, NVIDIA GPU bandwidth-latency curves, and a new graphical interface.
From servers to laptops and desktops
Mess now covers consumer laptops, desktops and workstations alongside server systems. Apple Silicon Macs use native IOReport bandwidth measurement without an additional counter driver. Linux PCs benefit from improved Intel and AMD backend discovery and fallback. Windows x86-64 backends use Intel VTune, Intel client IMC or AMD uProf, depending on the processor and available vendor tools.
Bandwidth measurement depends on the counters supported by your processor, operating system and selected backend. Windows requires vendor tools and Administrator access; native Windows driver and benchmark-curve validation remains pending. macOS bandwidth measurement requires Apple Silicon.
NVIDIA GPU bandwidth-latency curves
The new --gpu workflow generates GPU bandwidth-latency curves with configurable read/write ratios and adaptive pause discovery. It requires an NVIDIA driver and CUDA Toolkit 12.6 or newer, including nvcc and CUPTI. GPU kernels compile at runtime, and CUPTI collects DRAM counters.
A desktop interface for Mess
Mess GUI makes it easier to configure benchmark runs, launch bandwidth-latency curves and inspect results. It uses the same benchmark engine as the command line. On first launch, choose an existing Mess checkout or let the app clone the repository, then compile the benchmark through the setup workflow.
Visit the Mess GUI download page for Linux, macOS and Windows package availability, installation guidance and platform requirements.
Better measurements and richer results
- Automatic traffic-array calibration chooses array sizes using measured bandwidth, cache capacity and memory limits.
- Mess Score summarizes bandwidth efficiency and latency growth under load, normalized against theoretical bandwidth and baseline latency.
- Refined adaptive sampling offers lite, standard and detailed run tiers, plus custom point budgets.
- A reworked plotter improves curve processing and comparison, with CSV/JSON data exports, PDF/PNG figures and instruction-sampling percentile plots.
- Improved hardware and ISA support includes AMD Zen 5 counter discovery, expanded scalar and vector kernels, and improved ARM SVE handling.
- Stability and correctness fixes improve CPU-frequency calibration, perf interval accounting, low-throughput measurements, curve processing and builds in paths containing spaces.
Optional instruction-latency sampling also gains expanded Intel PEBS and ARM SPE workflows on supported Linux systems, including improved filtering for HBM and CXL measurements.
A simpler repository name
With this release, we are moving away from the old bsc-mem/Mess-2.0 GitHub name to bsc-mem/Mess. The repository name now follows the project rather than a specific release. Update your bookmarks and repository links to the new address.
Find the source code, setup instructions and full changelog in the Mess repository.