blog

Lecture 4: PROFET: No-Stress Performance Prediction

PROFET predicts how changing main memory affects application performance and system power and energy consumption, without simulating the CPU.

Lecture 4: PROFET: No-Stress Performance Prediction recording
Session Recording

Suppose you want to know what happens to your application if you change the server's main memory. The traditional answer is detailed hardware simulation, and anyone who has done it knows the experience: pain, agony and torture. It's slow, and to study a memory system you first have to simulate an entire, already existing CPU, just so it can generate the traffic that drives the memory simulator.

There has to be a better way to answer the question: how much will my application's performance, power and energy change if we change the main memory?

In the fourth lecture of It's the Memory, Stupid!, we introduce PROFET (Performance and Energy Prediction). PROFET was developed at BSC to quantify how main memory affects application performance and system power and energy consumption. It targets mainstream high-end CPUs and lets you explore different main memory options without simulating anything.

PROFET was the first paper from the BSC memory team to use bandwidth–latency curves to characterize a system. Those curves are now obtained with the Memory Stress Framework (Mess), which we covered in the previous lecture.

How does PROFET work? First, you profile your application on a baseline system: a real server you already have, such as an Intel Granite Rapids machine with RDIMM-6400. Then you define a target system. The CPU stays the same and only the memory changes: for example, you might ask how much your application would gain from upgrading that server from RDIMM-6400 to MRDIMM-8800. Finally, PROFET answers that question: it predicts the application's performance, power and energy on the MRDIMM-8800 system.

What goes into PROFET

PROFET uses memory bandwidth–latency curves, CPU parameters and application hardware counters to estimate performance.
Diagram of the whole process of performance estimation.

Here's the surprising part: PROFET does not need a simulator for any of its three inputs. Everything it uses already exists, sitting in datasheets, vendor reports and hardware counters.

The memory: bandwidth–latency curves. One set describes a familiar baseline, such as directly attached DDR4 or DDR5. The other describes the target, and here's where it gets interesting. The target curves can come from almost anywhere: real hardware, a development board, a prototype, a manufacturer's simulations or estimates, or even a sensitivity analysis. That means PROFET can explore memory systems that haven't been built yet.

The CPU: numbers from the datasheet. Re-order buffer (ROB) capacity, miss status holding register (MSHR) capacity, minimum theoretical CPI, and last-level cache (LLC) latency. No microarchitectural deep dive required.

The application: a profile from real hardware. A run on the baseline system, say Intel Granite Rapids with RDIMM-6400, captures a handful of standard counters over time: read and write memory bandwidth (just like in Mess application profiling), CPU cycles, instruction count, and LLC misses.

Memory, CPU, application. Just three inputs and the model has everything it needs.

Inside the model

The PROFET performance model follows an old-school electrical engineering mindset: model the circuit as a system of equations, then solve the system. This makes it an unconventional approach to memory design-space exploration, and it rests on two ideas.

Idea 1: an application can be positioned on the memory curves. If you've followed the series this far, we probably agree on this one already.

Idea 2: we can predict where the application will land on the target system's curves. The full system of equations is in the paper and the Git repository. If you'd like the details, or want to discuss PROFET, feel free to reach out to us (mariana.carmin@bsc.es).

PROFET prediction for the LMbench memory copy kernel moving from RDIMM to MRDIMM: predicted performance difference 33%, measured 34%.
Example with RDIMM to MRDIMM prediction of the LMbench benchmark memory copy kernel.

Does it actually work?

Yes. Across very different platforms and memory technologies, mispredictions stay in the low single digits on all evaluated systems, the sole expectation is DRAM → Optane whose behavior differs fundamentally from DRAM:

  • Intel Sandy Bridge (DDR3:800/1066/1333/1600): 5% mispredictions
  • Intel Knights Landing (DDR4-2400/MCDRAM): 4% mispredictions
  • Huawei Kunpeng 920 (DDR4 1600/1866/2933): 4.5% mispredictions
  • Intel Cascade lake DDR4 alone → DDR4+Optane: 0.7% mispredictions
    • DDR4 2666 → 2133: 0.7% mispredictions
    • DDR4 → Optane: 26% mispredictions
  • Intel Emerald Rapids (DDR5 4800 → 3200): 3% mispredictions
  • Intel Granite Rapids (RDIMM-6400 → MRDIMM-8800): 4% mispredictions

The story isn't finished yet. Work in progress includes Intel Max with DDR5-4800 + HBM2, and Intel Emerald Rapids with DDR5-4800 + a CXL memory expander.

No-Stress Performance Prediction!

If you thought predicting the impact of a new memory system meant weeks of simulating a CPU you already own, take a look at PROFET. Profile once and find out whether that memory upgrade is really worth it. No pain, no agony, no torture.

One model. Many memory systems. Openly available.

PROFET is open source and available in the BSC Memory team repository. The package also includes some memory system profiles, CPU parameters and application profiles. That means you can reproduce the results. The work is described in PROFET: Modeling System Performance and Energy Without Simulating the CPU, published at ACM SIGMETRICS 2019. For a closer look at the model in practice, the full lecture is available online.