King

Ramulator-Gate

When a benchmark exposes an uncomfortable truth, the community has two choices: Examine the evidence, or attack the mirror.

June 2026

What is this?

At MICRO 2024, we published our paper on the Mess benchmark, showing how it can be used to evaluate the memory performance of both real systems and simulators.

The work was built on years of experience in memory-systems research, including numerous industry collaborations. Our evaluation followed a realistic user scenario: we relied on publicly available installation and setup information, and unless we encountered major issues, we did not request additional support from simulator developers.

Some memory models performed well.

Some did not. Before publication, the work was presented at the MEMSYS 2024 panel. What was planned as a five-minute opening statement turned into a forty-minute technical discussion. I later received the award for “Most controversial, truth-telling panel discussion.”

Many in the community understood the importance of the problem and gave a name to what the results exposed:

"The emperor is naked."

The paper later received the Best Paper Runner-Up Award at MICRO 2024. The message resonated.

We opened a technical discussion with the Ramulator team even before the paper was published. Over the following year, we provided detailed technical support, addressed the technical questions raised, shared the files and instructions needed to reproduce the experiments, and proposed constructive ways forward.

What followed, however, was not merely a technical disagreement. The issue was progressively reframed away from the evidence and into a narrative that misrepresents the technical record.

We chose not to respond publicly to earlier preprints, despite what we consider serious misrepresentations of the Mess study. However, the publication of this work in an IEEE venue changes the situation. The technical record matters.

We will not allow carefully crafted narratives to replace open technical discussion.

It is rebuttal time.

Section 01

Memory bandwidth

A simulator result can look right for the wrong reason.

That is where the first chapter of Ramulator-gate begins.

At MICRO 2024, we published our paper on the Mess benchmark, showing how it can be used to evaluate the memory performance of real systems and simulators [1].

A year later, the Ramulator team published a preprint that characterized the Mess paper’s results and conclusions as incorrect due to “multiple trivial human errors” and “gross configuration errors” [2]. In 2026, this work was accepted to ISPASS [3]. The wording about “trivial” and “gross” errors was removed, but the main narrative remained.

We chose not to respond publicly to earlier preprints, despite what I consider serious misrepresentations of the Mess study. However, publication in an IEEE venue changes the situation.

In this series of posts, we will challenge several statements and results presented by the Ramulator team.

The first chapter of Ramulator-gate shows why better-looking results are not always the right results.

On one technical point, we agree with the Ramulator team: to resemble the bandwidth of the actual 8xDDR5-4800 Amazon Graviton3 and Intel Sapphire Rapids systems, Ramulator 2.0 requires a 16-channel DDR5-4800 configuration. With 8 channels, Ramulator 2.0 simulates roughly half of the expected bandwidth.

Where we disagree is whether the 16-channel Ramulator 2.0 configuration was obvious in this context.

This distinction matters.

The Ramulator team does not provide actual system measurements for the evaluated baseline and does not explicitly state that the baseline real systems are used to guide the simulator configuration. As a result, the 16-channel simulator configuration is presented in a way that appears more natural for an unspecified “ARM-based real system” than for Amazon Graviton3 with 8-channel DDR5-4800 main memory.

This is where narrative starts to diverge from the technical record.

And this is only the beginning.

In the next chapters, we will directly challenge the Ramulator team’s position that “channel” and “subchannel” are not distinct concepts in the DDR5 context.

Stay tuned!

  1. Mess paper evaluation

    The Mess paper measures two real systems and simulates Ramulator 2.0 against them. The detailed single-channel run is scaled ×8, and the simulated peak lands at roughly half of the measured bandwidth.

    Real-system measurements
    Amazon Graviton 3: 8×DDR5-4800
    Intel Sapphire Rapids: 8×DDR5-4800
    Ramulator 2.0 in the Mess paper
    1×DDR5-4800 simulated in detail, bandwidth scaled ×8
    SystemMax BW
    Amazon Graviton 3 · 8×DDR5-4800292 GB/s
    Intel Sapphire Rapids · 8×DDR5-4800264 GB/s
    Simulation: 1×DDR5-4800 scaled ×8: gem5 + Ramulator 2.0141 GB/s
    Simulation: 1×DDR5-4800 scaled ×8: Trace-driven Ramulator 2.0126 GB/s

    Conclusion

    Max simulated BW ≈ ½ of the actual one.

  2. Ramulator team re-evaluation

    The Ramulator team does not provide actual system measurements for the evaluated baseline, and does not state that the baseline real systems guide the simulator configuration. As a result, the 16-channel configuration is presented as more natural for an unspecified “ARM-based real system” than for the 8×DDR5-4800 Amazon Graviton 3 it actually reproduces.

    Real system under study
    NOT SPECIFIED
    Real-system measurements
    NONE.
    Reproducing “Mess results for an ARM-based real system” to Amazon Graviton 3: 8×DDR5-4800
    SystemMax BW
    Reproduced from the Mess paper: Amazon Graviton 3 · 8×DDR5-4800292 GB/s
    Simulation: 16×DDR5-4800 Trace-driven Ramulator 2.0281 GB/s

    Conclusion

    “By correctly configuring Ramulator 2.0, simulated memory performance resembles real system characteristics well.”

    Correctly configuring here means using a 16×DDR5-4800 Ramulator 2.0 setup to simulate the 8×DDR5-4800 actual system.

    From the Ramulator re-evaluation paper:

  3. Discussion

    On one technical point, we agree: to resemble the bandwidth of the actual 8×DDR5-4800 systems, Ramulator 2.0 needs a 16-channel DDR5-4800 configuration. Where we disagree is whether that 16-channel configuration was obvious in this context. This distinction matters.

    We agree
    To resemble the bandwidth of the actual Amazon Graviton 3 and Intel Sapphire Rapids systems, Ramulator 2.0 needs a 16-channel DDR5-4800 configuration.
    We disagree
    On whether the 16-channel Ramulator 2.0 configuration was obvious in this context.
    To date
    The Ramulator team has not acknowledged that Amazon Graviton 3 and Intel Sapphire Rapids have 8 memory channels, although Intel, Amazon and numerous independent resources state this. See the sources below.

    Evidence library

    Public sources on the evaluated platforms

Take out

To resemble the bandwidth of the actual 8×DDR5-4800 Amazon Graviton 3 and Intel Sapphire Rapids systems, Ramulator 2.0 needs a 16-channel DDR5-4800 configuration.

With 8 channels, Ramulator 2.0 simulates roughly half of the expected bandwidth.

References

  1. [1]Esmaili-Dokht et al. “A Mess of Memory System Benchmarking, Simulation and Application Profiling,” MICRO 2024. Best paper runner-up award.
  2. [2]Haocong Luo, Ataberk Olgun, Maria Makeenkova, F. Nisa Bostancı, Geraldo F. de Oliveira, A. Giray Yağlıkçı, Onur Mutlu, “Cleaning up the Mess,” arXiv Oct. 2025.
  3. [3]F. Nisa Bostancı, Haocong Luo, Ataberk Olgun, Maria Makeenkova, Geraldo F. de Oliveira, A. Giray Yağlıkçı, Onur Mutlu, “Cleaning up the Mess: Re-Evaluating the Real-System Modeling Accuracy of Ramulator 2.0”. ISPASS 2026.

DDR5 DIMMs: Channel vs. sub-channel

DDR5 channels and sub-channels: Terminology matters

In the previous post, we saw that some memory simulators require 16 DDR5 memory channels to simulate actual 8-channel systems, such as Intel Sapphire Rapids and Amazon Graviton 3.

In the next two posts, I will explain why this happens.

It took us 10 months of email exchange with the Ramulator team to understand the technical reasoning behind this setup: channel and sub-channel terminology in DDR5 DIMMs.

A common understanding today is that a DDR5 DIMM is connected to the CPU through a DDR5 channel, which is partitioned into two sub-channels. In the HPC domain, which is the focus of our studies, this terminology is broadly used.

At the same time, some sources use a different convention: 1 DIMM = 2 channels.

To make things even more interesting, these two conventions are sometimes mixed within the same document.

My objective here is not to criticize these inconsistencies.

The documents discussed in this post cover everything from DDR5 chips to HPC servers, so it is understandable that terminology may vary. Also, the sub-channel concept was introduced with DDR5 DIMMs, so the community may simply need more time to converge on a consistent convention.

My objective is to point to a practical problem:

To prevent confusion, especially when the intended meaning differs from the mainstream convention, the term “DDR5 channel” should be defined explicitly.

In the next post, I will explain why this discussion is directly relevant to the Ramulator-gate saga.

Stay tuned.

DDR5 DIMMs: Channel vs. sub-channel terminology map

Memory companies

Rambus and Micron use different terminology in different documents

Samsung, SK Hynix and Kingston consistently use one-DIMM/one-channel terminology

System view: CPU + DDR5 DIMMs

All CPU manufacturers consistently use one-DIMM/one-channel terminology

Hundreds of sources for:

Scientific studies

A vast majority of studies follows one-DIMM/one-channel terminology

Evidence library

DDR5 DIMMs: Channel vs. sub-channel terminology map

JESD401-5C

"Sub-channel: The data across the width of the DDR5 standard DIMMs is divided into two independent sub-channels, each with half the data width of the full DDR5 channel."

JESD79-5C: DDR5 SDRAM

"To achieve efficient module power supply design for JEDEC-standard DDR5 DIMMs, minimum timings as well as limitations in the number of DRAMs are provided for Refresh, and Write operations occurring on a single module. As well, since these modules are organized as two independent 36-bit or 40-bit channels (32 bits for non-ECC DIMMs), additional restrictions apply in order to limit localized power delivery noise on the module"

Mess and Ramulator views

It is OK to disagree. The question is what you do when you disagree.

In the previous post, we discussed DDR5 DIMM channel and sub-channel terminology, and the inconsistencies that appear across different sources.

The Mess paper followed the commonly used terminology: a DDR5 DIMM is connected to the CPU through a single DDR5 channel, which is partitioned into two sub-channels. This is also how Intel and Amazon describe the memory organization of the servers that we benchmarked.

Ramulator 2.0, however, assumes two channels per DDR5 DIMM. Although this is not the mainstream use of the term “channel”, this convention does appear in some documents, as discussed in the previous post. The Ramulator authors also expressed their view that channel and sub-channel are not two distinct concepts in the DDR5 context. We still find this interpretation difficult to reconcile with the terminology used by the other sources we have found.

This difference between the mainstream understanding and the Ramulator 2.0 convention explains why “Number of channels: 16” is needed to simulate actual 8-channel DDR5 systems.

It took us 10 months of email exchange with the Ramulator team to detect this issue.

Some people argue that detecting it should have been obvious.

We would argue that it is obvious only in hindsight.

This issue could have been solved with a clarifying comment on “DDR5 channel” in the Ramulator 2.0 configuration or README files.

This was our proposal for an easy way forward.

The Ramulator team, however, took a different path.

Despite knowing that our DDR5-channel reasoning followed a mainstream view in academia and industry, the Ramulator team publicly presented our channel configuration as a “trivial human error” and a “gross configuration error”.

The most striking part is that the Ramulator team also states that “the misconfiguration of the number of channels in DDR5 could have been easily resolved by emails and collegial discussions.”

Indeed.

That is exactly what we tried to do.

For 10 months.

And this is precisely where the record matters more than the narrative.

Mess / BSC Memory Team

ONE DIMM: ONE CHANNEL,
TWO SUB-CHANNELS
DDR5 DIMM diagram labelled as one channel with two sub-channels

Ramulator / ETH Zurich Ramulator Team

ONE DIMM: TWO CHANNELS
DDR5 DIMM diagram labelled as two channels
In addition: DDR5 channel and subchannel are not two distinct concepts

We started the email communication with the Ramulator team in October 2024, even before the Mess paper was published. After ten months of email exchange, we found the reason for the reported low memory bandwidth of Ramulator2.

1 27 Aug 2025 BSC Memory team → Ramulator team

We informed the Ramulator team that we found the reason for low Ramulator2 bandwidth reported in the Mess paper:

Ramulator2 does not differentiate between DDR5 channel and sub-channel. In the configuration file, setting the number of channels actually sets the number of sub-channels.

The Ramulator team's reply was unexpected.

2 12 Sep 2025 Ramulator team → BSC Memory team

The Ramulator team replied that, in their view, DDR5 channel and subchannel are not two distinct concepts.

BSC Memory team presented its position on DDR5 channel/sub-channel terminology and proposed a constructive way to resolve the issue.

3 16 Oct 2025 BSC Memory team → Ramulator team
Terminology position

We explained to the Ramulator team that in the context of the HPC servers with DDR5 DIMMs, which are the systems tested in the Mess paper, the community agrees that the memory channel is the interface between a DIMM and CPU memory controller, and that each DDR5 DIMM channel is divided into two independent sub-channels.

Constructive way forward

We also suggested that a straightforward way to ensure that all Ramulator2 users interpret the term “DDR5 channel” consistently with its intended meaning would be to include a clarifying comment in the simulator configuration or README files. We concluded that this would be a very nice outcome of this discussion that would improve the Ramulator2 setup.

These statements are limited to restoring a verifiable fact regarding the existence of this email discussion prior to the publication of the “Cleaning up the Mess” paper, with the aim of preserving a respectful, rigorous, and community-useful environment for scientific discussion.

Out of respect for the confidentiality of communications, we will not publish the full emails or their attachments, recipient lists, email addresses, literal excerpts, or screenshots of the thread that would allow the full content of the email exchanges between the two teams to be reconstructed. We have limited ourselves to minimal evidence consisting of communication dates, with the sender and recipient fields and the main messages exchanged redacted, while keeping the remaining information confidential.

In addition, we have also raised this matter through the internal channels made available by the organizations involved.

Ramulator-Gate paper · 17 Oct 2025

Their response

The question is what to do when you disagree …

View paper ↗

Indeed.
That is exactly what we tried to do.
For 10 months.

Coming next

Next chapterRamulator configuration issues

The full response is written. We are making it public one claim at a time, each documented to the same standard of evidence. Check back for the next one.