Memory Performance of AMD EPYC Rome and Intel Cascade Lake SP Server Processors

Markus Velten; Robert Schöne; Thomas Ilsche; Daniel Hackenberg

doi:10.1145/3489525.3511689

Memory Performance of AMD EPYC Rome and Intel Cascade Lake SP Server Processors

Research output: Contribution to book/Conference proceedings/Anthology/Report › Conference contribution › Contributed › peer-review

Contributors

Markus Velten - , Center for Information Services and High Performance Computing (ZIH), Chair of Computer Architecture (Author)
Robert Schöne - , Chair of Computer Architecture, Center for Information Services and High Performance Computing (ZIH) (Author)
Thomas Ilsche - , Center for Information Services and High Performance Computing (ZIH) (Author)
Daniel Hackenberg - , Center for Information Services and High Performance Computing (ZIH) (Author)

Abstract

Modern processors, in particular within the server segment, integrate more cores with each generation. This increases their complexity in general, and that of the memory hierarchy in particular. Software executed on such processors can suffer from performance degradation when data is distributed disadvantageously over the available resources. To optimize data placement and access patterns, an in-depth analysis of the processor design and its implications for performance is necessary. This paper describes and experimentally evaluates the memory hierarchy of AMD EPYC Rome and Intel Xeon Cascade Lake SP server processors in detail. Their distinct microarchitectures cause different performance patterns for memory latencies, in particular for remote cache accesses. Our findings illustrate the complex NUMA properties and how data placement and cache coherence states impact access latencies to local and remote locations. This paper also compares theoretical and effective bandwidths for accessing data at the different memory levels and main memory bandwidth saturation at reduced core counts. The presented insight is a foundation for modeling performance of the given microarchitectures, which enables practical performance engineering of complex applications. Moreover, security research on side-channel attacks can also leverage the presented findings.

Details

Original language	English
Title of host publication	ICPE 2022 - Proceedings of the 2022 ACM/SPEC International Conference on Performance Engineering
Pages	165–175
Number of pages	11
ISBN (electronic)	9781450391436
Publication status	Published - 9 Apr 2022
Peer-reviewed	Yes

External IDs

Scopus	85128652028
Mendeley	ff193789-b072-3738-a096-7141bda87960
dblp	conf/wosp/VeltenSIH22
WOS	000883411400019
ORCID	/0000-0002-8491-770X/work/141543305
ORCID	/0000-0002-2730-0308/work/142245960
ORCID	/0009-0003-0666-4166/work/151475605
ORCID	/0000-0002-5437-3887/work/154740524

Keywords

ASJC Scopus subject areas

Keywords

amd zen 2, cache coherence, amd epyc rome, memory hierarchy, intel xeon skylake, intel xeon cascade lake, AMD EPYC Rome, Intel Xeon Cascade Lake, AMD Zen 2, Intel Xeon Skylake

Research Portal of the TU Dresden

Contributors

Abstract

Details

External IDs

Keywords

ASJC Scopus subject areas

Keywords