Modern processors are becoming increasingly faster, while memory speed
and bandwidth continue to lag behind, creating the so-called memory wall.
AMD addresses this problem with 3D V-Cache technology, which increases L3
cache capacity by vertically stacking additional silicon layers above the cache.
This thesis investigates the impact of this technology on the performance
of three computationally intensive algorithms with different cache access
patterns: numerical solution of the two-dimensional Laplace equation using
Jacobi iteration, PageRank for graph processing, and heapsort. The C++
implementations were evaluated on two high-performance computing clusters:
FRIDA, with an AMD EPYC 9684X processor, and Arnes HPC, with an
AMD EPYC 9534 processor. These systems were selected because of their
comparable architectures and membership in the same processor family. The
results show that the benefit of a larger L3 cache depends mainly on memory
access patterns and working-set size. The 3D V-Cache processor outperformed
the other system for all three algorithms, with the largest gains for problem
sizes whose active data fit more effectively in cache. For very small or very
large working sets, the performance advantage was generally smaller.
|