Towards I/O Analysis of HPC Systems and a Generic Architecture to Collect Access Patterns

Marc C. Wiedemann; Julian M. Kunkel; Michaela Zimmer; Thomas Ludwig; Michael Resch; Thomas Bönisch; Xuan Wang; Andriy Chut; Alvaro Aguilera; Wolfgang E. Nagel; Michael Kluge; Holger Mickler

doi:10.1007/s00450-012-0221-5

Towards I/O Analysis of HPC Systems and a Generic Architecture to Collect Access Patterns

Research output: Contribution to journal › Research article › Contributed › peer-review

Contributors

Marc C. Wiedemann - , German Climate Computing Center (DKRZ) (Author)
Julian M. Kunkel - , German Climate Computing Center (DKRZ) (Author)
Michaela Zimmer - , German Climate Computing Center (DKRZ) (Author)
Thomas Ludwig - , German Climate Computing Center (DKRZ) (Author)
Michael Resch - , University of Stuttgart (Author)
Thomas Bönisch - , University of Stuttgart (Author)
Xuan Wang - , University of Stuttgart (Author)
Andriy Chut - , University of Stuttgart (Author)
Alvaro Aguilera - , Center for Information Services and High Performance Computing (ZIH) (Author)
Wolfgang E. Nagel - , Chair of Computer Architecture (Author)
Michael Kluge - , Center for Information Services and High Performance Computing (ZIH) (Author)
Holger Mickler - , Center for Information Services and High Performance Computing (ZIH) (Author)

Abstract

In high-performance computing applications, a high-level I/O call will trigger activities on a multitude of hardware components. These are massively parallel systems supported by huge storage systems and internal software layers. Their complex interplay currently makes it impossible to identify the causes for and the locations of I/O bottlenecks. Existing tools indicate when a bottleneck occurs but provide little guidance in identifying the cause or improving the situation.
We have thus initiated Scalable I/O for Extreme Performance to find solutions for this problem. To achieve this goal in SIOX, we will build a system to record access information on all layers and components, to recognize access patterns, and to characterize the I/O system. The system will ultimately be able to recognize the causes of the I/O bottlenecks and propose optimizations for the I/O middleware that can improve I/O performance, such as throughput rate and latency. Furthermore, the SIOX system will be able to support decision making while planning new I/O systems.
In this paper, we introduce the SIOX system and describe its current status: We first outline our approach for collecting the required access information. We then provide the architectural concept, the methods for reconstructing the I/O path and an excerpt of the interface for data collection. This paper focuses especially on the architecture, which collects and combines the relevant access information along the I/O path, and which is responsible for the efficient transfer of this information. An abstract modelling approach allows us to better understand the complexity of the analysis of the I/O activities on parallel computing systems, and an abstract interface allows us to adapt the SIOX system to various HPC file systems.

Details

Original language	English
Pages (from-to)	241–251
Number of pages	11
Journal	Software-Intensive Cyber-Physical Systems
Volume	28
Issue number	2–3
Publication status	Published - 1 May 2013
Peer-reviewed	Yes

External IDs

Scopus	84882904614
ORCID	/0000-0002-1686-8440/work/142240062

Keywords

I/O analysis, I/O path, Causality tree

Research Portal of the TU Dresden

Contributors

Abstract

Details

External IDs

Keywords

Keywords