| Home > In process > Bridging Grid and HPC Computing for Data-Intensive Science: Scaling dCache for Modern Workflows |
| Contribution to a conference proceedings/Journal Article | PUBDB-2026-01969 |
; ; ; ; ; ;
2026
Springer
Heidelberg
ISBN: 978-3-032-29911-6 (print), 978-3-032-29912-3 (electronic)
This record in other databases:
Please use a persistent id in citations: doi:10.1007/978-3-032-29912-3_40
Abstract: The increasing integration of high-performance computing (HPC) resources into data-intensive scientific workflows places new demands on storage systems traditionally developed for grid and distributed computing environments. dCache, a mature, exascale storage system jointly developed by Deutsches Elektronen-Synchrotron (DESY), Fermi National Accelerator Laboratory (FNAL), and the Nordic e-Infrastructure Collaboration (NeIC), has evolved to support a broad range of scientific communities beyond its origins in High-Energy Physics, including astrophysics, Photon science, and AI training. These communities increasingly rely on HPC systems for large-scale data analysis, exposing scalability and performance challenges at the metadata and data access layers.This paper presents recent development efforts to make dCache more HPC-friendly in interdisciplinary scientific environments. By aligning dCache’s architecture and development practices more closely with HPC requirements, we enable transparent access to shared data infrastructures from HPC clusters while preserving dCache’s strengths in data management, federation, and long-term preservation. We focus on optimizing metadata access to improve scalability and reduce latency under highly parallel workloads, as well as on substantial enhancements to dCache’s NFSv4.1/pNFS implementation. These include improved pNFS layout handling, read delegation, and zero-copy data paths to reduce CPU overhead and memory copies, resulting in significantly improved I/O performance for HPC applications. We evaluate these enhancements using representative HPC workloads at DESY and FNAL, assessing their impact on throughput, latency, and overall system scalability. The benchmark results demonstrate that the proposed changes can significantly improve application performance, depending on the workflow in use.
|
The record appears in these collections: |