The Data Management System Development Unit conducts research and development about a large-scale data integration system on RIKEN Data Science Infrastructure. Also, it provides user support for the system. We will push forward with Open Science by development and research about: How to manage research data from the perspective of middleware technique and operation policy and how to achieve the researcher-friendly data collection and distribution methods.
Unit Overview
Main Research Field
Informatics
Related Research Fields
Research Keywords
Unit Leader Interview
Interview with the Unit Leader
~Creating a Foundation for Autonomous, Circular, Data-Driven Research~
Beyond Storage: Powering the Loop of Data-Driven Research
Currently, Research Data Management (RDM) at RIKEN stands at a historic turning point. RIKEN has cultivated a "library-style" RDM—an approach centered on categorizing and storing collected results according to rigorous rules so far. While this sophisticated classification system remains vital, we are now facing a paradigm shift toward "data-driven" research, where the exploding volume of data is transformed into a direct source of scientific discovery.
In this context, our unit’s objective is to further evolve the traditional role of a data repository and draw the blueprint for a "major hub for data-driven research"—a place where data circulates autonomously and continuously generates new knowledge.
Data-driven research consists of three phases: the "collection phase," where experimental and observational data are aggregated and made shareable; the "computation phase," where data is processed at high speeds using resources like HPC; and the "interpretation phase," where researchers analyze and sublimate these results into scientific insight. These are not independent stages; they form a single, systematically integrated loop. To keep this loop spinning smoothly, it is essential to combine traditional data management expertise with HPC insights that treat computational resources and data movement as a unified whole.
The Fusion of Library Science and HPC: "Full-Stack Perspective"
HPC experts possess a "full-stack" perspective, overseeing systems from the top layer to the bottom—ranging from network and storage configurations to software and the behavior of the researchers who are the end users. I believe that this comprehensive expertise will become the core of complex research data management infrastructures.
Until now, HPC and library-style data management have evolved along separate paths. Today, however, both fields are approaching the same common challenge—the efficient handling of massive data—from different angles. We aim to establish next-generation data management as a field of systems research by fusing the two: the steadfast expertise of library science in "accurately classifying and structuring information" and the HPC approach of "comprehensively capturing everything from infrastructure to user behavior."
Research Data Management: Four Innovations and the Power of AI
Modern research data is exploding in volume, making it increasingly difficult to keep up using only traditional, manual classification methods. Furthermore, data have to meet the FAIR principles—the requirement that it is "always findable there." However, current infrastructure faces challenges with sustainability, as it often depends on annual budgets and project-based cycles. The key to solving this lies in the "separation of location and ID." By ensuring that data remains permanently accessible through a persistent identifier (such as a DOI) regardless of its physical storage location, we can guarantee long-term data reliability.
Including achieving this, we have set four innovative goals for building our research data management infrastructure. The first goal is the "dynamic placement of data." By optimizing the allocation between shared and local data and abstracting the storage hierarchy, we aim to achieve both global data sharing and local high-speed analysis simultaneously.
The second goal is "automated metadata annotation." In this system, AI automatically extracts features from data to generate search indexes. By fostering cooperation between AI and expert curators, we can automatically structure the "meaning" of massive datasets. This approach combines the precise classification methods pioneered by library science with the "probabilistic search" of AI—which allows for the discovery of target data at an overwhelming scale, even when tolerating some noise. By bringing flexibility and scalability to data management, we aim to liberate researchers from administrative burdens, allowing them to return to the essential "interpretation phase."
The third goal is the "programmability of research workflows." By describing the entire sequence of actions—from data movement to analysis—as reusable code, we make it significantly easier to conduct cross-disciplinary scientific validation.
And the fourth goal is "integration with cutting-edge technologies." We ensure a flexible scalability that allows us to immediately incorporate the latest computing resources—such as AI models or quantum computers that evolve drastically every three months. Our unit’s ultimate mission is not simply to line up RIKEN’s world-class supercomputers, quantum computers, and advanced AI side-by-side, but to integrate them into a single, massive data-driven ecosystem.
Ideal candidate profile: New professional linking systems and science
To achieve this vision, we are looking for talents who can offer a perspective that spans multiple layers—from network and hardware to software and the specific requirements of various research fields. Of course, you don’t need to be an expert in every layer from day one. It is enough to have a core strength in one area while maintaining the curiosity to expand your focus into adjacent fields, growing step-by-step through daily operations and design feedback. We also heartily welcome those who have experience in data analysis in a specific research domain, who wish to transition into system development.
In academia, there is a persistent challenge where the achievements in system operation and construction are not always directly reflected in traditional paper-based evaluations. However, it is entirely possible to establish a global presence through the development of infrastructure for large-scale international collaborations and through participation in standardization activities. We aim to establish the role of the "research system administrator"—a highly specialized position that combines system management with active research, a career path already well-recognized abroad—here in Japan. We look forward to welcoming those who can find joy in building the next-generation infrastructure for the advancement of science as a whole.
Hideyuki Jitsumoto (Ph.D.)
Unit Leader
Specialized in computer systems. Earned Ph.D. in Science from the Department of Mathematical and Computing Science, Graduate School of Information Science and Engineering at the Tokyo Institute of Technology (now Institute of Science Tokyo). After serving as an Assistant Professor at the Global Scientific Information and Computing Center (GSIC) at the Tokyo Institute of Technology and in other academic roles, he has held the current position since 2020.
Selected Publications
- *Hideyuki Jitsumoto.: "System log for Resilience from our experience in TSUBAME2.5" The International Conference for High Performance Computing, Networking, Storage and Analysis (SC17). Nov 2017. (BoF Session: Characterizing Faults, Errors, and Failures in Extreme-Scale Systems),
- *Hideyuki Jitsumoto, Yuya Kobayashi, Akihiro Nomura, and Satoshi Matsuoka.: "MH-QEMU: Memory-State-Aware Fault Injection Platform" In Supercomputing Frontiers Asia (SCFA), March 2019
Members
| Position | Name |
|---|---|
| Unit Leader | Hideyuki Jitsumoto |
| Senior Technical Scientist | Shinji Kikuchi |
| Technical Scientist | Hiroo Hayashi |