Why an HPC workload can struggle before it starts moving much data and how repeated discovery and first-access latency contribute.
When an HPC job runs slowly, we often look first at storage throughput.
Are the network links saturated? Would faster SSDs help? Does the filesystem need more bandwidth?
Those are useful questions. But sometimes the application has barely started reading its data.
It is still finding files, checking attributes and opening them.
That preparation can become a bottleneck of its own.
When every worker asks the same questions
Imagine 1,000 worker processes starting together. Each scans the same directory tree and checks 10,000 files before choosing its portion of the dataset.
That produces ten million application-level file checks.
Not every check necessarily becomes a request to storage. Client caching and implementation details can reduce the actual traffic. Nevertheless, the application has repeated the same discovery work 1,000 times.
The data processing is distributed. The preparation is duplicated.
If those checks arrive together, they also create a burst of metadata demand. A filesystem may have plenty of bandwidth while clients wait for directory listings, path lookups or file attributes.
Moving bytes quickly and finding files quickly are different capabilities.
A related experience with NetApp FlexCache
I encountered a related issue with clients accessing NetApp FlexCache volumes.
Users complained that directory listings could take several seconds. Sometimes, they reported seeing no immediate output from an ls command.
We ran a process to traverse the relevant directories and access files ahead of subsequent use. In our case, warming the cache improved the client experience.
The lesson was that first access and repeated access can behave differently.
FlexCache caches content on demand rather than maintaining a complete replica of the origin. Initial access can therefore involve communication with the origin and fetching uncached content, adding latency depending on the environment.
NetApp provides a built-in prepopulation command for proactively warming selected paths:
volume flexcache prepopulate start -cache-vserver svm_cache /
-cache-volume cache_vol -path-list /dataset -isRecursion trueThe command requires advanced privilege. Replace the example SVM and volume names; /dataset is relative to the origin volume’s root.
Prepopulation crawls directories and reads files. It moves fetching work ahead of client access and is broader than a metadata-only scan.
It is an optional performance measure, rather than a requirement for every FlexCache deployment.
Improvement after warming does not, by itself, prove the exact cause of every delay. Also, an ls command waiting to display results is different from one that completes but incorrectly reports an empty directory. The latter warrants separate investigation.
Avoid repeated work—or move necessary work earlier
The HPC discovery example and the FlexCache experience suggest two different approaches.
First, avoid unnecessary repetition. Prepare a manifest—a list of input files—and give each worker its assigned entries instead of making every worker scan the entire dataset.
Workers still need to open their files, but they no longer need to rediscover the whole dataset independently.
Second, move appropriate preparation ahead of the workload. Where first-access latency matters, warming the required dataset before a large job starts can help.
Neither approach is free. A manifest must remain consistent with the dataset. Prepopulation generates traffic, places load on the origin and consumes cache capacity.
The aim is to prepare the relevant working set, rather than blindly scan or prepopulate everything.
Scale the right part of storage
Parallel filesystems can scale metadata services. BeeGFS distributes portions of its namespace across metadata servers, while suitable FSx for Lustre configurations allow metadata performance to scale independently of storage capacity.
However, more data targets or faster network links do not automatically solve concentrated directory activity.
Before adding hardware, separate the job into four phases:
Discovery → Opening → Data transfer → Output creation
Where is the time spent? Are workers repeating the same checks? Are requests arriving in bursts? Is activity concentrated in a few directories?
Then compare startup time and total job time after a targeted change—not just peak throughput.
The question before the bandwidth question
HPC performance depends on how quickly storage answers questions as well as how quickly it moves bytes.
Sometimes those questions are unavoidable. Sometimes the application asks the same ones thousands of times.
Before asking how to make storage faster, it is worth asking:
How much work does this application ask storage to do before useful processing begins?
-Ash
Sources and further reading
- NetApp: Prepopulate FlexCache volumes
- NetApp: FlexCache prepopulation command
- BeeGFS architecture
- FSx for Lustre metadata performance
Note: The worker-count example is illustrative; the FlexCache account reflects my own troubleshooting experience.
Comments
Post a Comment