14 Sep
|
Capital Executive Search
|
Sydney
14 Sep
Capital Executive Search
Sydney
We are partnering with a rapidly scaling AI infrastructure and cloud computing business to appoint a Senior Manager, Storage Engineering to its Australian technology organisation. This is a hands‑on technical leadership role responsible for the design, delivery and operational excellence of global storage infrastructure supporting some of the world's most demanding AI and high‑performance computing workloads.
Operating across petabyte‑scale environments, you will own and evolve storage platforms spanning high‑performance parallel file systems, scale‑out NAS and object storage, supporting NVIDIA DGX SuperPOD environments and next‑generation GPU infrastructure.
This role will suit a senior storage engineer or technical leader who wants to remain deeply hands‑on while influencing architecture, platform strategy and infrastructure investment as the organisation continues to rapidly scale its global AI infrastructure footprint.
Key Responsibilities
- Drive the storage platform roadmap supporting the continued expansion of large‑scale GPU clusters and AI infrastructure across NVIDIA, AMD and next‑generation compute architectures.
- Architect, deploy and optimise high‑performance storage platforms across technologies including Dell PowerScale, WEKA, VAST Data and other distributed storage technologies.
- Design storage architectures capable of supporting extreme AI/ML workloads, including distributed training, checkpoint I/O, dataset staging and inference workloads operating at hundreds of GB/s.
- Own operational excellence across storage infrastructure, including availability, lifecycle management, replication, snapshots,
ransomware resilience and RPO/RTO requirements.
- Develop strong observability across capacity, latency, throughput and overall storage health.
- Lead complex troubleshooting, root‑cause analysis and performance optimisation across distributed storage environments.
- Optimise high‑bandwidth GPU‑to‑storage data paths leveraging technologies including RDMA, InfiniBand, RoCEv2, NVMe-oF and GPUDirect Storage.
- Partner closely with AI Platform Engineering, Cloud Infrastructure, Networking, Operations and Security teams to integrate storage into broader AI platform and GPU cluster architecture.
- Provide technical leadership across major storage initiatives while mentoring and developing engineers within the broader infrastructure organisation.
About You
You will bring deep technical expertise in high‑performance storage and distributed infrastructure, ideally gained within AI/ML, HPC, cloud or large‑scale data centre environments.
You will ideally have:
- 7+ years' experience across storage engineering, infrastructure engineering or distributed systems.
- 2+ years' experience operating in a senior engineering, technical leadership or team leadership capacity.
- Deep hands‑on expertise with at least two of Dell PowerScale/OneFS, Pure Storage FlashBlade, WEKA Data Platform or VAST Data.
- Strong knowledge of distributed and parallel storage architectures, including striping, metadata architecture, erasure coding and consistency models.
- Experience designing and operating petabyte‑scale storage environments supporting demanding or multi‑tenant workloads.
- Practical experience with RDMA technologies including InfiniBand and RoCEv2, particularly across NVMe‑oF or GPU‑centric environments.
- Strong understanding of GPUDirect Storage and NVIDIA DGX/HGX infrastructure.
- Experience optimising storage for AI/ML workloads, including checkpointing, dataset sharding and distributed data pipelines.
- Strong understanding of object storage architectures, S3‑compatible APIs, lifecycle management and storage tiering.
Why Consider This Opportunity?
This is an prospect to take a senior technical role within an organisation making significant investments in next‑generation AI infrastructure globally. Rather than maintaining traditional enterprise storage, you will be solving storage challenges created by large‑scale GPU computing, distributed AI training and rapidly growing datasets, with the opportunity to influence architecture and technology decisions as the platform scales. You will work alongside specialist teams across AI infrastructure, cloud, networking and operations while remaining close to the technology and helping define how storage infrastructure is designed and operated at scale.
#J-18808-Ljbffr
📌 Storage Engineer Senior Manager (GPU/AI Infrastructure) (Sydney)
🏢 Capital Executive Search
📍 Sydney