Role Title: HPC Platform & Slurm Lead Engineer
Location: Sydney, NSW (Hybrid)
Start Date: ASAP
Duration: 12-Month Contract (with potential permanent conversion)
Working Flexibility: Hybrid
NTT DATA is seeking a highly skilled HPC Platform & Slurm Lead Engineer to lead the design, deployment, optimisation and operational management of a large-scale High Performance Computing (HPC) environment. This role will play a critical part in delivering a production-grade research computing platform supporting AI, machine learning, data analytics and advanced scientific workloads. The successful candidate will provide hands-on technical leadership across Slurm administration, Linux engineering, GPU infrastructure, high-performance storage, automation and platform operations.
Is innovation part of your DNA? Do you want to enable a connected future for people, organizations, and society?
Join our growing global NTT team and you’ll be part of the world’s largest ICT company (by revenue). We’ve combined the capabilities of 28 remarkable companies to become one, leading technology services provider. Together, we help our people, clients, and communities do great things with technology to create a more secure and connected future.
We employ 40,000 people across 57 countries. By bringing together the world’s best technology companies and emerging innovators, we work together to deliver sustainable outcomes to businesses and the world. Innovation is part of our DNA.
We believe it’s key to what makes us different. So, we strive to move forward, challenge the status quo, and drive excellence through the technologies we integrate and the services we deliver around the world. The result is connected cities, connected factories, connected healthcare,
connected agriculture, connected conservation, connected mobility, and connected sport.
Together we enable the connected future.
Key Responsibilities
- Design, deploy and support enterprise-scale HPC environments using Slurm Workload Manager.
- Architect and administer highly available Slurm controller infrastructure, scheduling policies, QoS frameworks and fairshare models.
- Build, configure and maintain Linux-based compute, login and management nodes.
- Deploy and support NVIDIA GPU platforms, including CUDA, NCCL, DCGM and GPU scheduling capabilities.
- Assist research and engineering teams with workload onboarding, performance optimisation and resource utilisation.
- Implement automation, monitoring and observability solutions using tools such as Grafana and Prometheus.
- Drive platform reliability, security, patching, operational processes, reporting and continuous improvement initiatives.
Skills & Experience
- Strong experience designing, implementing and supporting HPC platforms in production environments.
- Deep expertise administering Slurm Workload Manager, SlurmDBD, accounting, reservations, partitions, QoS and scheduler optimisation.
- Advanced Linux systems administration skills across RHEL, Rocky Linux or AlmaLinux environments.
- Experience supporting NVIDIA GPU infrastructure, AI/ML workloads and large-scale compute environments.
- Knowledge of HPC storage, container technologies, high-performance networking and cluster management solutions.
- Robust scripting, automation and infrastructure lifecycle management capabilities.
- Experience within research, university, scientific computing or large enterprise environments will be highly regarded.
Questions? Reach out:
[email protected]
📌 HPC Platform & Slurm Lead Engineer (Sydney)
🏢 Ntt Data
📍 Sydney