20 Aug
|
Firmus Technologies
|
New South Wales
20 Aug
Firmus Technologies
New South Wales
Job Description
Firmus is seeking a highly skilled and driven Kubernetes HPC Engineer to join our Software Defined Infrastructure team. In this role, you will build high-performance, fault-tolerant, and reliable infrastructure to support bare-metal provisioning, performance benchmarking, and platform validation.
n
You will be instrumental in ensuring the stability, performance, and continuous improvement of our complex and mission-critical bare-metal HPC GPU clusters.
n
Key Responsibilities
n
n
Design and implement bare-metal provisioning workflows using Ironic and Kubernetes CRDs.
n
Deploy and manage GPU-enabled AI compute nodes with RDMA, InfiniBand, and RoCE networking.
n
Optimise Kubernetes and Slurm platforms for multi-node AI training performance, including NCCL, UCX, GPUDirect, and fabric tuning.
n
Implement Kubernetes primitives for GPU scheduling, isolation, and resource management models.
n
Design, deploy, and fine‐tune Slurm GPU clusters with topology‐aware configurations.
n
Develop and execute performance benchmarking workloads, including MLPerf, NCCL tests, microbenchmarks, and throughput/latency validation.
n
Establish observability across GPU, InfiniBand fabric, storage, and provisioning components.
n
Document architecture designs, operational procedures, and performance results.
n
Collaborate with L2 SRE engineers, site operations, and networking teams to ensure platform reliability, reproducibility, and performance.
n
Support hardware bring‐up activities, including BIOS tuning, GPU topology verification, NUMA alignment, and PCIe/NVLink checks.
n
Contribute to continuous improvement in cluster validation, CI/CD automation,
and provisioning and testing frameworks.
n
Contribute to the development of custom Kubernetes operators and intelligent orchestration frameworks that optimise AI workload performance for large‐scale GPU cluster commissioning.
n
n
Skills & Experience
n
n
Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
n
Experience with bare‐metal cluster provisioning using tools such as Metal3, OpenStack Ironic, MaaS, xCAT, or similar.
n
Deep knowledge of Kubernetes internals, including CRDs, controllers, operators, and cluster lifecycle management.
n
Strong understanding of Slurm configuration and compiling AI and HPC applications.
n
Robust understanding of GPU systems (NVIDIA H100/H200 SXM platforms), CUDA/NCCL, and GPU topology (NVLink, NVSwitch, PCIe).
n
Familiarity with container runtimes for compute workloads, including Docker, Enroot, Singularity, and Podman.
n
Experience with benchmarking and performance validation for AI, HPC, or distributed training workloads.
n
Practical Linux systems engineering experience, including kernel, cgroups, system services, networking, and drivers.
n
Strong automation mindset using tools such as Ansible, Helm,
Terraform/OpenTofu,
or equivalent.
n
Understanding of firmware, BIOS, BMC/IPMI/Redfish, and low‐level system tuning.
n
Proficiency in one or more programming languages such as Go, Bash, Rust, or Python.
n
Excellent documentation skills with strong attention to detail.
n
Experience participating in an on‐call rotation supporting production services.
n
Proactive self‐starter with a drive for continuous technical improvement.
n
Systems Architecture: Ability to design and integrate bare‐metal, GPU, RDMA, and Kubernetes/Slurm platforms.
n
Infrastructure Automation: Skilled in automated provisioning and lifecycle management of hardware and clusters.
n
GPU and HPC Performance: Understanding of GPU systems, RDMA fabrics, and distributed AI workload performance.
n
Technical Communication: Ability to communicate technical concepts effectively across diverse engineering and operations teams.
n
Continuous Improvement: Demonstrates curiosity, proactive learning, and innovation in AI and HPC infrastructure.
n
n
Success Metrics
n
n
Reliable provisioning of Kubernetes and Slurm AI clusters.
n
Performance validation and optimisation.
n
Improved operational efficiency.
n
High‐quality documentation and effective knowledge transfer.
n
n
Location & Reporting
n
n
Australia (Sydney, NSW or Launceston, TAS)
n
Reporting to Senior Manager, Software Defined Infrastructure
n
n
Employment Basis
n
Full‐time
n
Diversity
n
At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.
#J-*****-Ljbffr
📌 Senior Hpc Infrastructure Engineer (New South Wales)
🏢 Firmus Technologies
📍 New South Wales