20 Aug
|
Firmus Technologies
|
New South Wales
20 Aug
Firmus Technologies
New South Wales
Job Description
Firmus is seeking a highly skilled and driven Senior Engineer to play a key role in designing, building, and operating software-defined infrastructure, including high-performance AI storage platforms. You will help evolve our Software Defined Infrastructure by building reliable, scalable solutions that power some of the world's largest and most innovative AI workloads.
n
You will be instrumental in ensuring the stability, performance, and continuous improvement of our mission-critical control plane and storage infrastructure.
n
Key Responsibilities
n
n
Design and implement a highly scalable, multi-tenant control plane that supports Firmus' growing AI and infrastructure needs.
n
Contribute to the development of exabyte-scale, S3-compatible object storage, distributed file systems, and high-performance filesystems.
n
Work with bare-metal provisioning tools such as Base Command Manager, Warewulf, Ironic, MaaS, and similar platforms.
n
Apply a deep understanding of operating systems, computer networks, software-defined storage, and high-performance applications.
n
Work with technologies including RDMA, GPU Direct Storage, RoCE, InfiniBand, DPDK, Ceph, Weka, DAOS, and others.
n
Collaborate with operations teams to monitor, analyse, and optimise internal clusters and storage platforms.
n
Document architecture designs, operational procedures, and performance results.
n
Collaborate with L2 SRE engineers, site operations, and networking teams to ensure platform reliability, reproducibility, and performance.
n
Contribute to continuous improvement in cluster validation, CI/CD automation, and provisioning and testing frameworks.
n
Apply knowledge of Kubernetes and composable storage clusters.
n
Contribute to the development of custom Kubernetes operators and intelligent orchestration frameworks to optimise AI workload performance for large-scale GPU cluster commissioning.
n
n
Skills & Experience
n
n
Bachelor's or Master's degree in Computer Science, Engineering, or a related field. 6–10 years of experience in infrastructure engineering and/or storage engineering.
n
Hands-on experience with bare-metal provisioning. Ability to operate software-defined storage platforms such as Ceph, Weka, Vast Data, DAOS, or Lustre.
n
Solid understanding of cloud-native infrastructure, Kubernetes, and scalable system architectures.
n
Strong debugging and problem-solving skills in distributed, high-performance environments.
n
Solid automation mindset using tools such as Ansible, Helm,
Terraform/OpenTofu,
or equivalent.
n
Understanding of firmware, BIOS, BMC/IPMI/Redfish, and low-level system tuning.
n
Proficiency in one or more programming languages such as Go, Bash, Rust, or Python.
n
Excellent documentation skills with strong attention to detail.
n
Experience participating in an on-call rotation supporting production services.
Proactive self-starter with a drive for continuous technical improvement.
n
Systems Architecture: Ability to design and integrate virtualisation, bare-metal, GPU, storage, and Kubernetes/Slurm platforms.
n
Infrastructure Automation: Expertise in automated provisioning and lifecycle management of hardware and clusters.
n
GPU and HPC Performance: Strong understanding of GPU systems, RDMA fabrics, and distributed AI workload performance.
n
Technical Communication: Ability to communicate complex technical concepts effectively across engineering and operations teams.
n
Continuous Improvement: Demonstrates curiosity, proactive learning, and innovation in AI and HPC infrastructure.
n
n
Success Metrics
n
n
Reliable provisioning and benchmarking of scalable, high-performance storage systems.
n
Performance validation and optimisation.
n
Operational efficiency improvements.
n
High-quality documentation and effective knowledge transfer.
n
n
Location & Reporting
n
n
Singapore or Australia (Melbourne, VIC or Sydney, NSW or Launceston, TAS)
n
Reporting to Senior Manager, Software Defined Infrastructure
n
n
Employment Basis
n
Full-time
n
Diversity
n
At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.
n
Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering.
#J-*****-Ljbffr
📌 Senior Ai Infrastructure Engineer (Virtualisation) (New South Wales)
🏢 Firmus Technologies
📍 New South Wales