14 Sep
|
Firmus Technologies
|
New South Wales
14 Sep
Firmus Technologies
New South Wales
Job Description
n Firmus Technologiesn
Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.
Firmus Technologiesn
Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.
n
Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.
n
At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure - the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.
Firmus AI Cloudn
Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.
n
It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.
Role Summaryn
The Senior Kubernetes Engineer, AI Infrastructure owns the technical design and delivery of the backend infrastructure that powers the Firmus Kubernetes platform. This is a hands-on principal-level individual contributor role, responsible for building production-grade cluster lifecycle, control-plane, networking, storage, security, observability, and automation capabilities across GPU-accelerated bare-metal environments.
n
They solve the hardest platform engineering problems, set Kubernetes engineering standards, and provide domain-level technical sign-off for platform designs. They work across AI Platforms, Solutions Architecture & Delivery, networking, security, and operations to create a secure, resilient, multi-tenant platform that can be deployed and operated consistently at AI-factory scale.
Key Responsibilitiesn
n
- Define and own the Kubernetes platform reference architecture across management and workload clusters, including control-plane topology, cluster lifecycle, multi-tenancy, workload isolation, and failure-domain design.
n
- Build and maintain the backend services, APIs, controllers, operators, and automation required to provision, configure, upgrade, scale, and retire Kubernetes clusters reliably.
n
- Engineer repeatable bare-metal Kubernetes deployment and lifecycle workflows using infrastructure-as-code and automated provisioning technologies such as Cluster API, kubeadm, Redfish, PXE, Ironic, or Metal3.
n
- Design and operate cluster networking across CNI, ingress, service discovery, DNS, load balancing, network policy, and service mesh; integrate Multus, SR-IOV, BGP, InfiniBand, or RoCE where required for high-performance AI workloads.
n
- Define persistent-storage and data-service patterns using CSI, Ceph, local NVMe, object storage, backup and restore, and disaster-recovery mechanisms appropriate for stateful platform and AI workloads.
n
- Integrate and productionise NVIDIA GPU and Network Operators, device plugins, drivers, DCGM telemetry, scheduling, quotas, and topology-aware placement for multi-node accelerated workloads.
n
- Establish GitOps and CI/CD patterns for platform software, configuration, policy, and release management, with safe testing, progressive rollout, rollback, and upgrade practices.
n
- Build platform security into the architecture through identity and access control, RBAC, secrets management, policy-as-code, image and software-supply-chain controls, tenant isolation, and auditable change management.
n
- Define service-level objectives and engineer observability for metrics, logs, traces, events, capacity, and performance; lead diagnosis of complex distributed systems failures and eliminate recurring operational toil.
n
- Set engineering standards, design patterns, review practices, and operational readiness criteria; mentor senior engineers and resolve cross-team technical decisions while remaining directly involved in implementation.
n
Skills & Experiencen n
- 7+ years of progressive infrastructure, systems, or platform engineering experience, including substantial ownership of production Kubernetes platforms and at least 3 years operating at senior staff, principal, or equivalent level.
n
- Deep knowledge of Kubernetes internals, including the API server, etcd, scheduler, controller manager, kubelet, admission, CRI, CNI, CSI, reconciliation patterns, cluster performance, upgrades, and control-plane failure modes.
n
- Demonstrated experience designing, building,
and operating highly available, large scale and multi-cluster Kubernetes platforms on bare metal, private cloud, or hybrid infrastructure.
n
- Strong software engineering ability in Go and/or Rust, with practical Python and Bash skills; experience building Kubernetes operators, controllers, admission webhooks, CLIs, or platform services.
n
- Expert Linux systems knowledge, including namespaces, cgroups, systemd, kernel, host networking and container runtime behaviour, performance analysis, and low-level troubleshooting.
n
- Strong Kubernetes networking expertise across Cilium, Calico, or equivalent CNI implementations, plus load balancing, DNS, ingress, BGP, network policy, and multi-network architectures.
n
- Strong infrastructure automation and GitOps experience with tools such as Terraform, Ansible, Argo CD, Flux, GitHub Actions, GitLab CI, or Jenkins.
n
- Practical experience with Kubernetes security and governance, including RBAC, OPA Gatekeeper or Kyverno, secrets management, certificate lifecycle, image security, and workload isolation.
n
- Experience implementing production observability with Prometheus, Grafana, OpenTelemetry, Loki, Elasticsearch, or equivalent technologies, and using telemetry to manage reliability, capacity, and performance.
n
- Experience with GPU-enabled Kubernetes infrastructure, NVIDIA GPU Operator, accelerator scheduling for AI workloads at large scale, RDMA networking, and distributed AI workload requirements.
n
- Experience with distributed storage and data services such as Ceph, CSI-backed storage, object storage, backup and restore, and disaster recovery.
n
- CKA-level expertise is expected; CKA, CKS, or relevant cloud-native certifications are strongly preferred.
n
- Bachelor’s degree in computer science, engineering, or a related discipline, or equivalent depth of practical engineering experience.
n
- Explicit technical judgement and communication, with a record of influencing architecture across software, networking, security, platform, and operations teams.
n
Location & Reportingn n
- Australia (Sydney, NSW or Launceston, TAS)
n
- Reporting to Head of AI Platform
n
Employment Basisn
Full-time
Diversityn
At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.
n
Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering.
n
#J-18808-Ljbffr
📌 Principal AI Infrastructure Engineer, Kubernetes (New South Wales)
🏢 Firmus Technologies
📍 New South Wales