Senior Platform Engineer, Monitoring & Telemetry (New South Wales)

Senior Platform Engineer, Monitoring & Telemetry (New South Wales)

26 Sep
|
SQC Silicon Quatum Computing
|
New South Wales

26 Sep

SQC Silicon Quatum Computing

New South Wales

SQC Silicon Quatum Computing - Sydney NSW
2d ago , from SQC Silicon Quatum Computing
Silicon Quantum Computing (SQC) is at the forefront of global efforts to build the world's first commercial-scale quantum computer, while delivering quantum-enhanced AI and simulation products to customers today.
Backed by over 25 years of technological excellence, SQC is a full-stack quantum computing company that leverages its proprietary manufacturing process to engineer atomic qubits in silicon with 0.13 nanometer precision. It is the most precise semiconductor manufacturing in the world, enabling systems with world-leading algorithmic fidelity, and a decisive advantage in the global quantum computing race.
Our products are commercially deployed and generating revenue. Watermelon, our quantum-enhanced AI system, is delivering superior results on real-world problems across energy, telecom and finance. Quantum Twins, our simulation platform, provides unparalleled ability to model quantum systems, accelerating molecule and materials discovery.
This is SQC: building the future of computing while delivering quantum impact today.
Silicon Quantum Computing (SQC) is at the forefront of global efforts to build the world's first commercial-scale quantum computer, while delivering quantum-enhanced AI and simulation products to customers today.
Backed by over 25 years of technological excellence, SQC is a full-stack quantum computing company that leverages its proprietary manufacturing process to engineer atomic qubits in silicon with 0.13 nanometer precision. It is the most precise semiconductor manufacturing in the world, enabling systems with world-leading algorithmic fidelity, and a decisive advantage in the global quantum computing race.
Our products are commercially deployed and generating revenue. Watermelon, our quantum-enhanced AI system, is delivering superior results on real-world problems across energy, telecom and finance. Quantum Twins, our simulation platform, provides unparalleled ability to model quantum systems, accelerating molecule and materials discovery.
This is SQC: building the future of computing while delivering quantum impact today.
About the role




We are hiring a Platform Engineer to own monitoring and telemetry in Platform & Infrastructure. The scope is everything the team runs: the Kubernetes clusters from small local sites to the central cluster, Ceph storage, core network services, the HPC and simulation environments, the quantum runtime cluster and its FPGA mesh, our AWS footprint, and the cluster that ships with each quantum computer.
The workloads are not a web estate. A FPGA synthesis run takes hours and holds a licence while it does. A simulation job may need to be reproducible years later. The realtime control plane cares about microsecond jitter rather than request percentiles. A physicist will want to correlate a device measurement with a calibration run and a cluster event, so telemetry from unrelated systems has to be joinable.
Compute is on-premise and finite. Utilisation numbers decide where the next hardware spend goes and which team is under-served, so they end up in procurement and scheduling decisions. You own the platform and set the practice: the conventions other teams instrument against, and alerting an on-call engineer acts on without checking it twice.
Based at our Sydney facility, you will work alongside the platform and infrastructure engineers who run the estate, and with the research teams instrumenting their own work. This is a role for someone who wants observability to carry real decisions, and who would rather retire a noisy alert than tune it out.
Role responsibilities
Own the monitoring and telemetry platform end to end: metrics, logs, traces and alerting across clusters, storage, network and cloud
Instrument the platform services the team runs, including Kubernetes, Ceph, the CI/CD platform, GitHub Enterprise, Artifactory and Vault
Make utilisation and capacity visible and trustworthy, so scheduling decisions and hardware purchases rest on measurement




Design alerting that names a specific action and keeps noise low, and retire alerts that no longer prompt one
Support the quantum runtime cluster's observability needs, including high-rate telemetry from the FPGA mesh, without disturbing the realtime path
Extend monitoring to every cluster we build, one per quantum computer, including air-gapped sites where telemetry cannot leave the building and has to be practical to whoever is standing next to it
Handle the data engineering of observability at scale: cardinality, sampling, retention and the cost of keeping it
Define and measure service levels for internal platform services, and report against them
Build the on-call tooling and runbooks, and improve them after every incident
Help other teams instrument their own systems, and provide the libraries, conventions and defaults that make that easy
Correlate telemetry across domains, so a research question spanning device, cluster and job data can be answered
Support incident response and post-incident review across the platform, and document the platform, its conventions and its limits, so observability does not depend on knowledge held by one person
Your experience
Essential
5+ years in platform, infrastructure or site reliability engineering, with direct ownership of an observability stack
Prometheus and Grafana at scale, including long-term storage with Thanos, Mimir, VictoriaMetrics or an equivalent
Log pipelines in production: Loki, OpenSearch or ELK, with a shipper such as Vector or Fluent Bit
OpenTelemetry, and distributed tracing in a real system rather than a demo
Alerting and on-call design, including direct on-call responsibility for systems you instrumented
Production Kubernetes depth: workloads, networking, storage, RBAC, operators and the failure modes of each
Strong Linux systems administration, with Python and Bash
Infrastructure-as-code and GitOps: Terraform or OpenTofu, with Argo CD or Flux
High-cardinality time series in practice, and the retention and cost decisions that come with it
Capacity planning on finite on-premise hardware
Clear technical writing, and the ability to support researchers and engineers instrument
#J-*****-Ljbffr

📌 Senior Platform Engineer, Monitoring & Telemetry (New South Wales)
🏢 SQC Silicon Quatum Computing
📍 New South Wales

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior platform engineer, monitoring & telemetry (new south wales) / new south wales

Subscribe to this job alert:

Get the latest job offers by email for: senior platform engineer, monitoring & telemetry (new south wales) / new south wales