03 Sep
|
Silicon Quantum Computing
|
Sydney
03 Sep
Silicon Quantum Computing
Sydney
Silicon Quantum Computing (SQC) is at the forefront of global efforts to build the world's first commercial-scale quantum computer, while delivering quantum-enhanced AI and simulation products to customers today.Backed by over 25 years of technological excellence, SQC is a full-stack quantum computing company that leverages its proprietary manufacturing process to engineer atomic qubits in silicon with 0.13 nanometer precision.
It is the most precise semiconductor manufacturing in the world, enabling systems with world-leading algorithmic fidelity, and a decisive advantage in the global quantum computing race.Our products are commercially deployed and generating revenue.
Watermelon, our quantum-enhanced AI system, is delivering superior results on real-world problems across energy, telecom and finance.
Quantum Twins, our simulation platform, provides unparalleled ability to model quantum systems, accelerating molecule and materials discovery.This is SQC: building the future of computing while delivering quantum impact today.About The RoleWe are hiring a Platform Engineer to own monitoring and telemetry in Platform & Infrastructure.
The scope is everything the team runs: the Kubernetes clusters from small local sites to the central cluster, Ceph storage, core network services, the HPC and simulation environments, the quantum runtime cluster and its FPGA mesh, our AWS footprint, and the cluster that ships with each quantum computer.Role responsibilitiesOwn the monitoring and telemetry platform end to end: metrics, logs, traces and alerting across clusters, storage, network and cloudInstrument the platform services the team runs, including Kubernetes, Ceph, the CI/CD platform, GitHub Enterprise, Artifactory and VaultMake utilisation and capacity visible and trustworthy, so scheduling decisions and hardware purchases rest on measurementDesign alerting that names a specific action and keeps noise low, and retire alerts that no longer prompt oneSupport the quantum runtime cluster's observability needs, including high-rate telemetry from the FPGA mesh, without disturbing the realtime pathExtend monitoring to every cluster we build, one per quantum computer,
including air-gapped sites where telemetry cannot leave the building and has to be practical to whoever is standing next to itHandle the data engineering of observability at scale: cardinality, sampling, retention and the cost of keeping itDefine and measure service levels for internal platform services, and report against themBuild the on-call tooling and runbooks, and improve them after every incidentHelp other teams instrument their own systems, and provide the libraries, conventions and defaults that make that easyCorrelate telemetry across domains, so a research question spanning device, cluster and job data can be answeredSupport incident response and post-incident review across the platform, and document the platform, its conventions and its limits, so observability does not depend on knowledge held by oneYour experienceEssential5+ years in platform, infrastructure or site reliability engineering, with direct ownership of an observability stackPrometheus and Grafana at scale, including long-term storage with Thanos, Mimir, VictoriaMetrics or an equivalentLog pipelines in production: Loki, OpenSearch or ELK, with a shipper such as Vector or Fluent BitOpenTelemetry, and distributed tracing in a real system rather than a demoAlerting and on-call design, including direct on-call responsibility for systems you instrumentedProduction Kubernetes depth: workloads, networking, storage, RBAC, operators and the failure modes of eachStrong Linux systems administration, with Python and BashInfrastructure-as-code and GitOps: Terraform or OpenTofu, with Argo CD or FluxHigh-cardinality time series in practice, and the retention and cost decisions that come with itCapacity planning on finite on-premise hardwareClear technical writing,
and the ability to support researchers and engineers instrumenting their own workNice to haveCeph monitoring and troubleshooting at the cluster levelNetwork telemetry: SNMP, flow data, switch and fabric monitoring, DHCP and DNS visibilityHPC and batch scheduler observability, including Slurm job accountingGPU telemetry, such as DCGM, and accelerator utilisation reportingeBPF-based tooling for host and kernel visibilityHardware and out-of-band monitoring: IPMI, Redfish, environmental and power telemetryRealtime or hard-deadline systems, where jitter matters more than throughputLaboratory or scientific instrument telemetry, and joining it to infrastructure dataMonitoring a fleet of similar installations rather than a single estate, including sites with no outbound networkAWS observability and cost visibilityIncident management practice, including running blameless reviewsGrafana dashboard and plugin development, or building internal observability toolingEqual opportunitySQC is an equal opportunity employer.
We value diverse perspectives and experiences, and encourage applications from candidates who may not meet every listed requirement.Export controlsThis position may require access to export-controlled information or technology.
Employment may be subject to applicable export control laws and may require eligibility assessment based on factors such as nationality, citizenship, or residency, and, where necessary, obtaining relevant export licenses or approvals.About SQCSQC was founded by renowned physicist and materials scientist Michelle Simmons, who pioneered the field of atomic electronics, including the development of the world's first single-atom transistor and the first integrated circuit built with atomic precision.
Our Chair, Simon Segars, former CEO of Arm, is a leader in the semiconductor industry and was instrumental in developing the processors that powered the mobile computing revolution.As a full-stack company with in-house QPU manufacturing, SQC can design, produce and test new quantum chips in under a week, enabling rapid iteration and a decisive advantage in the race to build the world's first commercial-scale quantum computer.
#J-*****-Ljbffr
📌 Senior Platform Engineer, Monitoring & Telemetry (Sydney)
🏢 Silicon Quantum Computing
📍 Sydney