Senior Platform Engineer, Monitoring & Telemetry

SQC Silicon Quatum Computing

Sydney

On-site

AUD 140,000 - 210,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

SQC Silicon Quatum Computing is hiring a Platform Engineer to own monitoring and telemetry across on-premise clusters, Ceph storage, and the FPGA-mesh quantum runtime. The role covers Kubernetes, CI/CD tooling, and Vault integration, with a focus on actionable alerts and minimal noise.

Based at our Sydney facility, you’ll collaborate with researchers and platform engineers to drive observability that informs hardware and procurement decisions.

Qualifications

  • 5+ years in platform, infrastructure or site reliability engineering with observability ownership.
  • Prometheus and Grafana at scale with long-term storage.
  • Log pipelines in production: Loki, OpenSearch or ELK with a shipper.
  • OpenTelemetry and distributed tracing in a real system.
  • Alerting and on-call design with ownership of instruments.
  • Production Kubernetes depth: workloads, networking, storage, RBAC, operators.
  • Strong Linux systems administration with Python and Bash.
  • Infrastructure-as-code and GitOps: Terraform or OpenTofu with Argo CD or Flux.
  • High-cardinality time series management and cost considerations.
  • Capacity planning on finite on-premise hardware.
  • Clear technical writing to support researchers and engineers instrument.

Responsibilities

  • Own the monitoring and telemetry platform end to end across clusters, storage, network and cloud.
  • Instrument Kubernetes, Ceph, CI/CD, GitHub Enterprise, Artifactory and Vault services.
  • Make utilisation visible to guide scheduling and hardware purchases.
  • Design alerting that triggers concrete actions and keeps noise low.
  • Support quantum runtime observability, including FPGA mesh telemetry, without disturbing real-time path.
  • Extend monitoring to every cluster, including air-gapped sites.
  • Handle data engineering of observability at scale: cardinality, sampling, retention and cost.
  • Define and measure SLAs for internal platform services and report on them.
  • Build on-call tooling and runbooks, improving after incidents.
  • Help teams instrument their own systems and provide libraries and conventions.
  • Correlate telemetry across domains for cross-cutting research questions.
  • Support incident response and post-incident review and document platform.

Skills

Observability stack ownership
Prometheus/Grafana at scale
Log pipelines (Loki/OpenSearch/ELK)
OpenTelemetry
Alerting and on-call design
Production Kubernetes depth
Linux systems administration
Python/Bash
IaC and GitOps (Terraform/OpenTofu, Ar

Tools

Prometheus
Grafana
Thanos/Mimir/VictoriaMetrics
Loki/OpenSearch/ELK
Vector/Fluent Bit
Terraform/OpenTofu
Argo CD/Flux

Job description

SQC Silicon Quatum Computing - Sydney NSW

2d ago , from SQC Silicon Quatum Computing

Silicon Quantum Computing (SQC) is at the forefront of global efforts to build the world's first commercial-scale quantum computer, while delivering quantum-enhanced AI and simulation products to customers today.

Backed by over 25 years of technological excellence, SQC is a full-stack quantum computing company that leverages its proprietary manufacturing process to engineer atomic qubits in silicon with 0.13 nanometer precision. It is the most precise semiconductor manufacturing in the world, enabling systems with world-leading algorithmic fidelity, and a decisive advantage in the global quantum computing race.

Our products are commercially deployed and generating revenue. Watermelon, our quantum-enhanced AI system, is delivering superior results on real-world problems across energy, telecom and finance. Quantum Twins, our simulation platform, provides unparalleled ability to model quantum systems, accelerating molecule and materials discovery.

This is SQC: building the future of computing while delivering quantum impact today.

Silicon Quantum Computing (SQC) is at the forefront of global efforts to build the world's first commercial-scale quantum computer, while delivering quantum-enhanced AI and simulation products to customers today.

Backed by over 25 years of technological excellence, SQC is a full-stack quantum computing company that leverages its proprietary manufacturing process to engineer atomic qubits in silicon with 0.13 nanometer precision. It is the most precise semiconductor manufacturing in the world, enabling systems with world-leading algorithmic fidelity, and a decisive advantage in the global quantum computing race.

Our products are commercially deployed and generating revenue. Watermelon, our quantum-enhanced AI system, is delivering superior results on real-world problems across energy, telecom and finance. Quantum Twins, our simulation platform, provides unparalleled ability to model quantum systems, accelerating molecule and materials discovery.

This is SQC: building the future of computing while delivering quantum impact today.

About the role

We are hiring a Platform Engineer to own monitoring and telemetry in Platform & Infrastructure. The scope is everything the team runs: the Kubernetes clusters from small local sites to the central cluster, Ceph storage, core network services, the HPC and simulation environments, the quantum runtime cluster and its FPGA mesh, our AWS footprint, and the cluster that ships with each quantum computer.

The workloads are not a web estate. A FPGA synthesis run takes hours and holds a licence while it does. A simulation job may need to be reproducible years later. The realtime control plane cares about microsecond jitter rather than request percentiles. A physicist will want to correlate a device measurement with a calibration run and a cluster event, so telemetry from unrelated systems has to be joinable.

Compute is on-premise and finite. Utilisation numbers decide where the next hardware spend goes and which team is under-served, so they end up in procurement and scheduling decisions. You own the platform and set the practice: the conventions other teams instrument against, and alerting an on-call engineer acts on without checking it twice.

Based at our Sydney facility, you will work alongside the platform and infrastructure engineers who run the estate, and with the research teams instrumenting their own work. This is a role for someone who wants observability to carry real decisions, and who would rather retire a noisy alert than tune it out.

Role responsibilities
  • Own the monitoring and telemetry platform end to end: metrics, logs, traces and alerting across clusters, storage, network and cloud
  • Instrument the platform services the team runs, including Kubernetes, Ceph, the CI/CD platform, GitHub Enterprise, Artifactory and Vault
  • Make utilisation and capacity visible and trustworthy, so scheduling decisions and hardware purchases rest on measurement
  • Design alerting that names a specific action and keeps noise low, and retire alerts that no longer prompt one
  • Support the quantum runtime cluster's observability needs, including high-rate telemetry from the FPGA mesh, without disturbing the realtime path
  • Extend monitoring to every cluster we build, one per quantum computer, including air-gapped sites where telemetry cannot leave the building and has to be useful to whoever is standing next to it
  • Handle the data engineering of observability at scale: cardinality, sampling, retention and the cost of keeping it
  • Define and measure service levels for internal platform services, and report against them
  • Build the on-call tooling and runbooks, and improve them after every incident
  • Help other teams instrument their own systems, and provide the libraries, conventions and defaults that make that easy
  • Correlate telemetry across domains, so a research question spanning device, cluster and job data can be answered
  • Support incident response and post-incident review across the platform, and document the platform, its conventions and its limits, so observability does not depend on knowledge held by one person
Your experience

Essential

  • 5+ years in platform, infrastructure or site reliability engineering, with direct ownership of an observability stack
  • Prometheus and Grafana at scale, including long-term storage with Thanos, Mimir, VictoriaMetrics or an equivalent
  • Log pipelines in production: Loki, OpenSearch or ELK, with a shipper such as Vector or Fluent Bit
  • OpenTelemetry, and distributed tracing in a real system rather than a demo
  • Alerting and on-call design, including direct on-call responsibility for systems you instrumented
  • Production Kubernetes depth: workloads, networking, storage, RBAC, operators and the failure modes of each
  • Strong Linux systems administration, with Python and Bash
  • Infrastructure-as-code and GitOps: Terraform or OpenTofu, with Argo CD or Flux
  • High-cardinality time series in practice, and the retention and cost decisions that come with it
  • Capacity planning on finite on-premise hardware
  • Clear technical writing, and the ability to support researchers and engineers instrument
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Platform Engineer: Telemetry & Observability for Quantum
Platform Engineer: Telemetry & Observability for Quantum

SQC Silicon Quatum Computing • Sydney

On-site
AUD 140,000 - 210,000
Platform Engineer
Platform Engineer

Silicon Quantum Computing • Sydney

On-site
AUD 120,000 - 170,000
Platform Engineer
Platform Engineer

SQC Silicon Quatum Computing • Sydney

On-site
AUD 120,000 - 180,000
Infrastructure Engineer, Hardware & Networks
Infrastructure Engineer, Hardware & Networks

Silicon Quantum Computing • Sydney

On-site
AUD 120,000 - 190,000
Lead Observability Platform Engineer
Lead Observability Platform Engineer

Cox Purtell Staffing Services • Sydney

On-site
AUD 140,000 - 190,000
Real bonus
Rewards package
Competitive salary
Platform Engineer: HPC, Kubernetes & DevOps
Platform Engineer: HPC, Kubernetes & DevOps

Silicon Quantum Computing • Sydney

On-site
AUD 120,000 - 170,000
Infrastructure Engineer, Hardware & Networks
Infrastructure Engineer, Hardware & Networks

SQC Silicon Quatum Computing • Sydney

On-site
AUD 110,000 - 150,000
ML Ops Engineer, Quantum Machine Learning
ML Ops Engineer, Quantum Machine Learning

SQC Silicon Quatum Computing • Sydney

On-site
AUD 140,000 - 170,000
Platform Engineer: HPC & DevOps for Quantum Compute
Platform Engineer: HPC & DevOps for Quantum Compute

SQC Silicon Quatum Computing • Sydney

On-site
AUD 120,000 - 180,000
Software Engineer, Control & Error Correction
Software Engineer, Control & Error Correction

SQC Silicon Quatum Computing • Sydney

On-site
AUD 150,000 - 210,000