Lead SRE (Site Reliability Engineer)

SQD

Schweiz

Remote

CHF 120,000 - 180,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Fully remote
Flexible hours
Token incentives
Ownership of reliability

Job summary

SQD is seeking an experienced SRE to scale Portal with enterprise-grade reliability. You will own CI/CD, orchestration, and infrastructure-as-code, building a high-availability ingestion pipeline with sub-second latency and 99.9% uptime.

Internally, you’ll optimize hosting strategies across cloud and bare metal, monitor systems, and define SRE metrics, logging, and alerting. This fully remote role offers ownership of reliability and cost-effective infra for a frontier AI/Web3 platform.

Qualifications

  • 3+ years of experience as an SRE, DevOps Engineer, or similar role with reliability achievements.
  • Experience running production services with on-call, incident response, and post-mortems.
  • Experience defining metrics, logging, and alerting to keep production healthy.
  • Proficiency in Kubernetes, Terraform, Prometheus, Grafana or equivalent tools.
  • Strong understanding of distributed systems, data pipelines, and failure modes.
  • Cloud infrastructure knowledge (AWS, GCP) and cost optimization.
  • Ability to balance performance, reliability and cost, choosing hosted vs self-hosted.
  • Programming skills (Python, Go, Rust, or Bash) for automation.
  • Experience monitoring blockchain nodes or working with node providers.
  • Willingness to learn internals of blockchain nodes and EVM/SVM data.

Responsibilities

  • Design, build, and optimize a high-availability data ingestion pipeline with sub-second latency and 99.9% uptime.
  • Own CI/CD, orchestration, and infrastructure-as-code layers.
  • Identify and implement tools to run and monitor blockchain nodes with failover.
  • Assess infrastructure trade-offs to optimize performance, reliability, and cost.
  • Build and maintain a public status page and incident management process.
  • Collaborate with engineers integrating new chains and patches to ingestion pipeline.
  • Define and maintain SRE metrics, logging, and alerting.
  • Improve observability, automation, and fault tolerance.
  • Contribute to incident response and on-call rotations.

Skills

SRE experience
On-call experience
Kubernetes
Terraform
Prometheus/Grafana
Cloud infra (AWS/GCP)
Automation (Python/Go/Rust/Bash)
Blockchain nodes/Web3
Incident post-mortems
Observability

Tools

Kubernetes
Terraform
Prometheus
Grafana
CI/CD tooling

Job description

About SQD

At SQD, we're redefining the database layer for the AI and Web3 world. Our flagship product, Portal, streams validated real-time and historical data from 120+ networks to developers building across DeFi, AI agents, and Web3 ecosystems. As we scale, we need an experienced SRE to help take our systems to enterprise-grade infrastructure.

What you'll do
  • Help design, build, and optimize a high-availability blockchain data ingestion pipeline with sub-second latency and 99.9% uptime.
  • Own our CI/CD, orchestration, and infrastructure-as-code layers.
  • Identify and implement the right tools to run and monitor blockchain nodes, with failover solutions using multiple node providers.
  • Continuously assess infrastructure trade-offs (hosted nodes, subscriptions, bare metal) to achieve optimal performance, reliability, and cost-efficiency.
  • Build and maintain a public status page, and the incident management process behind it: severities, escalation paths, and post-mortems that change how we build.
  • Work closely with engineers integrating new chains, making necessary patches to the ingestion pipeline.
  • Define and maintain key SRE metrics, logging, and alerting, proactively identifying and resolving reliability risks.
  • Stay ahead of past incidents, continuously improving observability, automation, and fault tolerance.
  • Contribute to incident response, troubleshooting, and on-call rotations.
Requirements
  • 3+ years of experience as an SRE, DevOps Engineer, or similar role, with reliability achievements you can point to and talk through.
  • Experience running production services against real SLAs, including on-call, incident response, and post-mortems.
  • Experience defining and implementing metrics, logging, and alerting to keep production healthy and prevent incidents.
  • Proficiency in Kubernetes, Terraform, Prometheus, Grafana, or equivalent monitoring tools.
  • Strong understanding of distributed systems, streaming data pipelines, and the failure modes that come with them.
  • Deep knowledge of cloud infrastructure (AWS, GCP, or bare metal setups) and cost optimization strategies.
  • Ability to balance performance, reliability, and cost, assessing when to use hosted nodes, subscriptions, or self-hosted setups.
  • Programming skills (Python, Go, Rust, or Bash) for automation and infrastructure tooling.
  • Experience monitoring and running blockchain nodes (Ethereum, Solana, etc.), and/or working with node providers as a backup solution.
  • Willingness to learn and understand the internals of blockchain nodes, EVM/SVM data, and work with engineers to integrate new chains.
  • Previous experience in Web3 is preferred.
Benefits
  • Competitive salary + token incentives
  • Fully remote with flexible hours
  • High-impact role with ownership where your work directly shapes the reliability of the onchain data layer
  • Build the operational foundations for a frontier AI/Web3 company
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote SRE Lead: Blockchain Data Reliability & Ops
Remote SRE Lead: Blockchain Data Reliability & Ops

SQD • Switzerland

Remote
CHF 120,000 - 180,000
Fully remote
Flexible hours
Token incentives
+1
Senior Site Reliability Engineer (SRE) New Lausanne, Switzerland (Hybrid)
Senior Site Reliability Engineer (SRE) New Lausanne, Switzerland (Hybrid)

SpotMe • Lausanne

Remote
CHF 140,000 - 210,000
Site Reliability Engineer - Cloud Operations
Site Reliability Engineer - Cloud Operations

Swissquote • Gland

On-site
CHF 120,000 - 180,000
Site Reliability Engineer - Cloud Operations
Site Reliability Engineer - Cloud Operations

SmartRecruiters, Inc. • Gland

On-site
CHF 110,000 - 150,000
Staff Engineer (Core & MLOps)
Staff Engineer (Core & MLOps)

Jobgether SRL • Switzerland

Remote
CHF 140,000 - 200,000
Fully remote
Flexible hours
Global engineering community
Site Reliability Engineer
Site Reliability Engineer

DeepJudge • Zürich

On-site
CHF 120,000 - 180,000
DevOps Engineer
DevOps Engineer

Blockstream • Zürich

On-site
CHF 95,000 - 125,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • Zürich

On-site
CHF 120,000 - 180,000
Quant DevOps Engineer
Quant DevOps Engineer

SCOR • Zürich

On-site
CHF 180,000 - 260,000
Competitive salary
Benefits package
International team
Senior SRE Engineer
Senior SRE Engineer

Jobgether SRL • Switzerland

On-site
CHF 120,000 - 180,000
Fully remote working model
20 paid working days off
Medical insurance