Site Reliability Engineer

DeepJudge AG

Zürich

Vor Ort

CHF 120.000 - 180.000

Vollzeit

Vor 12 Tagen
Bewerbungsgenerator

Hebe dich für diese Rolle von der Masse ab — erstelle in etwa einer Minute einen maßgeschneiderten Lebenslauf und ein Anschreiben.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

DeepJudge AG in Switzerland is seeking a Site Reliability Engineer to own the reliability of our AI-powered search platform used by leading law firms. You will manage production environments, across multi-cloud deployments, ensuring speed, availability and trust in data processing and retrieval at scale.

The role emphasizes infrastructure-as-code, observability, SLOs, and automation with Python or Go, collaborating with product engineering to prevent outages and improve performance in a

Qualifikationen

  • 3+ years in SRE, production/platform engineering, or similar role.
  • Experience running production systems with reliability targets.
  • Strong container orchestration in production.
  • Experience with cloud providers and infrastructure as code.

Aufgaben

  • Own the availability, latency and performance of production environments.
  • Manage multi-cloud infrastructure, networking, IAM and storage.
  • Improve deployment via GitOps and infrastructure-as-code with safe changes.
  • Define and instrument SLOs and dashboards for observability.
  • Automate operational toil with Python and/or Go.
  • Provision and standardize new production environments to a reliability bar.

Kenntnisse

SRE / production engineering
Container orchestration
Python
Go
GitOps
Observability
Cloud infrastructure

Tools

Kubernetes
Docker
Terraform
CI/CD

Jobbeschreibung

About Us:

DeepJudge is Switzerland's leading AI and ICT scale-up, transforming how law firms and legal departments access and leverage their knowledge. Founded by former Google search engineers with PhDs in AI from ETH Zurich, and backed by top-tier Silicon Valley investors, we are building the intelligence layer that powers the legal industry.

Every firm can license the same AI models, but no two firms share the same institutional knowledge. Decades of experience, work product, and precedent are often trapped across systems, underused and inaccessible. DeepJudge makes this expertise instantly available through world-class enterprise search and AI infrastructure, enabling firms to automate workflows, build knowledge-powered applications, and turn experience into a lasting advantage.

Our technology is trusted by many of the world's leading law firms, including Freshfields, Gunderson Dettmer, Holland & Knight, Arent Fox, and Cozen O'Connor. Headquartered in Switzerland with a growing team across North America, we are expanding rapidly and shaping the future of how professional knowledge is discovered and applied.

At DeepJudge, we move fast, think rigorously, and care deeply about what we build. We combine cutting-edge research with close collaboration with our clients to deliver solutions that truly make an impact. If you want to be part of a team defining how AI transforms high-stakes knowledge work, this is the place to do it.

About the role:

As a Site Reliability Engineer at DeepJudge, you own the reliability of the systems that leading law firms depend on every day. We run our AI search platform in dedicated production environments, each a large-scale search and retrieval system operating at very high data volumes, alongside our ingestion pipelines and data stores, on container-orchestrated infrastructure across multiple public and private clouds. Keeping those environments fast, available, and trustworthy is a genuine engineering challenge, and it's your mission.

This is a technical, infrastructure-focused role: you'll spend your time in systems, code, and telemetry. You'll harden the platform so that classes of failure can't recur, build observability and SLOs to let us see problems before they surface, plan capacity ahead of relentless data growth, and automate away the toil of safely operating many production environments. You'll work across our deployment and infrastructure-as-code and partner closely with product engineering to feed reliability learnings back into the platform. If you like operating serious distributed systems at scale and turning firefighting into engineered reliability, this role is for you.

Your responsibilities will include:
  • Owning the availability, latency and performance of our production environments: large-scale search and retrieval services, ingestion and data stores on container-orchestrated infrastructure across multiple public and private clouds
  • Owning the cloud infrastructure across multiple public and private clouds, networking, identity and access, compute and storage, and managing cloud capacity, quotas and cost as the fleet grows
  • Improving deployment and infrastructure through code: GitOps-based configuration and deployment and infrastructure-as-code across our public and private clouds, with safe, previewed, recoverable production changes
  • Defining and instrumenting SLOs and error budgets and building the dashboards, metrics and alerting that make production observable and actionable
  • Reducing operational toil through automation and internal tooling written in Python and / or Go
  • Provisioning and standardizing new production environments to a defined reliability bar
You're a great fit if you:
  • Have 3+ years in SRE, production / platform engineering, or infrastructure-heavy software engineering, operating real production systems under reliability expectations
  • Are strong at container-orchestration operations in production and comfortable on at least one major public cloud
  • Have solid cloud infrastructure experience on a major public cloud and a working grasp of cloud capacity, quotas, and cost
  • Practice observability hands-on with metrics, logs and traces, and use it to root-cause live incidents
  • Work fluently with infrastructure-as-code and GitOps
  • Can code in Python and / or Go for automation, tooling, and reading and patching service code.
  • Bring sound production judgment and document your work clearly for other engineers
  • Bonus: experience operating a large-scale search or distributed data system, database operations, dedicated / enterprise-deployed software, or AI / enterprise-search infrastructure in regulated domains
What you can look forward to:
  • Owning the reliability of systems that the world's leading law firms rely on: real scale, real impact
  • Deep, hands-on work with modern infrastructure: distributed search at very large scale, container orchestration across multiple clouds, and a mature infrastructure-as-code
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Site Reliability Engineer
Site Reliability Engineer

DeepJudge • Zürich

Vor Ort
CHF 120.000 - 180.000
Forward Deployed Engineer
Forward Deployed Engineer

DeepJudge AG • Zürich

Vor Ort
CHF 110.000 - 170.000
Platform Engineer
Platform Engineer

DeepJudge AG • Zürich

Vor Ort
CHF 120.000 - 160.000
On-site in Zurich
Office near Zurich HB
Forward Deployed Engineer
Forward Deployed Engineer

DeepJudge • Zürich

Vor Ort
CHF 110.000 - 160.000
Production Engineer
Production Engineer

DeepJudge • Zürich

Vor Ort
CHF 120.000 - 180.000
Senior Site Reliability Engineer - AI Infra & Multicloud
Senior Site Reliability Engineer - AI Infra & Multicloud

DeepJudge • Zürich

Vor Ort
CHF 120.000 - 180.000
Site Reliability Engineer — AI Search Infra on Multi-Cloud
Site Reliability Engineer — AI Search Infra on Multi-Cloud

DeepJudge AG • Zürich

Vor Ort
CHF 120.000 - 180.000
Platform Engineer
Platform Engineer

DeepJudge • Zürich

Vor Ort
CHF 120.000 - 180.000
Office near Zurich
Collaborative environment
Professional growth opportunities
Machine Learning Engineer
Machine Learning Engineer

DeepJudge AG • Zürich

Vor Ort
CHF 120.000 - 170.000
Office near Zurich Central Station
Pre-Sales Engineer
Pre-Sales Engineer

DeepJudge AG • Zürich

Hybrid
CHF 130.000 - 180.000