Site Reliability Engineer (m/f/d)

Deepslate

Germany (OH)

On-site

USD 74,438 - 108,794

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A forward-thinking AI company is seeking an Infrastructure Engineer to enhance their Voice AI models. The ideal candidate will manage Kubernetes environments, ensuring reliable operation and scaling of cloud infrastructure. Candidates must have strong observability skills and a proactive approach to incident management. Proficiency in German is essential for this role. Join us in building the future of communication in a high-paced, innovative environment.

Qualifications

  • Deep hands-on experience in setting up, managing, and scaling Kubernetes clusters.
  • Strong experience with Infrastructure as Code, preferably using Pulumi.
  • Proficient with observability tools such as Datadog and OpenTelemetry.
  • Experience with incident management tools, notably PagerDuty.
  • Hands-on experience with CI/CD pipelines in microservice environments.
  • Proficient in written and spoken German.

Responsibilities

  • Build resilient infrastructure for AI workloads, ensuring uptime and performance.
  • Design and manage cloud infrastructure using Infrastructure as Code.
  • Orchestrate Kubernetes clusters for high efficiency and fault tolerance.
  • Implement monitoring systems for distributed systems.
  • Manage incident response processes and on-call culture.

Skills

Kubernetes
Infrastructure as Code
Observability
Monitoring with Datadog
Incident Management
Integration Testing & CI/CD
Fluent German Skills
Startup Mindset
Extreme Ownership

Tools

Pulumi
PagerDuty
OpenTelemetry

Job description

Location: Remote / Berlin (Office available) | Language Requirement: Fluent German

At Deepslate, we are building Speech to Speech Voice AI models that sound and act indistinguishable from a human. And we believe everyone should be able to use it.

When it comes to text and images, giants like OpenAI and Google have already cracked the code. With video, Veo3, Sora and others are closing the gap rapidly. But with its endless languages, dialects, accents, subtle intonations, and speech melodies voice remains a highly complex unsolved frontier.

That is exactly why we started Deepslate.

Backed by top-tier investors from the Tech and AI sectors, as well as a major German VC fund, we are incredibly well-funded and moving fast.

We are building the future of communication.

We aren\'t trying to build another standalone platform; instead, we are the intelligence engine powering countless other applications. Whether it\'s integrated as a module by a CRM provider, plugged into another Voice AI platform, or directly embedded into an enterprise system by our integration partners - our model is everywhere.

Your Role

Our Voice Models need to handle millions of calls, and you will be the guardian of their uptime, performance, and scalability. If our infrastructure goes down, our customers\' applications grind to a halt.

Your mission is to build an infrastructure so resilient that potential outages are caught and mitigated before they even happen. You are the bridge between development and operations, ensuring that our massive AI workloads run smoothly, efficiently, and with uncompromising high availability.

What You\'ll Do: Bulletproof Our Voice AI Engine

You don\'t build temporary workarounds; you build automated, scalable fortresses. Your focus is on absolute reliability, deep observability, and crafting an infrastructure that effortlessly keeps pace with our rapid growth.

  • Infrastructure as Code: Design, build, and manage our cloud infrastructure using modern tools (Pulumi) to ensure all infrastructure changes are reproducible, secure, and easily auditable.
  • Kubernetes : Orchestrate and optimize our Kubernetes clusters for complex, compute-heavy AI workloads, guaranteeing maximum efficiency and fault tolerance.
  • Deep Observability & Monitoring: Implement a flawless monitoring setup. Using Datadog and OpenTelemetry, you will make the black box of our distributed systems transparent, hunting down latency spikes or bottlenecks before they impact users.
  • Incident Response & Reliability: Establish and manage our on-call and alerting processes (using PagerDuty) and champion a culture of blameless post-mortems so the same mistake never happens twice.
  • Release Confidence: Build and maintain highly automated Integration Testing and deployment pipelines. No code goes live without rigorous validation of its impact on system stability.
Elevate Our Engineering Quality:
  • SLAs, SLOs & SLIs: Define and monitor our service-level metrics, turning reliability into a measurable, core component of our product development cycle.
  • Automation First: Ruthlessly automate away toil (repetitive, manual work) so the engineering team can focus on innovation instead of maintenance.
  • Security & Compliance: Ensure our infrastructure is not only highly available but also locked down and hardened against external threats.
What We’re Looking For:

Must-Haves:

  • Kubernetes: Deep, hands-on experience in setting up, managing, and scaling self-hosted Kubernetes clusters in production.
  • Infrastructure as Code: Strong experience with modern IaC, ideally with Pulumi (or deep Terraform knowledge alongside a willingness to adopt Pulumi).
  • Observability: You are a pro with Datadog and OpenTelemetry. You know exactly how to effectively monitor distributed systems across tracing, metrics, and logs.
  • Alerting & Incident Management: Proven experience with PagerDuty (or similar tools) and a track record of building a healthy, sustainable on-call culture.
  • Integration Testing & CI/CD: Hands-on experience setting up robust testing and deployment pipelines for complex, microservice-based architectures.
  • Fluent German skills: (spoken and written).
  • Startup Mindset: You are comfortable navigating the chaos of an early-stage codebase. If a process is missing or unstructured, you roll up your sleeves and build it.
  • Extreme Ownership: You aren\'t just looking to blindly process Jira tickets. You proactively identify where the infrastructure is burning—or where it will burn in the future—and you take action.

Nice-to-Haves:

  • Experience managing GPU workloads and scaling AI/ML infrastructures.
  • Background in network optimization (highly critical for latency-sensitive voice streaming protocols like WebRTC or WebSockets).
  • Previous experience building high-availability systems in a fast-paced B2B/API-first environment.

Deepslate is an equal opportunity employer. We welcome applications from all qualified candidates regardless of gender, nationality, ethnic origin, religion, disability, age, or sexual orientation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: AI Voice Infra & Kubernetes
Senior SRE: AI Voice Infra & Kubernetes

Deepslate • Germany (OH)

Remote
USD 74,000 - 109,000
Frontend Engineer (m/f/d)
Frontend Engineer (m/f/d)

Deepslate • Germany (OH)

On-site
USD 80,000 - 120,000
Backend Engineer (m/f/d)
Backend Engineer (m/f/d)

Deepslate • Germany (OH)

Hybrid
USD 93,000 - 127,000
Flexible work hours
Virtual open office environment
Total autonomy without hierarchy
Frontend Engineer — Real-Time AI Voice UI (Remote)
Frontend Engineer — Real-Time AI Voice UI (Remote)

Deepslate • Germany (OH)

Hybrid
USD 80,000 - 120,000
Remote Backend Engineer (Kotlin) for Scalable AI APIs
Remote Backend Engineer (Kotlin) for Scalable AI APIs

Deepslate • Germany (OH)

Remote
USD 93,000 - 127,000
Flexible work hours
Virtual open office environment
Total autonomy without hierarchy
Remote Voice AI Model Research Engineer
Remote Voice AI Model Research Engineer

Deepslate • Germany (OH)

Remote
USD 90,000 - 150,000
True R&D Freedom
Competitive Compensation
Serious Compute
+1
Model Research Engineer (m/f/d)
Model Research Engineer (m/f/d)

Deepslate • Germany (OH)

Remote
USD 90,000 - 150,000
True R&D Freedom
Competitive Compensation
Serious Compute
+1
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)

Deepgram • United States

Hybrid
USD 120,000 - 150,000
Medical, dental, vision benefits
Unlimited PTO
Generous paid parental leave
+1
Backend Software Engineer - Engine Team (Voice Agent)
Backend Software Engineer - Engine Team (Voice Agent)

Deepgram • United States

Remote
USD 100,000 - 130,000
Medical, dental, vision benefits
Unlimited PTO
401(k) plan with company match
+1
Backend Software Engineer - Engine Team (Voice Agent)
Backend Software Engineer - Engine Team (Voice Agent)

Deepgram, Inc. • United States

Remote
USD 100,000 - 140,000
Medical, dental, vision benefits
Unlimited PTO
401(k) plan with company match
+2