Senior Site Reliability Engineer

Replicant

United States

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Offsites
Tech & learning stipend
Remote by design
Health & wellness benefits
Competitive salaries
Equity with upside

Job summary

Replicant is seeking a Site Reliability Engineer to join our growing SRE team and own domains like platform engineering, CI/CD, and observability. You will help design scalable systems that support real-time AI traffic, improve tooling, and reduce toil across production.

You’ll work with TypeScript/Node, Python, Terraform, Kubernetes on GCP, and collaborate with remote teams to keep our AI-native platform available and secure. Remote-friendly, compensation reflects impact and equity.

Qualifications

  • 6+ years in software or site reliability roles.
  • Experience owning CI/CD platforms end-to-end and developer self-service.
  • Familiar with AI tools for coding and debugging; strong opinions on tool use.

Responsibilities

  • Contribute to patterns, design, and implementation of platform domains.
  • Build and improve systems to reduce toil and maintain availability under real-time AI traffic.
  • Extend and iterate the agent harness with CI, sandboxes, and guardrails.
  • Own and improve CI/CD pipelines and developer tooling for new services.
  • Participate in on-call rotation and incident management to ensure uptime.

Skills

CI/CD ownership
AI tooling (Claude/Cursor)
Node/TypeScript
Python
Terraform
Kubernetes/Helm
Observability
Remote teamwork

Tools

Datadog
Prometheus
Grafana
GitLab CI
Helm
Terraform

Job description

At Replicant, we believe AI should work for people, starting with customer service. That’s why we built a platform that helps contact centers resolve more requests, proactively identify issues, and improve agent performance with AI-powered conversation intelligence and AI agents that act like your best reps.

Our AI agents handle millions of calls every month for Fortune 500 companies and high-growth innovators. From processing payments to booking appointments and authenticating users, they help customers get what they need instantly, 24/7. Meanwhile, our real-time conversation insights help contact center leaders coach better and improve every interaction.

We are leading the shift from legacy systems to AI-first service, powered by large language models (LLMs) and designed for enterprise scale, security, and empathy. If you’re excited by the potential of LLMs, voice AI, and building category-defining technology with a kind, ambitious team, you’ll love it here.

Our SRE team builds - not just supports - an AI-native platform, and they own exciting domains like platform and harness engineering, site reliability, cloud infrastructure, CI/CD, DevEx, observability, incident management, and COGS (e.g. cloud spend visibility). We’re looking for a Site Reliability Engineer who has opinions about how these domains should work and wants agency in shaping where they go. If you are energized by enabling teams to succeed through systems- and patterns-level work, come build Replicant’s platform with us!

What You’ll Do
  • Contribute to patterns, design, and implementation of our domains; help shape the future of platform engineering at Replicant.
  • Build and improve systems that help reduce toil and enable Replicant's production infrastructure to remain available and operable under large-scale, real-time conversational AI traffic.
  • Extend and iterate our agent harness: Unsupervised AI agents are currently used by about 10% of the dev team – help us grow that number. The agent harness includes CI, sandboxes, guardrails, and validation (e.g. agent-first eval loops).
  • Own and improve our CI/CD pipelines and surrounding developer tooling: build and test performance, deployment ergonomics, and paved paths for new services.
  • Participate in on-call rotation and incident management to ensure platform uptime and quality. (SRE owns the base infrastructure, not the applications; non-business-hours pages are rare)
What You’ll Bring
  • 6+ years’ experience in software development enablement roles.
  • Solid experience owning CI/CD platforms end to end – including domains like caching, architecture, and developer self-service.
  • Effective use of AI tools such as Claude and Cursor for coding, troubleshooting, and reasoning. You pair these skills with a defensible opinion on where to avoid using AI tools.
  • Familiarity with Node/TypeScript including making code changes (e.g. exposing new metrics), Python and Terraform for automation, and developing in a Kubernetes/Helm ecosystem.
  • Practical experience with observability: logs/metrics/tracing, monitoring/alerting, incident management process, and tooling.
  • Experience working in fully remote teams – tell us how you’ve made one work better.
  • Bonus:
    • Harness engineering experience – building platforms for autonomous agents.
    • Production-at-scale experience with GCP.
    • Telephony and SIP architectures, FreeSWITCH in particular.

Our stack is TypeScript/Node and Python running on Kubernetes – primarily on GCP (we are multi-cloud), with GitLab CI, Helm, Terraform, Datadog, Prometheus, and Grafana.

For All Full-time Employees, We Offer
  • In-person connection that counts: company-wide offsites and smaller team gatherings designed to make remote work feel personal
  • Tech & learning stipend: Conferences, books, courses – interested? We’ll fund them
  • Remote by design: We’re distributed – no guilt about life events, we trust you to manage your calendar
  • Health & wellness: Flexible vacations, paid sabbatical after 5 years, comprehensive benefits, plus a stipend to support your physical and mental well-being
  • Compensation that matches your impact: competitive salaries in the company you’re helping to build
  • Equity with upside: We believe in shared ownership—You’ll own a real piece of a fast-growing AI company
Our Values

Replicant has three core values. It is critical that everyone who joins the team feels excited and moved by these values as every new team member makes an impact on our culture.

Blade Runners

We take ownership and pride to influence the outcomes of our goals. We are successful, and like a Blade Runner, use the tools at our disposal to reach our objectives. We value open and honest communication and proactively seek feedback along the way. We are a company driven to grow and achieve both individually and as a team.

Bread Makers

We are humble and strive toward an egalitarian culture. No task is too big or too small. We work together to achieve our goals and develop our company mission. We believe that the whole is greater than the sum of its parts in everything that we do.

Självdistans (Self-Distance)

Självdistans is Swedish for self-distance. It's the ability to critically reflect on oneself and one's relations from an external perspective. With this in mind, we act with objectivity and always remember that we are not our work. There's no perfect science to growing a team or business, but we trust everyone at Replicant to point out our blind spots and humbly admit their own.

Replicant is proud to be an equal opportunity employer. We are committed to fostering an inclusive, diverse and equitable workplace that is built on trust, support and respect. We welcome all individuals and do not discriminate on the basis of gender identity and expression, race, ethnicity, disability, sexual orientation, colour, religion, creed, gender, national origin, age, marital status, pregnancy, sex, citizenship, education, languages spoken or veteran status. Accommodation is available upon request at any point during our recruitment process. If you require an accommodation, please speak to your talent acquisition partner or email us at talent@replicant.ai and we’ll work to meet your needs.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bot Jobs • Myrtle Point (OR)

Remote
USD 170,000 - 210,000
In-person connection offsites
Tech & learning stipend
Remote by design
+2
AI Enablement Engineer
AI Enablement Engineer

Replicant • United States

Remote
USD 140,000 - 190,000
In-person offsites
Tech & learning stipend
Remote by design
+2
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Replicant • United States

Remote
USD 120,000 - 180,000
In-person offsites for remote teams
Tech & learning stipend
Remote by design
+2
Senior Director, Revenue Operations
Senior Director, Revenue Operations

Replicant • Northern (KY)

Hybrid
USD 180,000 - 240,000
In-person offsites
Tech & learning stipend
Remote by design
+2
Senior Site Reliability Engineer (Fully Remote)
Senior Site Reliability Engineer (Fully Remote)

Replicant • United States

Remote
USD 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Replit • Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive Salary & Equity
401(k) 4% match (US)
Health, Dental, Vision & Life
+7
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Replit • United States

Remote
USD 120,000 - 210,000
Competitive Salary
Equity
401(k) Match
+16
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Replit • Foster City (CA)

On-site
USD 180,000 - 260,000
Competitive Salary & Equity
401(k) with 4% match
Health, Dental, Vision and Life Ins.
+2
Engineering Manager, Site Reliability Engineering
Engineering Manager, Site Reliability Engineering

Replit • Foster City (CA)

On-site
USD 190,000 - 230,000
Competitive salary
Equity
401(k) match
+4
Staff Software Engineer, Agent Platform
Staff Software Engineer, Agent Platform

Replit • Foster City (CA)

On-site
USD 130,000 - 160,000
Competitive Salary & Equity
401(k) Program with a 4% match
Health, Dental, Vision and Life Insurance
+9