Senior Site Reliability Engineer

Camunda

New York (NY)

Remote

USD 140,000 - 180,000

Full time

26 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Camunda is a fully remote, global enterprise platform for agentic orchestration, enabling AI agents, people, and systems to coordinate complex business processes.

We’re seeking a Senior Site Reliability Engineer to design and maintain a Kubernetes-based multi-cloud platform, improve observability, and drive automation. You’ll own end-to-end systems, mentor others, and collaborate with product teams to ship reliable software at speed.

Qualifications

  • : Extensive hands-on Kubernetes in production environments with scale.
  • :
  • We want Terraform or similar IaC proficiency for safe, repeatable deployments.
  • We expect experience with monitoring tools like Prometheus and Grafana to instrument systems.
  • Strong incident response skills with RCA experience and clear communication during outages.
  • A passion for automation and raising system quality through reliable, maintainable architectures.
  • Experience using AI tools responsibly, with governance and human accountability.

Responsibilities

  • Design and maintain Kubernetes-based multi-cloud platform architecture, ensuring availability, scalability, and fault tolerance.
  • Build observability that matters by implementing monitoring and alerting with clear visibility across the stack.
  • Own your systems end-to-end, participate in on-call rotations, and develop runbooks for fast recovery.
  • Collaborate across product engineering, product management, and support to deliver impactful features.
  • Automate repetitive tasks and share learnings to raise the team's overall capability.
  • Mentor teammates and help less experienced engineers tackle complex infrastructure challenges.

Skills

Kubernetes
IaC
Monitoring/Observability
Incident response
Automation
Responsible AI usage
AWS/GCP experience
GitOps (ArgoCD)

Tools

Terraform
Prometheus
Grafana
AWS (EKS)
GCP (GKE)
ArgoCD
GitOps

Job description

Camunda is the enterprise platform for agentic orchestration, enabling organizations to coordinate AI agents, people, and systems across complex, end-to-end business processes. With built-in governance, auditability, and human oversight, Camunda gives enterprises the control they need to move AI from pilots to production — safely and at scale. Trusted by over 700 organizations worldwide, including 9 of top 10 US banks, Camunda helps enterprises boost operational efficiency, accelerate time-to-value, and deliver better customer experiences.

Fully remote and global, we are in the middle of something bigger: transforming into an AI‑first organisation, built on our own platform. We use Agentic AI to automate, orchestrate intelligent processes, and elevate human contribution across every team.

Named GP Bullhound’s Top 100 Next Unicorn list, 2025 Great Place to Work certified. Visionary in 2025 Gartner® Magic Quadrant™ for Business Orchestration and Automation Technologies. Ranked 3rd in Flexa's 2026 Most Flexible Companies, We’re growing fast and looking for top talent to join our team. If you want meaningful work, visible impact and put something genuinely rare on your CV, keep reading.

About The Role

We're looking for a Senior Site Reliability Engineer who's passionate about building reliable, scalable infrastructure that helps developers ship better software faster. You'll design and maintain our Kubernetes-based multi‑cloud platform, improve our monitoring and observability tools, and collaborate with product and engineering teams to solve real problems in real time. This is a role where you'll own the systems that power Camunda, drive automation that raises the bar for everyone, and mentor others who want to do the same. If you thrive on building things that work well and don't break when they shouldn't, we'd love to talk.

What You’ll Be Doing
  • Design and maintain our infrastructure – You'll evolve our Kubernetes-based, multi‑cloud platform architecture, ensuring it's available, scalable, and fault‑tolerant. You'll establish configuration best practices and network services that our teams rely on.
  • Build observability that matters – Implement and improve monitoring and alerting tools that give both SREs and developers real visibility into system health and performance. Make it easy for teams to understand what's happening across our stack.
  • Own your systems end‑to‑end – You'll adopt a "you build it, you run it" mentality, which means participating in on‑call rotations and being the go‑to person when things need quick fixes. You'll create runbooks and automation that turn complex problems into manageable processes.
  • Ship features and improvements with product teams – Work cross‑functionaly with product engineering, product management, and support to define and deliver features that move the needle. Bring your expertise to the table early and often.
  • Push automation to the next level – Identify repetitive work and automate it away. Share what you learn with your teammates so everyone gets better at their craft. We measure progress by the quality of our systems, not hours worked.
  • Be the expert others learn from – Help less experienced engineers tackle complex infrastructure challenges. Break down technical problems into clear steps and support others in growing their skills.
What You Bring
Must Haves:
  • Deep hands‑on experience with Kubernetes – You've built, deployed, and maintained Kubernetes clusters in production environments. You understand how to manage workloads, networking, and storage at scale.
  • Infrastructure as code expertise – You're fluent in tools like Terraform (or similar IaC tools) and know how to version, test, and safely deploy infrastructure changes.
  • Demonstrated experience in monitoring and observability – You've worked with tools like Prometheus, Grafana, or similar platforms to instrument systems and alert on what matters.
  • Strong 3rd‑level support and incident response skills – You've responded to production incidents, diagnosed complex issues, and communicated clearly with stakeholders under pressure. You understand root cause analysis and how to prevent issues from happening again.
  • A passion for automation and raising the quality bar – You see manual work as a problem to be solved. You care deeply about building systems that are reliable, maintainable, and easy to understand.
  • Responsible use of AI tools – You leverage AI for research, code review, documentation, test generation, and automation to improve your effectiveness. You know how to validate AI outputs against requirements, never share confidential or personal data with AI systems without authorization, and always retain human accountability for decisions and deliverables.
Nice-to-haves
  • Experience with major cloud providers – You've worked with AWS (EKS), Google Cloud Platform (GKE), or similar managed Kubernetes services.
  • ArgoCD or GitOps workflows – You've used declarative
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Camunda • United States

Remote
USD 150,000 - 242,000
Remote & Flexible
Annual Kickoff & team offsites
Health & Wellbeing
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Camunda • Atlanta (GA)

Remote
USD 150,000 - 242,000
Remote work
Annual company events
Health & wellbeing
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Camunda • Boston (MA)

Remote
USD 150,000 - 242,000
Remote & Flexible
Annual Kickoff & travel
Health & Wellbeing
+2
Engineering Manager - Infrastructure
Engineering Manager - Infrastructure

Camunda • United States

On-site
GBP 108,400 - 178,300
Remote & Flexible work arrangements
Health & Wellbeing support
Professional growth budget
Senior Software Engineer, Kubernetes
Senior Software Engineer, Kubernetes

Camunda • Atlanta (GA)

Remote
USD 150,000 - 242,000
Remote & Flexible
Annual Kickoff travel & team events
Health & wellbeing programs
+2
Senior Software Engineer, Kubernetes
Senior Software Engineer, Kubernetes

Camunda • United States

On-site
USD 150,000 - 242,000
Remote & Flexible
Annual Kickoff
Live Well Lifestyle Spending Account
+1
Senior Software Engineer, Kubernetes
Senior Software Engineer, Kubernetes

Camunda • Austin (TX)

Remote
USD 150,000 - 242,000
Remote & Flexible
Remote Senior SRE — Kubernetes, Observability & Automation
Remote Senior SRE — Kubernetes, Observability & Automation

Camunda • New York (NY)

Remote
USD 140,000 - 180,000
Senior SRE - Kubernetes, Observability & Automation (Remote)
Senior SRE - Kubernetes, Observability & Automation (Remote)

Camunda • Atlanta (GA)

Remote
USD 150,000 - 242,000
Remote work
Annual company events
Health & wellbeing
+2
Senior SRE: Remote, Scalable Kubernetes & Automation Lead
Senior SRE: Remote, Scalable Kubernetes & Automation Lead

Camunda • Boston (MA)

Remote
USD 150,000 - 242,000
Remote & Flexible
Annual Kickoff & travel
Health & Wellbeing
+2