Remote SRE: Scale Cloud Infra with Kubernetes & Golang

ArangoDB, Inc.

United States

Remote

USD 9,400 - 16,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

ArangoDB, Inc. is seeking a Site Reliability Engineer to strengthen cloud infrastructure and reliability for our distributed database systems.

You’ll work on Kubernetes-based services across AWS and Google Cloud, building automation in Golang and Python, and enhancing CI/CD and observability. You’ll collaborate with product and engineering teams to optimize performance, implement robust monitoring, and scale our platforms for high availability.

Qualifications

  • 4–7 years of SRE/DevOps experience in cloud-native environments.
  • Strong experience with AWS or GCP and Kubernetes at scale.
  • Proficient in Go or Python for automation.
  • Experience with CI/CD, monitoring and logging stacks.
  • Solid networking, security best practices, and incident response.

Responsibilities

  • Design, implement, and maintain cloud infrastructure on AWS and Google Cloud platforms.
  • Ensure the scalability, performance, and reliability of our Kubernetes-based distributed database systems.
  • Collaborate with developers to write efficient, production-grade code in Golang to automate infrastructure management and improve system operations.
  • Optimize and automate CI/CD pipelines, deployment processes, and monitoring systems to support our production environment.
  • Develop strategies for disaster recovery, high availability, and fault tolerance.
  • Proactively identify system bottlenecks, troubleshoot, and resolve issues across the stack (network, OS, cloud infrastructure).
  • Implement monitoring, logging, and alerting systems to ensure visibility into system health and performance.
  • Participate in on-call rotations to support critical production systems and respond to incidents.
  • Collaborate with cross-functional teams to improve overall system reliability and scalability.
  • Collaborate with the Customer Success team to resolve customer issues.

Skills

AWS
GCP
Kubernetes
Docker
Golang
Python
CI/CD
Prometheus
Grafana
ELK stack
Networking
Security

Tools

Terraform
Git
Jenkins
CircleCI

Job description

ArangoDB, Inc. is seeking a Site Reliability Engineer to strengthen cloud infrastructure and reliability for our distributed database systems.

You’ll work on Kubernetes-based services across AWS and Google Cloud, building automation in Golang and Python, and enhancing CI/CD and observability. You’ll collaborate with product and engineering teams to optimize performance, implement robust monitoring, and scale our platforms for high availability.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Go & Kubernetes Engineer — Operator Lead (Remote)
Senior Go & Kubernetes Engineer — Operator Lead (Remote)

ArangoDB, Inc. • United States

Remote
USD 140,000 - 210,000
Remote Senior Go & Kubernetes Operator Engineer
Remote Senior Go & Kubernetes Operator Engineer

Rollbar, Inc. • San Francisco (CA)

On-site
USD 150,000 - 190,000
Senior Go & Kubernetes Operator Engineer (Remote)
Senior Go & Kubernetes Operator Engineer (Remote)

Arango • United States

On-site
USD 140,000 - 230,000
Remote SRE: Infrastructure Security & Cloud Platform Ops
Remote SRE: Infrastructure Security & Cloud Platform Ops

The available sources do not contain information about the company name for rounx.com. • Austin (TX)

On-site
USD 127,000 - 249,000
Remote SRE: AWS, Kubernetes & Observability
Remote SRE: AWS, Kubernetes & Observability

Cloudbeds • United States

Remote
USD 120,000 - 150,000
Remote First
PTO
Home office stipend
+2
Golang/Kubernetes Engineer
Golang/Kubernetes Engineer

ArangoDB, Inc. • United States

Remote
USD 140,000 - 210,000
Golang/Kubernetes Engineer
Golang/Kubernetes Engineer

Arango • United States

On-site
USD 140,000 - 230,000
Remote SRE: Blockchain Infrastructure & Scale
Remote SRE: Blockchain Infrastructure & Scale

Embedded Shishya • United States

On-site
USD 150,000 - 210,000
Remote‑first global workforce
Professional reimbursement program
Medical, dental & vision coverage
+3
Blockchain SRE: Scale, Security & Reliability (Remote)
Blockchain SRE: Scale, Security & Reliability (Remote)

cyber • Northern (KY)

Hybrid
USD 140,000 - 190,000
Remote-first global workforce
Professional reimbursement program
Medical, dental & vision coverage (US)
+3
Remote SRE: Scale Reliability with Cloud Automation
Remote SRE: Scale Reliability with Cloud Automation

Glassbox • New York (NY)

On-site
USD 120,000 - 150,000