Remote SRE - AWS, Terraform & Kubernetes

GiveCampus

United States

Remote

USD 120,000 - 180,000

Full time

11 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Remote-friendly
Office in Washington, DC

Job summary

GiveCampus is seeking a hands-on Site Reliability Engineer to improve the reliability, performance, and operational maturity of our platform. With our migration to AWS complete, you will focus on operating and strengthening production: observability, automation, incidents, and collaboration with product engineers to build resilient systems.

You will own reliability projects, automate infrastructure with Terraform, maintain Kubernetes/EKS, and enhance dashboards, alerts, and CI/CD pipelines.

Qualifications

  • Approximately 5+ years of related experience in software engineering, infrastructure, systems engineering, SRE, Platform Engineering, DevOps, or equivalent practical experience.
  • Hands-on experience operating production workloads in AWS.
  • Experience building or maintaining infrastructure using Terraform or a similar infrastructure-as-code tool.
  • Experience with New Relic, Datadog, or another modern observability platform.
  • Experience troubleshooting production incidents and participating in an on-call rotation.
  • Experience building or maintaining CI/CD pipelines.
  • Software development or scripting experience, with the ability to read, debug, and make targeted changes to application or automation code.
  • Working knowledge of Linux, networking, distributed systems, and relational databases.
  • Ability to articulate root causes, explain technical tradeoffs, and translate findings into practical solutions.
  • Ability to manage a well-scoped project with general direction and provide timely updates at key milestones.
  • Strong written and verbal communication skills and a collaborative approach to working across engineering disciplines.
  • A habit of automating repetitive work and improving the reliability of the systems you support.

Responsibilities

  • Operate and improve production infrastructure in AWS.
  • Build and maintain infrastructure as code with Terraform.
  • Support workloads on Kubernetes and Amazon EKS.
  • Improve dashboards and alerts using New Relic or Datadog.
  • Investigate production issues, identify root causes, and implement durable fixes.
  • Participate in the shared 24/7 on-call rotation and contribute to effective incident response.
  • Collaborate with product engineers to troubleshoot performance and reliability issues throughout the application stack.
  • Improve application resilience using timeouts, retries, queuing, backpressure, and idempotency.
  • Maintain and improve CI/CD pipelines and deployment workflows (GitHub Actions, CircleCI).
  • Automate repetitive operational tasks and identify opportunities to reduce engineering toil.
  • Create and maintain runbooks, system diagrams, troubleshooting guides, and production documentation.
  • Contribute to capacity planning, performance testing, database reliability, and production-readiness reviews.
  • Apply security, access-control, logging, and compliance practices to infrastructure work.
  • Own small-to-medium reliability improvements from design to delivery.
  • Communicate progress, risks, and blockers clearly with engineering partners.

Skills

AWS
Terraform
Observability
CI/CD
Linux
On-call
Troubleshooting
Communication
SRE practices

Tools

New Relic
Datadog
CircleCI
GitHub Actions
Kubernetes
EKS
OpenSearch
PostgreSQL

Job description

GiveCampus is seeking a hands-on Site Reliability Engineer to improve the reliability, performance, and operational maturity of our platform. With our migration to AWS complete, you will focus on operating and strengthening production: observability, automation, incidents, and collaboration with product engineers to build resilient systems.

You will own reliability projects, automate infrastructure with Terraform, maintain Kubernetes/EKS, and enhance dashboards, alerts, and CI/CD pipelines.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Senior SRE: AWS, Kubernetes & Terraform
Remote Senior SRE: AWS, Kubernetes & Terraform

Motion Recruitment Partners LLC • Chicago (IL), Northern (KY)

Hybrid
USD 140,000 - 170,000
Senior SRE: Cloud Reliability, Terraform & Kubernetes (Remote)
Senior SRE: Cloud Reliability, Terraform & Kubernetes (Remote)

Motion Recruitment • Chicago (IL)

On-site
USD 140,000 - 190,000
Remote SRE: AWS, Kubernetes & Observability
Remote SRE: AWS, Kubernetes & Observability

Cloudbeds • United States

Remote
USD 120,000 - 150,000
Remote First
PTO
Home office stipend
+2
Site Reliability Engineer
Site Reliability Engineer

NextGen | GTA: A Kelly Telecom Company • Mount Laurel Township (NJ)

On-site
USD 110,000 - 170,000
Remote SRE Engineer — AWS, Kubernetes & CI/CD
Remote SRE Engineer — AWS, Kubernetes & CI/CD

ecosio • United States

Remote
USD 120,000 - 170,000
Remote-first culture
Flexible working hours
Personal development budget
+3
Senior SRE: Cloud, Kubernetes & Terraform (Remote)
Senior SRE: Cloud, Kubernetes & Terraform (Remote)

Motion Recruitment • Mount Laurel Township (NJ)

Remote
USD 130,000 - 180,000
Medical, dental, and vision benefits
Equity / Stock Options
Remote equipment stipend
+3
Site Reliability Engineer
Site Reliability Engineer

Motion Recruitment Partners LLC • Chicago (IL), Northern (KY)

On-site
USD 140,000 - 170,000
Remote Kubernetes SRE — Automation & Cloud Reliability
Remote Kubernetes SRE — Automation & Cloud Reliability

Eitacies Inc • Santa Clara (CA)

Remote
USD 110,000 - 160,000
Senior SRE - Kubernetes, Observability & Automation (Remote)
Senior SRE - Kubernetes, Observability & Automation (Remote)

Camunda • Atlanta (GA)

Remote
USD 150,000 - 242,000
Remote work
Annual company events
Health & wellbeing
+2
Senior SRE: Build Resilient AWS/Kubernetes Infrastructure
Senior SRE: Build Resilient AWS/Kubernetes Infrastructure

Kontakt.io • United States

On-site
USD 150,000 - 190,000