Site Reliability Engineer

Amplifire AI Learning

United States

On-site

USD 120,000 - 170,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Amplifire is seeking a Site Reliability Engineer to improve the reliability, scalability, performance, and operational efficiency of our cloud-based platform. You will work with DevOps within Platform Operations to elevate observability, automate workflows, and reduce incidents.

The role emphasizes collaboration with software engineering, QA, security, and DevOps, to embed reliability practices and enable safe, confident changes.

Qualifications

  • 4+ years of experience in Site Reliability Engineering, DevOps, cloud infrastructure, or related role.
  • Experience supporting production cloud environments.
  • Hands-on experience designing, deploying, and operating production workloads on AWS.
  • Experience with Infrastructure as Code, observability solutions, incident response, and CI/CD automation.
  • Experience reducing operational toil through automation and self-service tooling.

Responsibilities

  • Establish and maintain SLIs, SLOs, error budgets, and health metrics.
  • Build and improve monitoring, logging, tracing, dashboards, and alerts for visibility.
  • Respond to production incidents with urgency and sound judgment.
  • Develop runbooks, recovery procedures, and operational practices to improve resilience.
  • Embed reliability and observability throughout the development lifecycle with software teams.

Skills

SRE experience
AWS
CI/CD
Incident response
Automation
On-call
Monitoring & alerting
System observability
Problem solving

Tools

Docker
Kubernetes
Terraform
AWS CDK
Linux

Job description

Amplifire is seeking a Site Reliability Engineer to improve the reliability, scalability, performance, and operational efficiency of our cloud-based platform. Working alongside DevOps engineers within the Platform Operations team, this role combines software engineering and systems operations with a focus on observability, automation, incident reduction, and operational excellence.

You will partner closely with software engineering, QA, security, and DevOps to establish reliability practices, improve production visibility, strengthen incident response, automate operational workflows, and help engineering teams deliver changes safely and confidently.

This role supports systems operating under regulatory and compliance requirements, including FedRAMP and SOC 2. The ideal candidate understands that reliability, traceability, security, and change management enable sustainable development velocity rather than compete with it.

Amplifire expects all technical team members to leverage AI-assisted tools and workflows as force multipliers for productivity, learning, automation, and problem-solving while maintaining strong engineering skills, sound judgment, security standards, and operational accountability.

Reliability, Observability & Performance
  • Establish and maintain service-level indicators (SLIs), service-level objectives (SLOs), error budgets, and other measures of system health.
  • Build and continuously improve monitoring, logging, tracing, dashboards, and alerting that provide actionable visibility into application and infrastructure health.
  • Analyze system behavior, performance, capacity, and reliability trends to proactively identify risks and improvement opportunities.
  • Partner with engineering teams to define reliability requirements and improve the availability, scalability, and performance of production systems.
Incident Response & Operational Readiness
  • Participate in the on-call rotation and respond to production incidents with urgency and sound technical judgment.
  • Improve incident detection, triage, escalation, communication, mitigation, and recovery processes.
  • Lead or contribute to blameless post-incident reviews, root cause analysis, and corrective actions.
  • Develop and maintain runbooks, recovery procedures, and operational practices that improve service resilience and reduce mean time to detect and restore service.
Automation, Infrastructure & Delivery
  • Identify and eliminate operational toil through automation, self-service tooling, and continuous improvement.
  • Build and maintain cloud infrastructure using Infrastructure as Code practices and tools such as Terraform or AWS CDK.
  • Improve CI/CD pipelines, deployment safeguards, rollback capabilities, and progressive delivery practices.
  • Develop internal tools and automation that improve reliability, resilience, and engineering productivity.
  • Support scalable, secure, and cost-effective production environments.
  • Partner with software engineers to embed reliability, operability, and observability throughout the development lifecycle while reducing operational friction and helping teams safely own their services in production.
  • Help engineering teams diagnose complex production issues across application and infrastructure layers.
  • Contribute to security, compliance, capacity-planning, cloud cost optimization, and operational standards.
  • Use AI-assisted tools responsibly to improve troubleshooting, automation, documentation, and operational efficiency while sharing knowledge and continuously improving team practices.
Requirements
Required Experience
  • 4+ years of experience in Site Reliability Engineering, DevOps, cloud infrastructure, systems engineering, software engineering, or a related production-focused role.
  • Demonstrated experience supporting reliable, customer-facing applications in a production cloud environment.
  • Hands-on experience designing, deploying, and operating production workloads on AWS.
  • Experience building or maintaining Infrastructure as Code, observability solutions, incident response processes, and CI/CD automation.
  • Experience identifying and reducing operational toil through automation and continuous improvement.
  • Experience participating in on-call rotations, troubleshooting production issues, performing root cause analysis, and implementing preventive improvements.
Technical Skills
  • Working knowledge of Linux systems, networking, DNS, HTTP, load balancing, and common cloud architecture patterns.
  • Experience with containers and container-based deployment practices; Docker experience is required, while Kubernetes or similar orchestration experience is beneficial.
  • Ability to analyze logs, metrics, traces, and system behavior to troubleshoot issues across application and infrastructure layers.
  • Understanding of reliability concepts such as SLIs, SLOs, error budgets, availability, latency, capacity planning, and graceful degradation.
  • Familiarity with secure configuration, secrets management, access controls, vulnerability remediation, and other operational security fundamentals.
AI & Modern Tooling
  • Experience using AI-assisted development or operational tools to improve productivity, automation, troubleshooting, documentation, or engineering workflows.
  • Ability to critically evaluate AI-generated outputs and apply appropriate validation before production use.
  • Interest in adopting emerging tools and practices that improve engineering effectiveness while maintaining operational excellence.
Collaboration & Problem-Solving Skills
  • Strong debugging and systems-thinking skills, including the ability to work through ambiguous, cross-service production issues.
  • Ability to communicate clearly during incidents and translate technical findings for engineering and business stakeholders.
  • Experience collaborating with software engineering, QA, security, support, and product teams.
  • Ability to independently own reliability improvements while seeking input and alignment when appropriate.
  • A proactive approach to problem-solving and a willingness to challenge existing practices constructively.
What Success Looks Like

Within your first year:

  • Service-level indicators and objectives are established for critical systems.
  • Monitoring, alerting, and observability provide actionable visibility across the platform.
  • Incident response processes are more structured, repeatable, and measurable.
  • Mean time to detect and restore service trends improve through better tooling and operational practices.
  • Engineering teams have greater self-service access to operational insights and reliability tooling.
  • Manual operational work is reduced through automation and platform improvements.
  • Reliability, performance, and scalability risks are identified proactively rather than reactively.
  • Demonstrates a proactive approach to problem solving and does not accept existing processes simply because they have historically been done that way.

Amplifire is the leading AI Learning Platform built on brain science that delivers the proven results of 1:1 expert instruction at enterprise scale. We deliver one-on-one AI instruction at scale, detect where people are confidently wrong, and fix it before it becomes a mistake - reducing training time by 50-80% while driving measurable performance outcomes. Trusted by leading organizations in healthcare, accounting, life sciences and other high-stakes industries, Amplifire enables teams to achieve verified competency faster, reduce risk, and perform at the highest level when it matters most.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Site Reliability Engineer: Observability & Automation
Cloud Site Reliability Engineer: Observability & Automation

Amplifire AI Learning • United States

On-site
USD 120,000 - 170,000
Sr. QA Automation Engineer
Sr. QA Automation Engineer

Amplifire • Boulder (CO)

On-site
USD 120,000 - 160,000
Sr. Manager, Product Marketing
Sr. Manager, Product Marketing

Amplifire • Denver (CO)

On-site
USD 120,000 - 140,000
Health insurance
Dental insurance
Vision insurance
+6
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Analytic Partners • Dallas (TX)

On-site
USD 100,000 - 130,000
Competitive salary
Opportunities for professional growth
Supportive work environment focused on DEI
Senior Forward Deployed Engineer (DevOps/SRE)
Senior Forward Deployed Engineer (DevOps/SRE)

LeoForce • Pleasanton (CA)

On-site
USD 300,000 - 350,000
Medical benefits
401(k) plan
Free meals and snacks
+2
Site Reliability Engineer
Site Reliability Engineer

Instrumental Inc. • Palo Alto (CA)

On-site
USD 140,000 - 165,000
Health insurance
Vision insurance
Dental plan
+2
Site Reliability Engineer
Site Reliability Engineer

FLUIX • Palo Alto (CA)

On-site
USD 120,000 - 150,000
Attractive compensation package including equity options
Comprehensive health, dental, and vision insurance
Opportunities for professional growth
3510- Site Reliability Engineer II
3510- Site Reliability Engineer II

Innovaccer • Dallas (TX)

On-site
USD 110,000 - 140,000
Generous Paid Time Off: 22 days per year plus company holidays
Best-in-Class Parental Leave
Comprehensive insurance coverage
Site Reliability Engineer
Site Reliability Engineer

Inclusion Services S.A • Chicago (IL)

On-site
USD 90,000 - 130,000
100% company-covered health insurance
401k plan with 4% match
15 days paid time off
+3
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 140,000 - 210,000