Infra Reliability SDE: AI-Powered Incident Platform

Amazon

Nashville (TN)

On-site

USD 137,000 - 185,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical, Dental, and Vision Coverage
Maternity and Parental Leave Options
Paid Time Off (PTO)
401(k) Plan

Job summary

Amazon.com Services LLC in Nashville, TN is hiring Reflex, the agentic incident-response platform for Amazon's fulfillment network. You'll design and ship production software and AI agents on Amazon Bedrock AgentCore to triage high-severity incidents and generate real-time call intelligence.

Join a software team in Operations Infrastructure Services, stay close to live operations, and help mature Reflex toward shared incident context across organizations with guarded automation of recovery and

Qualifications

  • 3+ years of non-internship professional software development experience.
  • 2+ years of non-internship design or architecture of new and existing systems experience
  • 1+ years of software development engineer or related occupational experience
  • 1+ years of designing and developing large-scale, multi-tiered, multi-threaded, embedded or distributed software applications, tools, systems, and services using languages such as C#, C++, Java, or Perl
  • 1+ years of Object Oriented Design experience
  • Bachelor's degree or foreign equivalent in Computer Science, Engineering, Mathematics, or a related field
  • Experience programming with at least one software programming language

Responsibilities

  • Design, build, test, deploy, and operate production services and AI agents on AWS (Amazon Bedrock AgentCore, serverless compute, event-driven pipelines) that automate incident triage, call intelligence, communications, and post-incident documentation and reporting
  • Own features end-to-end: from discovery with Incident Managers and resolver teams, through design, implementation, evaluation, deployment, and production operation
  • Build the foundations that gate agent autonomy: LLM output evaluation, observability and alerting for agents in production, and identity and access controls aligned with Amazon standards
  • Design the feedback loops through which agents learn: capturing human reviews, corrections, approvals, and incident outcomes as evaluation signal, and turning resolved incidents into structured history that improves recommendations over time
  • Integrate with the incident lifecycle (ticketing, chat, telemetry, detection feeds, and live call transcription) and model consistent incident state across those systems
  • Raise the bar on software quality, security, testing, and operational excellence for AI systems acting inside production incident workflows

Skills

3+ years software development
OO design
multi-tier distributed

Education

Bachelor's degree in CS or related field

Tools

Amazon Bedrock
LLM basics

Job description

Amazon.com Services LLC in Nashville, TN is hiring Reflex, the agentic incident-response platform for Amazon's fulfillment network. You'll design and ship production software and AI agents on Amazon Bedrock AgentCore to triage high-severity incidents and generate real-time call intelligence.

Join a software team in Operations Infrastructure Services, stay close to live operations, and help mature Reflex toward shared incident context across organizations with guarded automation of recovery and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Incident-Response Software Engineer
AI Incident-Response Software Engineer

Amazon • Arlington (VA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
Software Development Engineer, Infrastructure Reliability Engineering
Software Development Engineer, Infrastructure Reliability Engineering

Amazon • Nashville (TN)

On-site
USD 137,000 - 185,000
Medical, Dental, and Vision Coverage
Maternity and Parental Leave Options
Paid Time Off (PTO)
+1
Software Development Engineer, Infrastructure Reliability Engineering
Software Development Engineer, Infrastructure Reliability Engineering

Amazon • Arlington (VA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
Senior SDE: AI-Powered Resilience & Incident Prevention
Senior SDE: AI-Powered Resilience & Incident Prevention

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
SDE II: Real-Time AI Incident Response
SDE II: Real-Time AI Incident Response

Amazon • Seattle (WA)

On-site
USD 144,000 - 194,000
Senior SDE, AWS Resilience & Incident Prevention (AI)
Senior SDE, AWS Resilience & Incident Prevention (AI)

Amazon • Seattle (WA)

On-site
USD 168,000 - 227,000
Health insurance
401(k) matching
Paid time off
+1
Lead SDE - AI-Driven Incident Control
Lead SDE - AI-Driven Incident Control

Amazon • Seattle (WA)

On-site
USD 168,000 - 227,000
RSUs
Health insurance
401(k) matching
+1
AI-Driven DevOps Engineer for Incident Resolution
AI-Driven DevOps Engineer for Incident Resolution

Amazon • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
RSUs
Platform Engineer, No-Code AI Agents for Fulfillment
Platform Engineer, No-Code AI Agents for Fulfillment

Amazon.com Services LLC • Bellevue (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+2
Senior AI-Driven DevOps Engineer
Senior AI-Driven DevOps Engineer

Amazon • Seattle (WA)

On-site
USD 168,000 - 227,000
Health insurance
RSUs
Parental leave
+1