AI Site Reliability Engineer

Remote Jobs

United States

Remote

USD 101,000 - 188,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical, Dental, Vision
401(k) match
Paid time off

Job summary

The AI Factory is seeking a Site Reliability Engineer to enhance the reliability of platforms and services that support AI development and deployment. You will leverage software development and automation to identify operational problems, reduce repetitive work, and help teams deliver dependable systems.

You will collaborate with engineers and partners across the AI Factory to improve observability, incident response, release validation, and platform health, contributing to reliability platforms

Qualifications

  • Experience designing, developing, and maintaining production software or automation using Go, Python, or a comparable language.
  • Experience operating or engineering Kubernetes-based platforms, including troubleshooting complex service or infrastructure issues.
  • Experience building or improving CI/CD, GitOps, infrastructure automation, or deployment workflows.
  • Experience using observability data—including metrics, logs, or traces—to diagnose problems and improve system health.

Responsibilities

  • Improve the scalability, resilience, and reliability of existing platform services and middleware to ensure they remain dependable as usage and demand grow.
  • Develop and maintain automation and operational tooling that improve platform reliability and reduce recurring manual work.
  • Help teams investigate incidents, identify contributing factors, and implement fixes that prevent repeat issues.
  • Build and improve dashboards, alerts, and other observability capabilities using metrics, logs, and traces.
  • Create automated tests and validation workflows for upgrades, releases, and changes to platform services.
  • Contribute to CI/CD and GitOps workflows that support consistent, reliable deployments.
  • Assess platform health, document findings, and work with partner teams on practical reliability improvements.
  • Participate in design reviews, code reviews, testing, and incident reviews.
  • Contribute to reliability improvements for AI Factory services, including AIF Up and tools that support health validation and investigation.

Skills

Go/Python automation
Kubernetes platforms
CI/CD/GitOps
Observability data usage
OpenShift
GitLab CI/CD
Argo CD
Argo Rollouts
Prometheus
Grafana
OpenTelemetry

Tools

OpenShift
GitLab CI/CD
Argo CD
Argo Rollouts
Prometheus
Grafana
OpenTelemetry

Job description

Your Mission:

The AI Factory team is seeking a Site Reliability Engineer to improve the reliability of the platforms and services that support AI development and deployment. In this role, you will use software development and automation to identify operational problems, reduce repetitive work, and help teams deliver dependable systems.

You will work with engineers and other partners across the AI Factory to improve observability, incident response, release validation, and platform health. The role is broader than any single tool or product: you may contribute to reliability platforms such as Cluster Concierge, but your focus will be on solving reliability problems across the environment.

Key Responsibilities:
  • Improve the scalability, resilience, and reliability of existing platform services and middleware to ensure they remain dependable as usage and demand grow.
  • Develop and maintain automation and operational tooling that improve platform reliability and reduce recurring manual work.
  • Help teams investigate incidents, identify contributing factors, and implement fixes that prevent repeat issues.
  • Build and improve dashboards, alerts, and other observability capabilities using metrics, logs, and traces.
  • Create automated tests and validation workflows for upgrades, releases, and changes to platform services.
  • Contribute to CI/CD and GitOps workflows that support consistent, reliable deployments.
  • Assess platform health, document findings, and work with partner teams on practical reliability improvements.
  • Participate in design reviews, code reviews, testing, and incident reviews.
  • Contribute to reliability improvements for AI Factory services, including AIF Up and tools that support health validation and investigation.

Responsible for autonomy hardware and software system integration and testing including verification and validation.Translates customer requirements into product and systems specifications while addressing technical, schedule, and cost considerations; Establishes functional and technical specifications and standards for autonomous systems; Determines sensing hardware components; Recommends and selects appropriate control systems; Integrates and optimizes the autonomous compute and sensing hardware and software; Solves hardware/software interface problems; Develops plan(s) to integrate autonomous functionality into product(s) and platform(s) and other system(s); Develops test plans for validation and verification and procedures for test and evaluation requirements in collaboration with designers and developers; Performs integration testing and coordinates subsystem and/or system testing activities for autonomous programs; Documents and conducts analysis of test results and recommends fixes to software, hardware components, subsystems and systems; Interfaces with other teams involved the development lifecycle for perception

Basic Qualifications
  • Experience designing, developing, and maintaining production software or automation using Go, Python, or a comparable language.
  • Experience operating or engineering Kubernetes-based platforms, including troubleshooting complex service or infrastructure issues.
  • Experience building or improving CI/CD, GitOps, infrastructure automation, or deployment workflows.
  • Experience using observability data—including metrics, logs, or traces—to diagnose problems and improve system health.
Desired Skills
  • Experience with OpenShift, GitLab CI/CD, Argo CD, Argo Rollouts, or similar GitOps and progressive-delivery tooling.
  • Experience improving incident response, reducing operational toil, or defining actionable service-health measures.
  • Familiarity with Prometheus, Grafana, OpenTelemetry, or comparable observability tools.
  • Familiarity with designing automated reliability tests, upgrade validation, resilience tests, or failure-mode analyses.
  • Familiarity with AI/ML platforms, GPU-based infrastructure, or deployments in disconnected environments.
  • Strong oral and written communication skills, and ability to collaborate with cross-functional partners
  • Creative and resourceful when it comes to problem-solving
  • Ability to work with internal stakeholders to collect feedback, prioritize tasks, and manage the engineering backlog
  • Self-motivated, self-directed, and the ability to thrive in a fast-paced environment in an industry that constantly changes
Pay Information

GeoZone Definition: GeoZones are geographic groupings created by Lockheed Martin to align compensation ranges with regional labor markets and cost-of-labor differences across the United States. Locations are assigned a Geo Zone based on the primary work location of the role.

  • Full-time salary range (GEOZONE 1): $101400.00 - $188200.00 - Includes metropolitan areas such as Sunnyvale CA; Pal Alto, CA; New York City metropolitan area; Newark, New Jersey; etc.
  • Full-time salary range (GEOZONE 2): $91200.00 - $169400.00 - Includes metropolitan areas such as Denver, CO; King of Prussia, PA; Stratford, CT; Moorestown, NJ; etc.
  • Full-time salary range (GEOZONE 3): $81100.00 - $150500.00- Includes metropolitan areas such as Dallas–Fort Worth, TX; Orlando, FL; Grand Prairie, TX; Marietta, GA; etc.
  • Full-time salary range (GEOZONE 4): $72900.00 - $135500.00- Includes metropolitan areas such as Camden, AR; Lexington, KY; Lufkin, TX; etc.

At Lockheed Martin, we know mission success starts with taking care of our people. Our Total Rewards program is designed to attract top talent, support your well-being, and help you grow—both professionally and personally.

The salary range for this position is as listed on the requisition. Please note that the salary information listed is a general guideline only. Lockheed Martin considers factors such as (but not limited to) scope and responsibilities of the position, candidate's work experience, education/ training, key skills as well as market(work location) and business considerations when extending an offer.

Benefits offered:
  • Medical, Dental, Vision, Flexible work arrangements and schedules (e.g., 4x10), 401(k) match, Paid time off, Holidays, Parental Leave, EAP, Flexible Spending Accounts, Education Assistance, Life Insurance, Short-Term Disability, and Long-Term Disability.
  • Annual short-term and/or long-term incentive compensation programs may be offered depending on the position. Payments under these annual programs are not guaranteed and can vary from year to year and are tied to a range of performance metrics.
  • For (Washington state applicants only) Non-represented full-time employees: accrue at least 10 hours per month of Paid Time Off (PTO) to be used for incidental absences and other reasons; receive at least 90 hours for holidays. Represented full time employees accrue 6.67 hours of Vacation per month; accrue up to 52 hours of sick leave annually; receive at least 96 hours for holidays. PTO, Vacation, sick leave, and holiday hours are prorated based on start date during the calendar year.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

A/AI Research Engineer
A/AI Research Engineer

Lockheed Martin • Town of Texas (WI)

On-site
USD 101,000 - 188,000
Medical coverage
Dental coverage
Vision coverage
+1
A/AI Research Engineer Sr Mgr - L6
A/AI Research Engineer Sr Mgr - L6

Lockheed Martin • Palo Alto (CA)

On-site
USD 228,000 - 424,000
Medical, Dental, Vision
401(k) match
PTO & Holidays
+2
Transformation & Enablement Associate Manager
Transformation & Enablement Associate Manager

Lockheed Martin • Fort Worth (TX)

On-site
USD 105,000 - 195,000
Medical, Dental, Vision
Flexible work arrangements
401(k) match
+6
Radical Interoperability Capability Development Director
Radical Interoperability Capability Development Director

Lockheed Martin • Arizona

On-site
USD 180,000 - 325,000
Annual incentive programs
Paid time off and holidays
Vacation and sick leave policy
Engineering Aide - N1
Engineering Aide - N1

Lockheed Martin • United States

Remote
USD 42,000 - 97,000
Software Engineer Sr - E3
Software Engineer Sr - E3

Lockheed Martin • King of Prussia (PA)

On-site
USD 110,000 - 205,000
Medical
Dental
Vision
+11
Project Engineer Sr
Project Engineer Sr

Lockheed Martin • Highlands Ranch (CO)

On-site
USD 96,000 - 178,000
Medical insurance
Dental & Vision
Flexible schedules
+4
Project Engineer Sr Stf - E5
Project Engineer Sr Stf - E5

Remote Jobs • United States

Remote
USD 140,000 - 230,000
Medical, Dental, Vision
Flexible work arrangements
Life Insurance
Employment Law Compliance Counsel
Employment Law Compliance Counsel

Lockheed Martin • Fort Worth (TX)

On-site
USD 148,000 - 276,000
Metric & Data Analyst Sr - E3 Lockheed Martin · Denver, CO Full-time · On-site — 1 hour ago
Metric & Data Analyst Sr - E3 Lockheed Martin · Denver, CO Full-time · On-site — 1 hour ago

Emploive • Denver (CO)

Hybrid
USD 98,000 - 170,000
Medical Insurance
401(k) match
Paid time off
+2