Sr. SRE – AI Platforms

Mainz Brady Group

Portland (OR)

On-site

USD 140,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mainz Brady Group in Portland, OR is seeking a Senior Site Reliability Engineer to strengthen AI platforms and infrastructure reliability. You will troubleshoot production systems, design resilience improvements, and lead automation across services and clusters.

The role emphasizes building observability with metrics, logs, traces, and alerting, plus coordinating with infra and validation teams for readiness and safe production changes. Strong SRE experience and collaboration are essential.

Qualifications

  • 7+ years in SRE or reliability-focused roles.
  • Hands-on with Linux, Kubernetes, networking and distributed systems.
  • Experience with incident management and root cause analysis.

Responsibilities

  • Operate and improve reliability of AI platform services, cluster dependencies, and shared infrastructure.
  • Lead incident triage involving Kubernetes, Linux, storage, networking, scheduling and dependencies.
  • Define and improve SLIs, SLOs, alert thresholds, runbooks, escalation paths, and post-incident actions.
  • Automate manual processes and reduce production support overhead.
  • Build observability with metrics, logs, traces, and event correlation.
  • Troubleshoot performance and availability affecting AI/ML workloads.
  • Collaborate with infra and validation teams on production readiness and change management.
  • Drive operational reviews and readiness assessments.

Skills

Linux
Kubernetes
Networking
Distributed systems
Observability
Incident management
Scripting

Tools

Prometheus
Grafana
ELK Stack
OpenSearch
Loki
PagerDuty

Job description

Senior Site Reliability Engineer (SRE) - AI Platforms

We’re looking for a Senior SRE, AI Platforms to provide Site Reliability Engineering support for AI platforms and infrastructure environments. This role is focused on platform reliability, production troubleshooting, incident response, observability, and operational automation. The ideal candidate is a senior, hands-on engineer who can troubleshoot complex production environments while driving long-term improvements in platform availability and resilience.

What You’ll Be Doing
  • Operate and improve the reliability of AI platform services, cluster dependencies, and shared infrastructure.
  • Lead and support incident triage involving Kubernetes, Linux, storage, networking, scheduling platforms, job orchestration, and dependency failures.
  • Define and improve SLIs, SLOs, alerting thresholds, runbooks, escalation paths, and post-incident corrective actions.
  • Analyze recurring failures and turn manual operational processes into automation and preventative controls.
  • Build observability across systems and services using metrics, logs, traces, and event correlation.
  • Troubleshoot performance and availability issues impacting AI/ML training workloads, inference services, and internal platforms.
  • Partner with infrastructure and validation teams to improve production readiness and change-management safety.
  • Drive operational reviews, resilience testing, and readiness assessments.
What We’re Looking For
  • 7+ years of SRE, production operations, or reliability-focused infrastructure engineering experience.
  • Strong hands-on expertise with Linux, Kubernetes, networking, and distributed systems.
  • Experience building observability, monitoring, alerting, and incident-response workflows.
  • Strong scripting and automation experience with Python, Bash, Go, or similar technologies.
  • Experience with incident management, root cause analysis, and post-incident remediation.
  • Ability to automate repetitive operational processes and reduce production support overhead.
  • Strong troubleshooting skills across complex, highly available production environments.
  • Excellent communication and ability to work across technical teams.
Preferred Experience
  • AI platforms, machine learning infrastructure, or large-scale HPC environments.
  • Prometheus, Grafana, ELK Stack, OpenSearch, Loki, and PagerDuty.
  • Defining and managing SLIs, SLOs, and error budgets.
  • Supporting both batch and service-based workloads.
  • Cloud and data center infrastructure.
  • Logging, tracing, infrastructure monitoring, and service reliability tooling.
About Mainz Brady Group

Mainz Brady Group is a technology staffing firm with offices in California, Oregon, Washington, and Texas. We specialize in Information Technology and Engineering placements on a Contract, Contract-to-Hire, and Direct Hire basis. Mainz Brady Group is a recipient of multiple annual Excellence Awards from the TechServe Alliance, the leading association for IT and engineering staffing firms in the U.S.

Mainz Brady Group is an Equal Opportunity Employer. We are committed to Diversity & Inclusion and incorporate non-discrimination best practices in all our staffing processes.

Mainz Brady Group does not discriminate based on race, color, religion, sex, sexual orientation, gender identity, gender expression, age, disability, or any other protected class.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE — AI Platforms (Reliability & Automation)
Senior SRE — AI Platforms (Reliability & Automation)

Mainz Brady Group • Portland (OR)

On-site
USD 140,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Greenwood Village (CO)

On-site
USD 120,000 - 150,000
Senior Site Reliability Engineer – AI & Automation.
Senior Site Reliability Engineer – AI & Automation.

Veriipro • Miami (FL)

On-site
USD 130,000 - 170,000
Senior AI-Enabled Platform / SRE Engineer
Senior AI-Enabled Platform / SRE Engineer

ZipStaff Inc. • Dallas (TX)

Hybrid
USD 124,000 - 207,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Optomi • Seattle (WA)

On-site
USD 140,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Optomi • New York (NY)

On-site
USD 140,000 - 200,000