Senior Site Reliability Engineer - Kubernetes & AI Platform

Optomi

Seattle (WA)

On-site

USD 140,000 - 180,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Optomi, in partnership with a leading Entertainment industry client, is seeking an experienced SRE to support multiple critical platforms from Seattle/Orlando/New York/Burbank. The role centers on Kubernetes, AI/ML integration, and automation across distributed systems.

You will design, operate, and improve scalable infrastructure with strong focus on reliability, monitoring, and observability, while collaborating across engineering, architecture, and security teams.

Qualifications

  • 5+ years of experience in SRE, software engineering, platform/infrastructure engineering, systems administration, or related fields.
  • Deep, hands-on Kubernetes expertise is required, including operations, monitoring, capacity planning, troubleshooting, and performance management.
  • Strong Linux administration and distributed systems experience.
  • Experience with cloud platforms, particularly AWS, and container technologies such as Docker.
  • Strong infrastructure-as-code experience with Terraform, OpenTofu, or similar tools.
  • Hands-on experience with CI/CD pipelines and source control systems such as GitLab or GitHub Actions.
  • Experience with monitoring, logging, and observability platforms such as Datadog, Splunk, or CloudWatch.
  • Strong understanding of networking fundamentals, APIs, and infrastructure automation.
  • Experience building reusable tools, platforms, libraries, or modules used by multiple engineering teams.
  • Experience operating in large enterprise environments with cross-team collaboration and competing priorities.
  • Strong written communication skills and the ability to produce clear technical documentation and architecture proposals.
  • Demonstrated ability to influence technical direction and serve as a thought leader within SRE/platform engineering.
  • Experience with AI-assisted development tools such as Cursor, Claude Code, or GitHub Copilot is preferred.
  • Experience integrating AI/ML services into engineering or CI/CD workflows is a plus.

Responsibilities

  • Support and evolve critical platforms, including an AI platform, anexisting automation platform, and a new data and observability platform.
  • Provide technical leadership across SRE, infrastructure, platform reliability, and automation initiatives.
  • Design, operate, and improve scalable, highly available infrastructure and distributed systems, with a strong focus on Kubernetes.
  • Monitor system health, capacity, performance, SLAs/SLOs, and reliability; troubleshoot complex infrastructure and application issues.
  • Build and maintain infrastructure-as-code, automation, and CI/CD capabilities using tools such as Terraform/OpenTofu, Ansible, and GitLab CI/CD.
  • Develop reusable platform components, modules, and developer-facing tools that enable engineering teams to work more efficiently.
  • Implement observability, security scanning, monitoring, instrumentation, and telemetry across platforms and applications.
  • Partner across engineering, architecture, and security teams to establish technical standards and drive platform strategy.
  • Leverage AI-assisted development and AI/ML services to improve automation, code quality, documentation, and engineering workflows.

Skills

Kubernetes
Linux administration
AWS cloud
Docker
Terraform/OpenTofu
GitLab/GitHub Actions
Observability
Networking basics
SRE / Platform engineering
CI/CD pipelines

Tools

Datadog
Splunk
CloudWatch
OpenTofu
GitLab
GitHub Actions

Job description

Optomi, in partnership with a leading Entertainment industry client, is seeking an experienced SRE to support multiple critical platforms from Seattle/Orlando/New York/Burbank. The role centers on Kubernetes, AI/ML integration, and automation across distributed systems.

You will design, operate, and improve scalable infrastructure with strong focus on reliability, monitoring, and observability, while collaborating across engineering, architecture, and security teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Lead - Kubernetes, AI Platforms & Observability
SRE Lead - Kubernetes, AI Platforms & Observability

Optomi • New York (NY)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Optomi • New York (NY)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Optomi • Seattle (WA)

On-site
USD 140,000 - 180,000
Senior Platform Reliability Engineer (Kubernetes & CI/CD)
Senior Platform Reliability Engineer (Kubernetes & CI/CD)

Optomi • Orlando (FL)

Hybrid
USD 120,000 - 180,000
Senior Site Reliability Engineer – AI-Driven Reliability
Senior Site Reliability Engineer – AI-Driven Reliability

Optomi • United States

On-site
USD 120,000 - 180,000
Lead SRE for Generative AI Platform (Multi-Cloud)
Lead SRE for Generative AI Platform (Multi-Cloud)

Optomi • Orlando (FL)

Hybrid
USD 140,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Optomi • Dallas (TX)

Hybrid
USD 120,000 - 150,000
Senior SRE (Contract/Hybrid)
Senior SRE (Contract/Hybrid)

Optomi • Orlando (FL)

Hybrid
USD 120,000 - 180,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Optomi • Orlando (FL)

Hybrid
USD 140,000 - 200,000
Lead SRE, Generative AI Platform (Remote)
Lead SRE, Generative AI Platform (Remote)

Optomi • United States

On-site
USD 150,000 - 190,000
Remote flexibility
Cutting-edge AI platform experience
Mentoring engineers and shaping cloud架