Site Reliability Engineer

Optomi

Seattle (WA)

On-site

USD 140,000 - 180,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Optomi, in partnership with a leading Entertainment industry client, is seeking an experienced SRE to support multiple critical platforms from Seattle/Orlando/New York/Burbank. The role centers on Kubernetes, AI/ML integration, and automation across distributed systems.

You will design, operate, and improve scalable infrastructure with strong focus on reliability, monitoring, and observability, while collaborating across engineering, architecture, and security teams.

Qualifications

  • 5+ years of experience in SRE, software engineering, platform/infrastructure engineering, systems administration, or related fields.
  • Deep, hands-on Kubernetes expertise is required, including operations, monitoring, capacity planning, troubleshooting, and performance management.
  • Strong Linux administration and distributed systems experience.
  • Experience with cloud platforms, particularly AWS, and container technologies such as Docker.
  • Strong infrastructure-as-code experience with Terraform, OpenTofu, or similar tools.
  • Hands-on experience with CI/CD pipelines and source control systems such as GitLab or GitHub Actions.
  • Experience with monitoring, logging, and observability platforms such as Datadog, Splunk, or CloudWatch.
  • Strong understanding of networking fundamentals, APIs, and infrastructure automation.
  • Experience building reusable tools, platforms, libraries, or modules used by multiple engineering teams.
  • Experience operating in large enterprise environments with cross-team collaboration and competing priorities.
  • Strong written communication skills and the ability to produce clear technical documentation and architecture proposals.
  • Demonstrated ability to influence technical direction and serve as a thought leader within SRE/platform engineering.
  • Experience with AI-assisted development tools such as Cursor, Claude Code, or GitHub Copilot is preferred.
  • Experience integrating AI/ML services into engineering or CI/CD workflows is a plus.

Responsibilities

  • Support and evolve critical platforms, including an AI platform, anexisting automation platform, and a new data and observability platform.
  • Provide technical leadership across SRE, infrastructure, platform reliability, and automation initiatives.
  • Design, operate, and improve scalable, highly available infrastructure and distributed systems, with a strong focus on Kubernetes.
  • Monitor system health, capacity, performance, SLAs/SLOs, and reliability; troubleshoot complex infrastructure and application issues.
  • Build and maintain infrastructure-as-code, automation, and CI/CD capabilities using tools such as Terraform/OpenTofu, Ansible, and GitLab CI/CD.
  • Develop reusable platform components, modules, and developer-facing tools that enable engineering teams to work more efficiently.
  • Implement observability, security scanning, monitoring, instrumentation, and telemetry across platforms and applications.
  • Partner across engineering, architecture, and security teams to establish technical standards and drive platform strategy.
  • Leverage AI-assisted development and AI/ML services to improve automation, code quality, documentation, and engineering workflows.

Skills

Kubernetes
Linux administration
AWS cloud
Docker
Terraform/OpenTofu
GitLab/GitHub Actions
Observability
Networking basics
SRE / Platform engineering
CI/CD pipelines

Tools

Datadog
Splunk
CloudWatch
OpenTofu
GitLab
GitHub Actions

Job description

Optomi, in partnership with a leading client in the Entertainment industry, is seeking an experienced SRE for their, Orlando, Seattle, New York, OR Burbank location. The right candidate will have Kubernetes expertise, with experience in Monitoring, capacity planning, and SLA / SLO's. The SRE will support multiple critical platforms, including a new AI platform, and an existing automation platform.

Responsibilities
  • Support and evolve critical platforms, including an AI platform, anexisting automation platform, and a new data and observability platform.
  • Provide technical leadership across SRE, infrastructure, platform reliability, and automation initiatives.
  • Design, operate, and improve scalable, highly available infrastructure and distributed systems, with a strong focus on Kubernetes.
  • Monitor system health, capacity, performance, SLAs/SLOs, and reliability; troubleshoot complex infrastructure and application issues.
  • Build and maintain infrastructure-as-code, automation, and CI/CD capabilities using tools such as Terraform/OpenTofu, Ansible, and GitLab CI/CD.
  • Develop reusable platform components, modules, and developer-facing tools that enable engineering teams to work more efficiently.
  • Implement observability, security scanning, monitoring, instrumentation, and telemetry across platforms and applications.
  • Partner across engineering, architecture, and security teams to establish technical standards and drive platform strategy.
  • Leverage AI-assisted development and AI/ML services to improve automation, code quality, documentation, and engineering workflows.
Qualifications
  • 5+ years of experience in SRE, software engineering, platform/infrastructure engineering, systems administration, or related fields.
  • Deep, hands-on Kubernetes expertise is required, including operations, monitoring, capacity planning, troubleshooting, and performance management.
  • Strong Linux administration and distributed systems experience.
  • Experience with cloud platforms, particularly AWS, and container technologies such as Docker.
  • Strong infrastructure-as-code experience with Terraform, OpenTofu, or similar tools.
  • Hands-on experience with CI/CD pipelines and source control systems such as GitLab or GitHub Actions.
  • Experience with monitoring, logging, and observability platforms such as Datadog, Splunk, or CloudWatch.
  • Strong understanding of networking fundamentals, APIs, and infrastructure automation.
  • Experience building reusable tools, platforms, libraries, or modules used by multiple engineering teams.
  • Experience operating in large enterprise environments with cross-team collaboration and competing priorities.
  • Strong written communication skills and the ability to produce clear technical documentation and architecture proposals.
  • Demonstrated ability to influence technical direction and serve as a thought leader within SRE/platform engineering.
  • Experience with AI-assisted development tools such as Cursor, Claude Code, or GitHub Copilot is preferred.
  • Experience integrating AI/ML services into engineering or CI/CD workflows is a plus.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer - Kubernetes & AI Platform
Senior Site Reliability Engineer - Kubernetes & AI Platform

Optomi • Seattle (WA)

On-site
USD 140,000 - 180,000
Sr. Director, Site Reliability and Platform Engineering
Sr. Director, Site Reliability and Platform Engineering

Optomi • Tacoma (WA)

On-site
USD 150,000 - 200,000
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Senior SRE (Contract/Hybrid)
Senior SRE (Contract/Hybrid)

Optomi • Orlando (FL)

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Greenwood Village (CO)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Associate Engineer, Site Reliability
Associate Engineer, Site Reliability

R&D • United States

On-site
USD 90,000 - 140,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Optomi • Orlando (FL)

Hybrid
USD 140,000 - 200,000