Site Reliability Engineer

Optomi

New York (NY)

On-site

USD 140,000 - 200,000

Full time

42 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Optomi, in partnership with a leading client in the Entertainment industry, seeks an experienced SRE for locations in Orlando, Seattle, New York, or Burbank. The right candidate will align Kubernetes operations with monitoring, capacity planning, and SLA/SLO management to support multiple critical platforms including an AI platform.

You will design scalable infrastructure, implement IaC and CI/CD pipelines, and collaborate across engineering, architecture, and security teams to drive platform

Qualifications

  • 5+ years of SRE, systems/infrastructure experience.
  • Deep Kubernetes ops, monitoring, capacity planning, troubleshooting.
  • Experience with cloud platforms (AWS) and containers.

Responsibilities

  • Support and evolve critical platforms, including an AI platform and automation platform.
  • Provide technical leadership across SRE, infra, platform reliability, and automation.
  • Design, operate and improve scalable, highly available infrastructure and distributed systems focusing on Kubernetes.
  • Monitor system health, capacity, performance, SLAs/SLOs, and reliability; troubleshoot complex issues.
  • Build infrastructure-as-code and CI/CD capabilities using Terraform/OpenTofu, Ansible, and GitLab CI/CD.
  • Develop reusable platform components and developer-facing tools to improve efficiency.
  • Implement observability, security scanning, monitoring, instrumentation, and telemetry.
  • Partner with engineering, architecture, and security teams to set technical standards and drive platform strategy.
  • Leverage AI-assisted development to enhance automation and workflows.

Skills

Kubernetes
Linux
AWS
Terraform/OpenTofu
CI/CD (GitLab/GitHub Actions)
Observability (Datadog/Splunk/CloudW)
Networking fundamentals
AI/ML tooling (Cursor/Claude/ Copilot)

Tools

Docker

Job description

Optomi, in partnership with a leading client in the Entertainment industry, is seeking an experienced SRE for their, Orlando, Seattle, New York, OR Burbank location. The right candidate will have Kubernetes expertise, with experience in Monitoring, capacity planning, and SLA / SLO's. The SRE will support multiple critical platforms, including a new AI platform, and an existing automation platform.

Responsibilities
  • Support and evolve critical platforms, including an AI platform, anexisting automation platform, and a new data and observability platform.
  • Provide technical leadership across SRE, infrastructure, platform reliability, and automation initiatives.
  • Design, operate, and improve scalable, highly available infrastructure and distributed systems, with a strong focus on Kubernetes.
  • Monitor system health, capacity, performance, SLAs/SLOs, and reliability; troubleshoot complex infrastructure and application issues.
  • Build and maintain infrastructure-as-code, automation, and CI/CD capabilities using tools such as Terraform/OpenTofu, Ansible, and GitLab CI/CD.
  • Develop reusable platform components, modules, and developer-facing tools that enable engineering teams to work more efficiently.
  • Implement observability, security scanning, monitoring, instrumentation, and telemetry across platforms and applications.
  • Partner across engineering, architecture, and security teams to establish technical standards and drive platform strategy.
  • Leverage AI-assisted development and AI/ML services to improve automation, code quality, documentation, and engineering workflows.
Qualifications
  • 5+ years of experience in SRE, software engineering, platform/infrastructure engineering, systems administration, or related fields.
  • Deep, hands-on Kubernetes expertise is required, including operations, monitoring, capacity planning, troubleshooting, and performance management.
  • Strong Linux administration and distributed systems experience.
  • Experience with cloud platforms, particularly AWS, and container technologies such as Docker.
  • Strong infrastructure-as-code experience with Terraform, OpenTofu, or similar tools.
  • Hands-on experience with CI/CD pipelines and source control systems such as GitLab or GitHub Actions.
  • Experience with monitoring, logging, and observability platforms such as Datadog, Splunk, or CloudWatch.
  • Strong understanding of networking fundamentals, APIs, and infrastructure automation.
  • Experience building reusable tools, platforms, libraries, or modules used by multiple engineering teams.
  • Experience operating in large enterprise environments with cross-team collaboration and competing priorities.
  • Strong written communication skills and the ability to produce clear technical documentation and architecture proposals.
  • Demonstrated ability to influence technical direction and serve as a thought leader within SRE/platform engineering.
  • Experience with AI-assisted development tools such as Cursor, Claude Code, or GitHub Copilot is preferred.
  • Experience integrating AI/ML services into engineering or CI/CD workflows is a plus.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Optomi • Seattle (WA)

On-site
USD 140,000 - 180,000
Senior Site Reliability Engineer - Kubernetes & AI Platform
Senior Site Reliability Engineer - Kubernetes & AI Platform

Optomi • Seattle (WA)

On-site
USD 140,000 - 180,000
SRE Lead - Kubernetes, AI Platforms & Observability
SRE Lead - Kubernetes, AI Platforms & Observability

Optomi • New York (NY)

On-site
USD 140,000 - 200,000
Sr. Director, Site Reliability and Platform Engineering
Sr. Director, Site Reliability and Platform Engineering

Optomi • Tacoma (WA)

On-site
USD 150,000 - 200,000
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Senior SRE (Contract/Hybrid)
Senior SRE (Contract/Hybrid)

Optomi • Orlando (FL)

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Greenwood Village (CO)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Associate Engineer, Site Reliability
Associate Engineer, Site Reliability

R&D • United States

On-site
USD 90,000 - 140,000