Lead Site Reliability Engineer

Optomi

Orlando (FL)

Hybrid

USD 140,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Optomi in partnership with a leading entertainment company is seeking a hands-on Lead Site Reliability Engineer to own the enterprise Generative AI platform in a hybrid Orlando role. You will architect, build, and operate scalable cloud infrastructure across GCP, AWS, and Azure while guiding Kubernetes operations and CI/CD modernization.

The ideal candidate has deep SRE expertise, strong automation skills, and experience supporting large-scale distributed systems in production, with leadership

Qualifications

  • 7+ years in Site Reliability Engineering, DevOps, or Platform Engineering.
  • Kubernetes clusters and Helm Charts administration experience.
  • Terraform and IaC expertise with multi-cloud
  • Strong CI/CD experience (Harness, GitHub Actions, GitLab, Jenkins, or Azure DevOps).

Responsibilities

  • Lead design, implementation, and reliability of the enterprise Generative AI platform.
  • Build and optimize Kubernetes environments with IaC and Terraform.
  • Develop and enhance CI/CD pipelines and cloud-native services.
  • Improve platform scalability, observability, monitoring, and incident response (99.99% uptime).
  • Support production databases, messaging systems, and cloud infra across GCP, AWS, Azure.
  • Mentor engineers and partner with architecture, security, product, and engineering teams.

Skills

Kubernetes
Python
Bash
CI/CD
Multi-cloud
SRE leadership
Agile/Scrum
AI/Gen AI platforms
Troubleshooting

Tools

Terraform
Helm
GitHub Actions
Jenkins
Azure DevOps
Prometheus
OpenTelemetry
Splunk

Job description

Lead Site Reliability Engineer (Generative AI Platform)

Optomi, in partnership with a leading entertainment company, is seeking a Lead Site Reliability Engineer (Generative AI Platform) to join their team in Orlando, FL. This is a long-term hybrid opportunity (4 days onsite) supporting the organization’s enterprise Generative AI platform by driving cloud infrastructure strategy, automation, platform reliability, and scalability across multi-cloud environments. The Lead Site Reliability Engineer will play a key technical leadership role supporting the enterprise Generative AI platform (JedAI). This engineer will help architect, build, and operate highly available cloud infrastructure across GCP, AWS, and Azure while leading Kubernetes operations, Infrastructure-as-Code initiatives, observability, and CI/CD modernization. The ideal candidate is a hands‑on technical leader with deep SRE expertise, strong automation skills, and experience supporting large‑scale distributed systems in production.

What the right candidate would enjoy:
  • Building infrastructure that powers enterprise Generative AI applications!
  • Working with modern cloud-native technologies across GCP, AWS, and Azure!
  • Leading platform reliability and automation initiatives in a highly collaborative environment!
  • Influencing long‑term cloud architecture and engineering best practices!
  • Mentoring engineers while remaining hands‑on with cutting‑edge technologies!
  • Solving complex distributed systems challenges at enterprise scale!
Experience required:
  • 7+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering
  • Extensive production experience administering Kubernetes clusters and Helm Charts
  • Expert-level Terraform and Infrastructure-as-Code experience
  • Strong CI/CD experience with Harness, GitHub Actions, GitLab, Jenkins, or Azure DevOps
  • Multi-cloud experience with GCP (preferred), AWS, and Azure
  • Strong Python and Bash scripting skills
  • Experience supporting PostgreSQL, Redis, Kafka, MongoDB, and Vault in production
  • Hands‑on experience with observability tools including Splunk, Prometheus, OpenTelemetry, or AppDynamics
  • Strong troubleshooting skills supporting highly available distributed systems
  • Technical leadership, mentoring, and Agile/Scrum experience
  • Experience supporting AI, Machine Learning, or Generative AI platforms
Responsibilities:
  • Lead the design, implementation, and reliability of the enterprise Generative AI platform
  • Build, maintain, and optimize Kubernetes environments using Infrastructure-as-Code and Terraform
  • Develop and enhance CI/CD pipelines, deployment automation, and cloud-native platform services
  • Improve platform scalability, observability, monitoring, and incident response while maintaining 99.99% uptime
  • Support production databases, messaging systems, and cloud infrastructure across GCP, AWS, and Azure
  • Mentor engineers, establish SRE best practices, and partner with architecture, security, product, and engineering teams to deliver highly available AI solutions
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead SRE for Generative AI Platform (Multi-Cloud)
Lead SRE for Generative AI Platform (Multi-Cloud)

Optomi • Orlando (FL)

Hybrid
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Motion Recruitment Partners LLC • Lake Buena Vista (FL)

On-site
USD 120,000 - 150,000
Lead SRE, Generative AI Platform (Remote)
Lead SRE, Generative AI Platform (Remote)

Optomi • United States

On-site
USD 150,000 - 190,000
Remote flexibility
Cutting-edge AI platform experience
Mentoring engineers and shaping cloud架
Lead Site Reliability Engineer - Infrastructure & DevOps
Lead Site Reliability Engineer - Infrastructure & DevOps

SRI Tech Solutions Inc. • Orlando (FL)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Optomi • New York (NY)

On-site
USD 140,000 - 200,000
Site Reliability Engineer
Site Reliability Engineer

Optomi • Seattle (WA)

On-site
USD 140,000 - 180,000
Lead Site Reliability Engineer - Multi-Cloud & Kubernetes
Lead Site Reliability Engineer - Multi-Cloud & Kubernetes

SRI Tech Solutions Inc. • Orlando (FL)

On-site
USD 140,000 - 190,000
Generative AI Engineer
Generative AI Engineer

Arkhya Tech. Inc. • Bellevue (WA)

Hybrid
USD 138,000 - 248,000
Senior SRE (Contract/Hybrid)
Senior SRE (Contract/Hybrid)

Optomi • Orlando (FL)

Hybrid
USD 120,000 - 180,000
Sr. Director, Site Reliability and Platform Engineering
Sr. Director, Site Reliability and Platform Engineering

Optomi • Tacoma (WA)

On-site
USD 150,000 - 200,000