Lead Site Reliability Engineer - Multi-Cloud & Kubernetes

SRI Tech Solutions Inc.

Orlando (FL)

On-site

USD 140,000 - 190,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SRI Tech Solutions Inc. is seeking a Lead Site Reliability Engineer to drive reliability, scalability, and operational excellence for a rapidly growing Generative AI platform.

You will provide technical leadership while designing and supporting highly available cloud infrastructure powering modern AI and data-driven applications. The ideal candidate combines deep SRE expertise with cloud, Kubernetes, and Infrastructure as Code, and will mentor other engineers while improving automation and high

Qualifications

  • 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related infrastructure roles.
  • Expert-level experience administering and operating Kubernetes in production environments.
  • Strong experience with Helm for Kubernetes application management.
  • Advanced experience using Terraform for Infrastructure as Code.
  • Hands-on experience building automated deployment pipelines using Harness or comparable enterprise CI/CD platforms.
  • Experience supporting production workloads across Google Cloud Platform, AWS, and Azure.
  • Strong scripting and automation skills using Python, Bash, and YAML.
  • Experience supporting production databases and messaging technologies, including PostgreSQL, Redis, Kafka, MongoDB, and Vault.
  • Experience with enterprise CI/CD platforms such as GitHub Actions, GitLab CI, Jenkins, Azure DevOps, or Harness.
  • Experience implementing observability solutions using technologies such as OpenTelemetry, Prometheus, Splunk, AppDynamics, or similar platforms.
  • Strong troubleshooting skills within distributed systems and cloud-native environments.
  • Experience working within Agile development environments.
  • Excellent communication skills with the ability to explain complex technical concepts to both technical and non-technical audiences.

Responsibilities

  • Lead the design, implementation, and support of highly available cloud infrastructure across Google Cloud Platform (primary), AWS, and Azure.
  • Design, build, and maintain Kubernetes infrastructure using Helm and Terraform for Infrastructure as Code.
  • Develop scalable platform solutions capable of maintaining 99.99% service availability.
  • Lead and mentor Site Reliability Engineers and DevOps engineers by providing technical guidance and establishing engineering best practices.
  • Plan, prioritize, and coordinate infrastructure initiatives within Agile delivery teams.
  • Design and implement automated deployment pipelines using modern CI/CD tools, including Harness.
  • Implement progressive deployment strategies such as blue/green deployments, canary releases, and feature flag rollouts.
  • Build and enhance observability solutions using monitoring, logging, alerting, and distributed tracing technologies.
  • Partner with engineering teams to review infrastructure sizing, capacity planning, and scalability requirements.
  • Support production systems through backups, upgrades, patching, disaster recovery, and operational maintenance.
  • Troubleshoot complex production issues across distributed systems and cloud-native applications.
  • Evaluate emerging SRE and DevOps technologies and recommend improvements to platform reliability and operational efficiency.
  • Ensure infrastructure aligns with security, governance, and compliance standards.

Skills

Leadership
SRE
DevOps
Cloud architecture
Automation

Tools

Kubernetes
Helm
Terraform
Harness
OpenTelemetry
Prometheus
Splunk
AppDynamics

Job description

SRI Tech Solutions Inc. is seeking a Lead Site Reliability Engineer to drive reliability, scalability, and operational excellence for a rapidly growing Generative AI platform.

You will provide technical leadership while designing and supporting highly available cloud infrastructure powering modern AI and data-driven applications. The ideal candidate combines deep SRE expertise with cloud, Kubernetes, and Infrastructure as Code, and will mentor other engineers while improving automation and high

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead SRE: Cloud Infra, Kubernetes & CI/CD
Lead SRE: Cloud Infra, Kubernetes & CI/CD

HTC Global Services • Orlando (FL)

On-site
USD 150,000 - 210,000
Health Insurance
401(k) matching
Paid Time Off
Lead Site Reliability Engineer - Infrastructure & DevOps
Lead Site Reliability Engineer - Infrastructure & DevOps

SRI Tech Solutions Inc. • Orlando (FL)

On-site
USD 140,000 - 190,000
Lead SRE for AI Platform & Multi-Cloud Infra
Lead SRE for AI Platform & Multi-Cloud Infra

Motion Recruitment Partners LLC • Lake Buena Vista (FL)

On-site
USD 120,000 - 150,000
Senior SRE: Kubernetes, CI/CD & Cloud Reliability Leader
Senior SRE: Kubernetes, CI/CD & Cloud Reliability Leader

techchaintalent • San Francisco (CA)

On-site
USD 120,000 - 160,000
Lead SRE for Generative AI Platform (Multi-Cloud)
Lead SRE for Generative AI Platform (Multi-Cloud)

Optomi • Orlando (FL)

Hybrid
USD 140,000 - 200,000
Senior SRE Lead – Generative AI Cloud
Senior SRE Lead – Generative AI Cloud

KellyMitchell Group • United States

On-site
Medical, Dental, & Vision Insurance
ESOP
401K
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Optomi • Orlando (FL)

Hybrid
USD 140,000 - 200,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

KellyMitchell Group • United States

On-site
Medical, Dental, & Vision Insurance
ESOP
401K
Senior SRE – AI Infrastructure Reliability Leader
Senior SRE – AI Infrastructure Reliability Leader

Nscale • San Francisco (CA), Seattle (WA), Houston (TX)

On-site
USD 170,000 - 265,000
Equity
Ownership from start
Flexible schedule
Site Reliability Engineer
Site Reliability Engineer

Motion Recruitment Partners LLC • Lake Buena Vista (FL)

On-site
USD 120,000 - 150,000