Lead SRE: Cloud Infra, Kubernetes & CI/CD

HTC Global Services

Orlando (FL)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health Insurance
401(k) matching
Paid Time Off

Job summary

HTC Global Services is seeking a Lead Site Reliability Engineer to drive reliability, scalability, and operational excellence for a Generative AI platform. You will provide technical leadership, design highly available cloud infrastructure, and mentor a team of SREs and DevOps engineers.

The role emphasizes Kubernetes, IaC, and end-to-end observability in a fast-moving environment. You will lead infrastructure across Google Cloud Platform (primary), AWS, and Azure, building scalable solutions

Qualifications

  • 7+ years in SRE/DevOps or related infrastructure roles.
  • Expert Kubernetes in production, including Helm usage.
  • Terraform IaC and automated deployment pipelines experience.
  • Cloud experience across GCP, AWS and Azure.

Responsibilities

  • Lead design and support of highly available cloud infrastructure across GCP, AWS, and Azure.
  • Build and maintain Kubernetes infrastructure with Helm and Terraform.
  • Develop scalable platform solutions targeting 99.99% availability.
  • Mentor SRE/DevOps engineers and promote best practices.
  • Coordinate infrastructure initiatives within Agile teams.
  • Implement automated deployments using modern CI/CD tools (Harness, etc.).
  • Apply progressive deployment strategies (blue/green, canary, feature flags).
  • Enhance observability with monitoring, logging, tracing, and alerting.
  • Review sizing, capacity, and scalability with engineering teams.
  • Support production systems with backups, upgrades, DR, and maintenance.
  • Troubleshoot complex distributed systems and cloud-native apps.
  • Evaluate new SRE/DevOps tech for reliability improvements.
  • Ensure security, governance, and compliance alignment.

Skills

Kubernetes production experience
Helm
Terraform
CI/CD
Python scripting
Cloud platforms (GCP/AWS/Azure)

Tools

Harness
GitHub Actions
GitLab CI
Jenkins
OpenTelemetry

Job description

HTC Global Services is seeking a Lead Site Reliability Engineer to drive reliability, scalability, and operational excellence for a Generative AI platform. You will provide technical leadership, design highly available cloud infrastructure, and mentor a team of SREs and DevOps engineers.

The role emphasizes Kubernetes, IaC, and end-to-end observability in a fast-moving environment. You will lead infrastructure across Google Cloud Platform (primary), AWS, and Azure, building scalable solutions

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Platform Reliability Engineer
Senior AI Platform Reliability Engineer

HTC Global Services • Orlando (FL)

On-site
USD 150,000 - 210,000
Health Insurance
401(k) matching
Paid Time Off
Lead GCP SRE & Platform Reliability
Lead GCP SRE & Platform Reliability

Altice USA • Bethpage (NY)

On-site
USD 134,000 - 220,000
Hybrid Lead SRE: Scale, Automate & Stabilize Cloud Services
Hybrid Lead SRE: Scale, Automate & Stabilize Cloud Services

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Site Reliability Engineer II — Remote, GCP & Observability
Site Reliability Engineer II — Remote, GCP & Observability

HTC Global Services, Inc. • Dearborn (MI)

Hybrid
USD 110,000 - 140,000
Hybrid work
Work-Life Balance
Career development plan
+3
Lead Site Reliability Engineer, GCP & Cloud Platforms
Lead Site Reliability Engineer, GCP & Cloud Platforms

Optimum • Bethpage (NY)

On-site
USD 134,000 - 220,000
Senior SRE: AI Cloud Platform & Kubernetes
Senior SRE: AI Cloud Platform & Kubernetes

Lambda Inc. • San Francisco (CA)

Hybrid
USD 190,000 - 270,000
Health insurance
401k with company match
Flexible PTO
+2
Lead SRE - Remote/Hybrid, Enterprise-Scale Reliability
Lead SRE - Remote/Hybrid, Enterprise-Scale Reliability

Empower Retirement • Greenwood Village (CO)

Hybrid
USD 114,000 - 166,000
Medical insurance
401(k) with company match
Tuition reimbursement
+3
Lead SRE, Generative AI Platform (Remote)
Lead SRE, Generative AI Platform (Remote)

Optomi • United States

On-site
USD 150,000 - 190,000
Remote flexibility
Cutting-edge AI platform experience
Mentoring engineers and shaping cloud架
Senior SRE – AI Cloud Platform, Kubernetes Expert
Senior SRE – AI Cloud Platform, Kubernetes Expert

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health, dental, vision coverage for in
Wellness and commuter stipends
401k with 2% company match
+1
Remote SRE Technical Lead — Automate & Scale Reliability
Remote SRE Technical Lead — Automate & Scale Reliability

Bright Vision Technologies • New Albany (IN), City of Albany (NY)

On-site
USD 100,000 - 150,000