Senior AI Platform Reliability Engineer

HTC Global Services

Orlando (FL)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health Insurance
401(k) matching
Paid Time Off

Job summary

HTC Global Services is seeking a Lead Site Reliability Engineer to drive reliability, scalability, and operational excellence for a Generative AI platform. You will provide technical leadership, design highly available cloud infrastructure, and mentor a team of SREs and DevOps engineers.

The role emphasizes Kubernetes, IaC, and end-to-end observability in a fast-moving environment. You will lead infrastructure across Google Cloud Platform (primary), AWS, and Azure, building scalable solutions

Qualifications

  • 7+ years in SRE/DevOps or related infrastructure roles.
  • Expert Kubernetes in production, including Helm usage.
  • Terraform IaC and automated deployment pipelines experience.
  • Cloud experience across GCP, AWS and Azure.

Responsibilities

  • Lead design and support of highly available cloud infrastructure across GCP, AWS, and Azure.
  • Build and maintain Kubernetes infrastructure with Helm and Terraform.
  • Develop scalable platform solutions targeting 99.99% availability.
  • Mentor SRE/DevOps engineers and promote best practices.
  • Coordinate infrastructure initiatives within Agile teams.
  • Implement automated deployments using modern CI/CD tools (Harness, etc.).
  • Apply progressive deployment strategies (blue/green, canary, feature flags).
  • Enhance observability with monitoring, logging, tracing, and alerting.
  • Review sizing, capacity, and scalability with engineering teams.
  • Support production systems with backups, upgrades, DR, and maintenance.
  • Troubleshoot complex distributed systems and cloud-native apps.
  • Evaluate new SRE/DevOps tech for reliability improvements.
  • Ensure security, governance, and compliance alignment.

Skills

Kubernetes production experience
Helm
Terraform
CI/CD
Python scripting
Cloud platforms (GCP/AWS/Azure)

Tools

Harness
GitHub Actions
GitLab CI
Jenkins
OpenTelemetry

Job description

Job Title

Lead Site Reliability Engineer (SRE)

Overview / Summary

We are seeking a Lead Site Reliability Engineer to help drive the reliability, scalability, and operational excellence of a rapidly growing Generative AI platform. This role provides technical leadership while designing and supporting highly available cloud infrastructure powering modern AI and data‑driven applications.

Key Responsibilities
  • Lead the design, implementation, and support of highly available cloud infrastructure across Google Cloud Platform (primary), AWS, and Azure.
  • Design, build, and maintain Kubernetes infrastructure using Helm and Terraform for Infrastructure as Code.
  • Develop scalable platform solutions capable of maintaining 99.99% service availability.
  • Lead and mentor Site Reliability Engineers and DevOps engineers by providing technical guidance and establishing engineering best practices.
  • Plan, prioritize, and coordinate infrastructure initiatives within Agile delivery teams.
  • Design and implement automated deployment pipelines using modern CI/CD tools, including Harness.
  • Implement progressive deployment strategies such as blue/green deployments, canary releases, and feature flag rollouts.
  • Build and enhance observability solutions using monitoring, logging, alerting, and distributed tracing technologies.
  • Partner with engineering teams to review infrastructure sizing, capacity planning, and scalability requirements.
  • Support production systems through backups, upgrades, patching, disaster recovery, and operational maintenance.
  • Troubleshoot complex production issues across distributed systems and cloud-native applications.
  • Evaluate emerging SRE and DevOps technologies and recommend improvements to platform reliability and operational efficiency.
  • Ensure infrastructure aligns with security, governance, and compliance standards.
Required Qualifications
  • 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related infrastructure roles.
  • Expert‑level experience administering and operating Kubernetes in production environments.
  • Strong experience with Helm for Kubernetes application management.
  • Advanced experience using Terraform for Infrastructure as Code.
  • Hands‑on experience building automated deployment pipelines using Harness or comparable enterprise CI/CD platforms.
  • Experience supporting production workloads across Google Cloud Platform, AWS, and Azure.
  • Strong scripting and automation skills using Python, Bash, and YAML.
  • Experience supporting production databases and messaging technologies, including PostgreSQL, Redis, Kafka, MongoDB, and Vault.
  • Experience with enterprise CI/CD platforms such as GitHub Actions, GitLab CI, Jenkins, Azure DevOps, or Harness.
  • Experience implementing observability solutions using technologies such as OpenTelemetry, Prometheus, Splunk, AppDynamics, or similar platforms.
  • Strong troubleshooting skills within distributed systems and cloud‑native environments.
  • Experience working within Agile development environments.
  • Excellent communication skills with the ability to explain complex technical concepts to both technical and non‑technical audiences.
What Makes HTC a Great Place to Build Your Future

HTC Global Services wants you to join our team. Come build new things with us and advance your career. At HTC Global, you’ll collaborate with experts, work alongside clients, and be part of high‑performing teams driving success together. You’ll have long‑term opportunities to grow your career and develop skills in the latest emerging technologies.

At HTC Global Services, our employees have access to a comprehensive benefits package. Benefits can include Group Health (Medical, Dental, and Vision), Paid Time Off, Paid Holidays, 401(k) matching, Group Life and Disability insurance, Professional Development opportunities, Wellness programs, and a variety of other perks.

Our success as a company is built on inclusion and diversity. HTC Global Services is committed to providing a workplace free from discrimination and harassment, where every employee is treated with dignity and respect. We celebrate differences and believe that diverse cultures, perspectives, and skills drive innovation and success. HTC is an Equal Opportunity Employer and a proud National Minority Supplier. We seek to empower each individual, fostering an environment where everyone feels valued, included, and respected.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Backend Software Engineer – Cloud & AI
Senior Backend Software Engineer – Cloud & AI

HTC Global Services • Dearborn (MI)

On-site
USD 120,000 - 180,000
Group Health Insurance
401(k) matching
Paid Time Off
Site Reliability Engineer II
Site Reliability Engineer II

HTC Global Services, Inc. • Dearborn (MI)

On-site
USD 110,000 - 140,000
Hybrid work
Work-Life Balance
Career development plan
+3
AI/ML Engineer
AI/ML Engineer

HTC Global Services • Dearborn (MI)

On-site
USD 120,000 - 160,000
Health insurance
401(k) matching
Paid time off
Lead SRE: Cloud Infra, Kubernetes & CI/CD
Lead SRE: Cloud Infra, Kubernetes & CI/CD

HTC Global Services • Orlando (FL)

On-site
USD 150,000 - 210,000
Health Insurance
401(k) matching
Paid Time Off
Software Engineering Manager II, Site Reliability Engineering, AI Foundry SRE
Software Engineering Manager II, Site Reliability Engineering, AI Foundry SRE

Google • San Jose (CA)

On-site
USD 207,000 - 300,000
Equity
Bonus target
Benefits
Senior Site Reliability Engineer – Multi-Cloud Architecture
Senior Site Reliability Engineer – Multi-Cloud Architecture

HTC Global Services • Madison (WI)

On-site
USD 100,000 - 130,000
Paid-Time-Off
401K matching
Life Insurance
+1
Senior Backend Software Engineer
Senior Backend Software Engineer

HTC Global Services • Dearborn (MI)

On-site
USD 120,000 - 180,000
Health, Dental, Vision insurance
Paid time off
401(k) matching
+3
Full Stack Data Engineer
Full Stack Data Engineer

HTC Global • Dearborn (MI), Northern (KY)

Hybrid
USD 120,000 - 170,000
Hybrid and Workplace flexibility
Work-Life-Balance
Well-defined career development plan
+2
AI/ML Engineer
AI/ML Engineer

HTC Global Services, Inc. • Dearborn (MI), Northern (KY)

On-site
USD 120,000 - 170,000
Hybrid and Workplace flexibility
Work-Life-Balance
Career development program
+1
Senior Platform Engineer – VMware & Kubernetes
Senior Platform Engineer – VMware & Kubernetes

HTC Global Services • Dearborn (MI)

On-site
USD 110,000 - 150,000