Senior AI Platform Reliability Engineer

HTC Global Services

Orlando (FL)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health Insurance
401(k) matching
Paid Time Off

Job summary

HTC Global Services is seeking a Lead Site Reliability Engineer to drive reliability, scalability, and operational excellence for a Generative AI platform. You will provide technical leadership, design highly available cloud infrastructure, and mentor a team of SREs and DevOps engineers.

The role emphasizes Kubernetes, IaC, and end-to-end observability in a fast-moving environment. You will lead infrastructure across Google Cloud Platform (primary), AWS, and Azure, building scalable solutions

Qualifications

  • 7+ years in SRE/DevOps or related infrastructure roles.
  • Expert Kubernetes in production, including Helm usage.
  • Terraform IaC and automated deployment pipelines experience.
  • Cloud experience across GCP, AWS and Azure.

Responsibilities

  • Lead design and support of highly available cloud infrastructure across GCP, AWS, and Azure.
  • Build and maintain Kubernetes infrastructure with Helm and Terraform.
  • Develop scalable platform solutions targeting 99.99% availability.
  • Mentor SRE/DevOps engineers and promote best practices.
  • Coordinate infrastructure initiatives within Agile teams.
  • Implement automated deployments using modern CI/CD tools (Harness, etc.).
  • Apply progressive deployment strategies (blue/green, canary, feature flags).
  • Enhance observability with monitoring, logging, tracing, and alerting.
  • Review sizing, capacity, and scalability with engineering teams.
  • Support production systems with backups, upgrades, DR, and maintenance.
  • Troubleshoot complex distributed systems and cloud-native apps.
  • Evaluate new SRE/DevOps tech for reliability improvements.
  • Ensure security, governance, and compliance alignment.

Skills

Kubernetes production experience
Helm
Terraform
CI/CD
Python scripting
Cloud platforms (GCP/AWS/Azure)

Tools

Harness
GitHub Actions
GitLab CI
Jenkins
OpenTelemetry

Job description

Job Title

Lead Site Reliability Engineer (SRE)

Overview / Summary

We are seeking a Lead Site Reliability Engineer to help drive the reliability, scalability, and operational excellence of a rapidly growing Generative AI platform. This role provides technical leadership while designing and supporting highly available cloud infrastructure powering modern AI and data‑driven applications.

Key Responsibilities
  • Lead the design, implementation, and support of highly available cloud infrastructure across Google Cloud Platform (primary), AWS, and Azure.
  • Design, build, and maintain Kubernetes infrastructure using Helm and Terraform for Infrastructure as Code.
  • Develop scalable platform solutions capable of maintaining 99.99% service availability.
  • Lead and mentor Site Reliability Engineers and DevOps engineers by providing technical guidance and establishing engineering best practices.
  • Plan, prioritize, and coordinate infrastructure initiatives within Agile delivery teams.
  • Design and implement automated deployment pipelines using modern CI/CD tools, including Harness.
  • Implement progressive deployment strategies such as blue/green deployments, canary releases, and feature flag rollouts.
  • Build and enhance observability solutions using monitoring, logging, alerting, and distributed tracing technologies.
  • Partner with engineering teams to review infrastructure sizing, capacity planning, and scalability requirements.
  • Support production systems through backups, upgrades, patching, disaster recovery, and operational maintenance.
  • Troubleshoot complex production issues across distributed systems and cloud-native applications.
  • Evaluate emerging SRE and DevOps technologies and recommend improvements to platform reliability and operational efficiency.
  • Ensure infrastructure aligns with security, governance, and compliance standards.
Required Qualifications
  • 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related infrastructure roles.
  • Expert‑level experience administering and operating Kubernetes in production environments.
  • Strong experience with Helm for Kubernetes application management.
  • Advanced experience using Terraform for Infrastructure as Code.
  • Hands‑on experience building automated deployment pipelines using Harness or comparable enterprise CI/CD platforms.
  • Experience supporting production workloads across Google Cloud Platform, AWS, and Azure.
  • Strong scripting and automation skills using Python, Bash, and YAML.
  • Experience supporting production databases and messaging technologies, including PostgreSQL, Redis, Kafka, MongoDB, and Vault.
  • Experience with enterprise CI/CD platforms such as GitHub Actions, GitLab CI, Jenkins, Azure DevOps, or Harness.
  • Experience implementing observability solutions using technologies such as OpenTelemetry, Prometheus, Splunk, AppDynamics, or similar platforms.
  • Strong troubleshooting skills within distributed systems and cloud‑native environments.
  • Experience working within Agile development environments.
  • Excellent communication skills with the ability to explain complex technical concepts to both technical and non‑technical audiences.
What Makes HTC a Great Place to Build Your Future

HTC Global Services wants you to join our team. Come build new things with us and advance your career. At HTC Global, you’ll collaborate with experts, work alongside clients, and be part of high‑performing teams driving success together. You’ll have long‑term opportunities to grow your career and develop skills in the latest emerging technologies.

At HTC Global Services, our employees have access to a comprehensive benefits package. Benefits can include Group Health (Medical, Dental, and Vision), Paid Time Off, Paid Holidays, 401(k) matching, Group Life and Disability insurance, Professional Development opportunities, Wellness programs, and a variety of other perks.

Our success as a company is built on inclusion and diversity. HTC Global Services is committed to providing a workplace free from discrimination and harassment, where every employee is treated with dignity and respect. We celebrate differences and believe that diverse cultures, perspectives, and skills drive innovation and success. HTC is an Equal Opportunity Employer and a proud National Minority Supplier. We seek to empower each individual, fostering an environment where everyone feels valued, included, and respected.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead SRE: Cloud Infra, Kubernetes & CI/CD
Lead SRE: Cloud Infra, Kubernetes & CI/CD

HTC Global Services • Orlando (FL)

On-site
USD 150,000 - 210,000
Health Insurance
401(k) matching
Paid Time Off
Lead Site Reliability Engineer - Infrastructure & DevOps
Lead Site Reliability Engineer - Infrastructure & DevOps

SRI Tech Solutions Inc. • Orlando (FL)

On-site
USD 140,000 - 190,000
Senior Software Engineer – DevOps & CI/CD Automation
Senior Software Engineer – DevOps & CI/CD Automation

HTC Global Services • Dearborn (MI)

On-site
USD 130,000 - 180,000
Group Health (Medical, Dental, Vision)
Paid Time Off
401(k) matching
+2
Senior AI Software Engineer – Multi-Agent Systems (Python, GCP)
Senior AI Software Engineer – Multi-Agent Systems (Python, GCP)

HTC Global Services • Dearborn (MI)

On-site
USD 140,000 - 180,000
Group Health (Medical, Dental, Vision)
401(k) matching
Paid Time Off
+2
Senior Site Reliability Engineer – Multi-Cloud Architecture
Senior Site Reliability Engineer – Multi-Cloud Architecture

HTC Global Services • Madison (WI)

On-site
USD 100,000 - 130,000
Paid-Time-Off
401K matching
Life Insurance
+1
Java Software Engineer – Google Cloud Platform (GCP)
Java Software Engineer – Google Cloud Platform (GCP)

HTC Global Services • Dearborn (MI)

Hybrid
USD 90,000 - 130,000
Group Health (Medical, Dental, Vision)
Paid Time Off
401(k) matching
+3
Senior Full Stack Software Engineer (Java, Spring Boot, Angular, GCP)
Senior Full Stack Software Engineer (Java, Spring Boot, Angular, GCP)

HTC Global Services • Dearborn (MI)

On-site
USD 110,000 - 160,000
Full Stack Software Engineer (Java, Spring Boot, Angular, GCP)
Full Stack Software Engineer (Java, Spring Boot, Angular, GCP)

HTC Global Services • Dearborn (MI)

On-site
USD 110,000 - 140,000
Health benefits
Paid time off
401(k) matching
Senior Data Engineer
Senior Data Engineer

HTC Global Services • Dearborn (MI)

On-site
USD 120,000 - 170,000
Health insurance
Paid time off
401(k) matching
+4
Senior AI Software Engineer – Multi-Agent Systems (Python, GCP)
Senior AI Software Engineer – Multi-Agent Systems (Python, GCP)

HTC Global Services, Inc. • Dearborn (MI)

Hybrid
USD 130,000 - 185,000
Hybrid and Workplace flexibility
Work-Life-Balance
Well-defined career development plan
+3