Site Reliability Engineer II

HTC Global Services

Dearborn (MI)

On-site

USD 110,000 - 140,000

Full time

25 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Group Health (Medical, Dental, Vision)
401(k) matching
Paid Time Off
Wellness programs
Professional Development opportunities

Job summary

HTC Global Services is seeking a Site Reliability Engineer II to join our SRE team focused on observability, monitoring, and technical consulting across GCP-based data platforms. You will ensure reliability and performance of cloud and network systems through automation, monitoring, and optimization.

The role emphasizes hands-on work with cloud infrastructure, BigQuery workloads, CI/CD pipelines, and enterprise monitoring tools, while collaborating with cross-functional teams to improve incident

Qualifications

  • Bachelor's degree.
  • 4+ years of experience in IT.
  • 3+ years of development experience.
  • Practitioner-level experience with at least one coding language or framework.
  • Hands-on experience with Google Cloud Platform (GCP).
  • Experience with BigQuery.
  • Experience with Dynatrace.
  • Proficiency with monitoring and observability tools, ideally Dynatrace or comparable tools such as Datadog or New Relic.
  • Familiarity with ITSM tools such as ServiceNow, including incident, problem, and change management.

Responsibilities

  • Collaborate with infrastructure teams to automate routine tasks.
  • Monitor and manage production environments, proactively identifying and resolving issues.
  • Participate in building advanced tooling for system access monitoring, log session recording, and reliability administration across multiple geographically distributed data centers.
  • Engage with engineering teams to improve on-call efficiencies, incident management, and post-mortem analysis.
  • Perform capacity planning and optimization to support growing demands and traffic patterns.
  • Maintain monitoring and alerting systems for proactive system health checks.
  • Continuously improve system performance, stability, and security through data-driven analysis and optimization.
  • Create and maintain comprehensive documentation and diagrams to facilitate knowledge sharing.
  • Work hands-on with cloud infrastructure, BigQuery workloads, CI/CD pipelines, and enterprise monitoring tools to maintain critical systems at scale.

Skills

GCP
BigQuery
Dynatrace
Monitoring
ITSM
Coding language

Education

Bachelor's degree

Tools

Datadog
New Relic
ServiceNow

Job description

Overview / Summary

We are seeking a Site Reliability Engineer to join an SRE team focused on observability, monitoring, and technical consulting across GCP-based data platforms. This role is responsible for ensuring the availability, reliability, and performance of cloud and network systems and services through automation, monitoring, troubleshooting, and continuous optimization.

Job Description
Site Reliability Engineer II
Overview / Summary

We are seeking a Site Reliability Engineer to join an SRE team focused on observability, monitoring, and technical consulting across GCP-based data platforms. This role is responsible for ensuring the availability, reliability, and performance of cloud and network systems and services through automation, monitoring, troubleshooting, and continuous optimization.

Key Responsibilities
  • Collaborate with infrastructure teams to implement critical solutions by automating routine tasks.
  • Monitor and manage production environments, proactively identifying and resolving issues.
  • Participate in building advanced tooling for system access monitoring, log session recording, and reliability administration across multiple geographically distributed data centers.
  • Engage with engineering teams to improve on-call efficiencies, incident management, and post-mortem analysis.
  • Perform capacity planning and optimization to support growing demands and traffic patterns.
  • Maintain monitoring and alerting systems for proactive system health checks.
  • Continuously improve system performance, stability, and security through data-driven analysis and optimization.
  • Create and maintain comprehensive documentation and diagrams to facilitate knowledge sharing.
  • Work hands-on with cloud infrastructure, BigQuery workloads, CI/CD pipelines, and enterprise monitoring tools to maintain critical systems at scale.
Required Qualifications
  • Bachelor's degree.
  • 4+ years of experience in IT.
  • 3+ years of development experience.
  • Practitioner-level experience with at least one coding language or framework.
  • Hands-on experience with Google Cloud Platform (GCP).
  • Experience with BigQuery.
  • Experience with Dynatrace.
  • Proficiency with monitoring and observability tools, ideally Dynatrace or comparable tools such as Datadog or New Relic.
  • Familiarity with ITSM tools such as ServiceNow, including incident, problem, and change management.
Preferred Qualifications
  • Experience with GCP Cloud Run.
  • Experience with Python.
  • Strong troubleshooting and problem-solving skills.
  • Familiarity with AI tools, including agents, skills, LLMs, and copilots.
  • Experience defining and tracking SLAs, SLOs, and SLIs.
What Makes HTC A Great Place To Build Your Future

HTC Global Services wants you to join our team. Come build new things with us and advance your career. At HTC Global, you’ll collaborate with experts, work alongside clients, and be part of high-performing teams driving success together. You’ll have long-term opportunities to grow your career and develop skills in the latest emerging technologies.

At HTC Global Services, our employees have access to a comprehensive benefits package. Benefits can include Group Health (Medical, Dental, and Vision), Paid Time Off, Paid Holidays, 401(k) matching, Group Life and Disability insurance, Professional Development opportunities, Wellness programs, and a variety of other perks.

Our success as a company is built on inclusion and diversity. HTC Global Services is committed to providing a workplace free from discrimination and harassment, where every employee is treated with dignity and respect. We celebrate differences and believe that diverse cultures, perspectives, and skills drive innovation and success. HTC is an Equal Opportunity Employer and a proud National Minority Supplier. We seek to empower each individual, fostering an environment where everyone feels valued, included, and respected.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer II
Site Reliability Engineer II

HTC Global Services, Inc. • Dearborn (MI)

Hybrid
USD 110,000 - 140,000
Hybrid work
Work-Life Balance
Career development plan
+3
Senior AI Platform Reliability Engineer
Senior AI Platform Reliability Engineer

HTC Global Services • Orlando (FL)

On-site
USD 150,000 - 210,000
Health Insurance
401(k) matching
Paid Time Off
Site Reliability Engineer II — Remote, GCP & Observability
Site Reliability Engineer II — Remote, GCP & Observability

HTC Global Services, Inc. • Dearborn (MI)

Hybrid
USD 110,000 - 140,000
Hybrid work
Work-Life Balance
Career development plan
+3
GCP SRE II: Observability, Automation & Reliability
GCP SRE II: Observability, Automation & Reliability

HTC Global Services • Dearborn (MI)

On-site
USD 110,000 - 140,000
Group Health (Medical, Dental, Vision)
401(k) matching
Paid Time Off
+2
Senior Software Engineer – Full Stack / Cloud
Senior Software Engineer – Full Stack / Cloud

HTC Global Services • Dearborn (MI)

On-site
USD 120,000 - 160,000
GCP Cloud Engineer – Terraform / Java / Python
GCP Cloud Engineer – Terraform / Java / Python

HTC Global Services • Dearborn (MI)

On-site
USD 110,000 - 150,000
Health insurance
Paid time off
401(k) matching
+2
Senior Backend Software Engineer
Senior Backend Software Engineer

HTC Global Services • Dearborn (MI)

On-site
USD 120,000 - 180,000
Health, Dental, Vision insurance
Paid time off
401(k) matching
+3
Senior Backend Software Engineer – Cloud & AI
Senior Backend Software Engineer – Cloud & AI

HTC Global Services • Dearborn (MI)

On-site
USD 120,000 - 180,000
Group Health Insurance
401(k) matching
Paid Time Off
Data Engineer III
Data Engineer III

HTC Global Services • Dearborn (MI)

On-site
USD 110,000 - 150,000
Health benefits
401(k) matching
Paid time off
Site Reliability Engineer
Site Reliability Engineer

Apex Systems • Dearborn (MI)

Hybrid
USD 110,000 - 165,000