Site Reliability Engineer (SRE) – II

Huntington National Bank

Columbus (OH)

Hybrid

USD 90,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Huntington National Bank is looking for a Site Reliability Engineer Level II to maintain critical infrastructure's availability, scalability, and performance. You will respond to complex incidents and lead troubleshooting efforts while enhancing automation processes with tools like Terraform and Ansible.

The ideal candidate has a minimum of 5 years in site reliability engineering, with strong skills in cloud platforms and monitoring tools. This role offers a workplace type that combines in-office and flexible work options.

Qualifications

  • Minimum 5 years of experience in site reliability engineering, DevOps, or systems administration.
  • Strong experience with Linux/Unix administration and proficiency in scripting.
  • Deep understanding of cloud platforms and related services.

Responsibilities

  • Respond to complex incidents and ensure service uptime.
  • Lead troubleshooting for high-impact production issues.
  • Build and maintain automation scripts and infrastructure.

Skills

Linux/Unix administration
Scripting languages (Python, Bash, Go)
Cloud platforms (AWS, GCP, Azure)
Containerization (Docker, Kubernetes)
Monitoring tools (Dynatrace, Prometheus, Grafana)
Networking fundamentals (DNS, HTTP, TCP/IP)
CI/CD tools (Jenkins, GitLab CI, CircleCI)
Infrastructure automation (Terraform, Ansible, Puppet)
Distributed systems and microservices
Problem-solving and troubleshooting

Job description

Overview

As a Site Reliability Engineer (SRE) Level II, you will play a key role in maintaining the availability, scalability, and performance of critical infrastructure and services. You will build and automate solutions that enhance system reliability and support continuous delivery. In this role you will manage complex operational tasks and incidents, mentor junior SREs, and collaborate with development teams to ensure systems are designed for reliability from the ground up.

Responsibilities
  • Respond to complex incidents and ensure service uptime.
  • Lead troubleshooting for high‑impact production issues, performing root‑cause analysis and preventive measures.
  • Participate in on‑call rotations, acting as an escalation point for Level 1 SREs during major incidents.
  • Build and maintain automation scripts and infrastructure using Terraform, Ansible, or CloudFormation.
  • Implement automation to eliminate manual tasks and improve system reliability, scalability, and performance.
  • Analyze system performance and recommend optimizations for scalability and reliability.
  • Support capacity planning by monitoring metrics, traffic patterns, and usage trends to predict future resource needs.
  • Collaborate with software engineering teams to influence design of new services, ensuring they are scalable, reliable, and resilient.
  • Contribute to architectural decisions, aligning with best practices in fault tolerance, redundancy, and recovery.
  • Build and maintain robust monitoring, alerting, and observability solutions; optimize existing tools and build dashboards for better visibility.
  • Ensure systems and infrastructure are secure and compliant; assist with vulnerability management, patching, and security implementation.
  • Lead continuous improvement of operational processes, tools, and workflows.
  • Implement and enforce best practices in deployment, monitoring, and incident management to reduce downtime.
Qualifications
  • Minimum 5 years of experience in site reliability engineering, DevOps, systems administration, or related roles.
  • Strong experience with Linux/Unix administration and proficiency in scripting languages such as Python, Bash, or Go.
  • Deep understanding of cloud platforms (AWS, GCP, Azure) and related services (EC2, S3, Lambda, Kubernetes, etc.).
  • Experience with containerization and orchestration technologies like Docker and Kubernetes.
  • Proficiency with monitoring and observability tools such as Dynatrace, Prometheus, Grafana, Datadog, or ELK Stack.
  • Strong understanding of networking fundamentals (DNS, HTTP, TCP/IP), load balancing, and CDNs.
  • Experience with CI/CD tools (Jenkins, GitLab CI, CircleCI) and infrastructure automation (Terraform, Ansible, Puppet).
  • Familiarity with distributed systems and microservices architecture.
  • Excellent problem‑solving and troubleshooting skills, especially in diagnosing production issues in high‑scale environments.
Preferred
  • Background in MLOps, data engineering, and/or cloud‑native AI deployment.
  • Strong communication and documentation abilities.
  • Knowledge of security best practices for AI and cloud infrastructure.
  • Contributions to open‑source AI/SRE projects or relevant technical communities.
  • Proven track record of managing complex infrastructure, troubleshooting production issues, and optimizing system performance.
Exempt Status

Yes (not eligible for overtime pay)

Workplace Type

Office. Certain positions outside our branch network may be eligible for a flexible work arrangement, combining in‑office and work‑from‑home. Remote roles may also have the opportunity to come together in our offices for moments that matter. Specific work arrangements will be provided by the hiring team.

Huntington is an Equal Opportunity Employer.

Tobacco‑Free Hiring Practice: Visit Huntington’s Career Web Site for more details.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE) – II
Site Reliability Engineer (SRE) – II

Huntington Bank • Easton (PA)

Hybrid
USD 90,000 - 120,000
Site Reliability Engineer
Site Reliability Engineer

SRE • Puerto Rico

Hybrid
USD 120,000 - 180,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

OutSolve • Mission (KS)

Remote
USD 90,000 - 130,000
100% remote work environment
Competitive compensation
Professional development opportunities
+1
Site Reliability Engineer (Secret Clearance)
Site Reliability Engineer (Secret Clearance)

ROI Services LLC • Huntsville (AL)

On-site
USD 110,000 - 150,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • New Hampshire

On-site
USD 110,000 - 160,000
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

VITG • Ellicott City (MD)

Hybrid
USD 90,000 - 120,000
401(k) with employer contribution
Medical/Dental/Vision insurance
Paid vacation (PTO)
Senior Site Reliability Engineer II
Senior Site Reliability Engineer II

RELX • New York (NY)

Hybrid
USD 104,000 - 175,000
Competitive salary
Hybrid or remote options
Career growth in SRE and DevOps
+1