Senior Site Reliability Engineer

Electronic Transaction Consultants Corporation

Northern (KY)

Hybrid

USD 120,000 - 180,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Paid time off
Health and dental plans
Retirement plans
EFAP
Employee referral program

Job summary

Quarterhill is seeking a Senior Site Reliability Engineer to join our growing team. This role focuses on reliability and performance of a cloud-native tolling platform processing transactions and payments around the clock.

You will lead incident response, implement SRE practices, and collaborate with software, infra, and operations teams to ensure highly available systems and scalable infrastructure.

Qualifications

  • 5+ years of experience in Site Reliability Engineering or DevOps.
  • Hands-on cloud experience across AWS/Azure/GCP and orchestration (Kubernetes).
  • Strong knowledge of monitoring, logging, and tracing stacks.
  • Scripting skills in Python, Bash, or Go.
  • Linux and Windows administration basics in production environments.
  • Experience with relational databases and distributed systems.

Responsibilities

  • System Reliability: Monitor and maintain health, availability, and performance.
  • Incident Management: Lead response, RCA, and blameless post-incident reviews.
  • Automation: Develop scripts to streamline operations and reduce toil.
  • Monitoring & Performance: Set up Prometheus, Grafana, OpenTelemetry, and alerts.
  • Capacity Planning: Plan scaling for growing transaction volumes and distributed databases.
  • Collaboration: Work with software, infrastructure, and operations teams.
  • System Optimization: Identify bottlenecks across infra and applications.
  • Continuous Improvement: Improve processes, docs, and SRE practices.
  • Disaster Recovery: Design and test DR plans for critical services.

Skills

Cloud platforms AWS/Azure/GCP
Kubernetes
Monitoring & Observability
Scripting Python/Bash/Go
Linux/Windows administration
Relational databases

Tools

Docker
ELK stack
Grafana
Prometheus
OpenTelemetry
Datadog
Terraform
Ansible
Helm
Argo CD (GitOps)

Job description

Quarterhill is seeking a Senior Site Reliability Engineer (SRE) to join our growing team. This role is an exciting opportunity to contribute to the reliability and performance of smart transportation systems, including a next-generation, cloud-native tolling platform that processes roadway transactions and payments around the clock. As a Senior SRE, you will ensure our systems are highly available, resilient, and scalable, take a leading role in incident response and reliability engineering practices, and help optimize operations across infrastructure and applications.

Responsibilities
  • System Reliability: Monitor and maintain the health, availability, and performance of critical transportation services.
  • Reliability Standards: Define and track service-level objectives (SLOs), error budgets, and reliability metrics for revenue-critical services.
  • Incident Management: Lead incident response — perform root cause analysis, coordinate resolution across teams, and drive blameless post-incident reviews and follow-up actions.
  • Automation: Develop and implement automation scripts to streamline operational tasks and improve efficiency.
  • Monitoring & Performance: Set up and maintain monitoring, logging, tracing, and alerting tools (e.g., Prometheus, Grafana, OpenTelemetry) to track service health, performance, and resource utilization.
  • Capacity Planning: Help assess and plan for capacity, scaling infrastructure to meet growing transaction volumes — including stateful systems such as distributed SQL databases and event-streaming clusters (e.g., NATS, Kafka).
  • Collaboration: Work with software engineering, infrastructure, and operations teams to improve the reliability of systems and services.
  • System Optimization: Identify performance bottlenecks, troubleshoot issues, and work on optimizations at both infrastructure and application layers.
  • Continuous Improvement: Contribute to the ongoing improvement of operational processes, documentation, and best practices in the SRE team.
  • Disaster Recovery: Participate in designing and testing disaster recovery plans to ensure the continuity of critical services.

This list of responsibilities might not cover everything you'll end up doing.

Qualifications
  • Experience: 5+ years of experience in Site Reliability Engineering, DevOps, or a similar role, preferably in a mission-critical or large-scale environment, including experience leading incident response and mentoring other engineers.
  • Technical Skills:
    • Experience with cloud platforms (AWS, Azure, GCP) and container orchestration tools (Docker, Kubernetes).
    • Proficiency with monitoring and logging tools (Prometheus, Grafana, ELK stack, Datadog, etc.).
    • Strong scripting skills in Python, Bash, or Go.
    • Solid understanding of Linux and Windows administration.
  • Database Knowledge: Familiarity with relational databases (MySQL, PostgreSQL, etc.) and distributed systems.
  • Collaboration & Communication: Excellent teamwork and communication skills, with the ability to work across teams to improve service reliability.
  • Problem-Solving: Strong troubleshooting skills with a proactive, solution-oriented mindset.
  • Experience in Intelligent Transportation: While not required, familiarity with transportation systems, autonomous vehicles, or real-time data systems is a plus.
Preferred Qualifications
  • Experience with traffic management systems, sensor data processing, or other intelligent transportation systems.
  • Knowledge of infrastructure-as-code tools (e.g., Terraform, Ansible, Helm) and GitOps workflows (e.g., Argo CD).
  • Exposure to CI/CD pipelines, including pipeline-as-code (e.g., Dagger, GitHub Actions), and Git-based version control.
Benefits
  • Paid days off (i.e. vacation, sick days, bereavement leave)
  • Health and Dental plans
  • Retirement plans
  • Employee and Family Assistance Program (EFAP)
  • Employee referral program

We welcome applicants from all backgrounds, regardless of race, color, religion, sex, veteran status, sexual orientation, gender identity, national origin, age, or disability or any other protected characteristics in accordance with applicable federal, state/provincial, and local laws. We’re committed to creating a workplace where everyone feels valued and respected.

We appreciate all responses and will acknowledge only those being considered for an interview.

We respectfully request no calls or unsolicited resumes from Agencies.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Quarterhill Inc. • United States

On-site
USD 140,000 - 190,000
Paid time off
Health and Dental plans
Retirement plans
+2
DevOps Engineer
DevOps Engineer

Electronic Transaction Consultants Corporation • Frisco (TX)

On-site
USD 120,000 - 180,000
Paid days off
Health and Dental plans
Retirement plans
+2
Senior SRE - Scalable Cloud for Smart Transit
Senior SRE - Scalable Cloud for Smart Transit

Electronic Transaction Consultants Corporation • Northern (KY)

Hybrid
USD 120,000 - 180,000
Paid time off
Health and dental plans
Retirement plans
+2
Senior/Staff Software Engineer
Senior/Staff Software Engineer

Electronic Transaction Consultants Corporation • Northern (KY)

On-site
USD 150,000 - 190,000
Health and Dental plans
Retirement plans
EFAP
+2
Senior/Staff Software Engineer
Senior/Staff Software Engineer

Electronic Transaction Consultants Corporation • United States

On-site
USD 150,000 - 190,000
Paid days off
Health and Dental plans
Retirement plans
+3
Manager, Software Development
Manager, Software Development

Electronic Transaction Consultants Corporation • Northern (KY)

On-site
USD 140,000 - 190,000
Paid time off
Health and dental plans
Retirement plans
+2
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

On-site
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3
Site Reliability Engineer
Site Reliability Engineer

Stelvio Inc. • Town of Texas (WI)

On-site
USD 125,000 - 145,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

On-site
USD 146,032 - 162,257
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

SEI • Chicago (IL)

On-site
USD 140,000 - 170,000
Comprehensive healthcare benefits
401(k) match
Paid Time Off (PTO)
+2