Quarterhill is seeking a Senior Site Reliability Engineer (SRE) to join our growing team. This role is an exciting opportunity to contribute to the reliability and performance of smart transportation systems, including a next-generation, cloud-native tolling platform that processes roadway transactions and payments around the clock. As a Senior SRE, you will ensure our systems are highly available, resilient, and scalable, take a leading role in incident response and reliability engineering practices, and help optimize operations across infrastructure and applications.
Responsibilities
- System Reliability: Monitor and maintain the health, availability, and performance of critical transportation services.
- Reliability Standards: Define and track service-level objectives (SLOs), error budgets, and reliability metrics for revenue-critical services.
- Incident Management: Lead incident response — perform root cause analysis, coordinate resolution across teams, and drive blameless post-incident reviews and follow-up actions.
- Automation: Develop and implement automation scripts to streamline operational tasks and improve efficiency.
- Monitoring & Performance: Set up and maintain monitoring, logging, tracing, and alerting tools (e.g., Prometheus, Grafana, OpenTelemetry) to track service health, performance, and resource utilization.
- Capacity Planning: Help assess and plan for capacity, scaling infrastructure to meet growing transaction volumes — including stateful systems such as distributed SQL databases and event-streaming clusters (e.g., NATS, Kafka).
- Collaboration: Work with software engineering, infrastructure, and operations teams to improve the reliability of systems and services.
- System Optimization: Identify performance bottlenecks, troubleshoot issues, and work on optimizations at both infrastructure and application layers.
- Continuous Improvement: Contribute to the ongoing improvement of operational processes, documentation, and best practices in the SRE team.
- Disaster Recovery: Participate in designing and testing disaster recovery plans to ensure the continuity of critical services.
This list of responsibilities might not cover everything you'll end up doing.
Qualifications
- Experience: 5+ years of experience in Site Reliability Engineering, DevOps, or a similar role, preferably in a mission-critical or large-scale environment, including experience leading incident response and mentoring other engineers.
- Technical Skills:
- Experience with cloud platforms (AWS, Azure, GCP) and container orchestration tools (Docker, Kubernetes).
- Proficiency with monitoring and logging tools (Prometheus, Grafana, ELK stack, Datadog, etc.).
- Strong scripting skills in Python, Bash, or Go.
- Solid understanding of Linux and Windows administration.
- Database Knowledge: Familiarity with relational databases (MySQL, PostgreSQL, etc.) and distributed systems.
- Collaboration & Communication: Excellent teamwork and communication skills, with the ability to work across teams to improve service reliability.
- Problem-Solving: Strong troubleshooting skills with a proactive, solution-oriented mindset.
- Experience in Intelligent Transportation: While not required, familiarity with transportation systems, autonomous vehicles, or real-time data systems is a plus.
Preferred Qualifications:
- Experience with traffic management systems, sensor data processing, or other intelligent transportation systems.
- Knowledge of infrastructure-as-code tools (e.g., Terraform, Ansible, Helm) and GitOps workflows (e.g., Argo CD).
- Exposure to CI/CD pipelines, including pipeline-as-code (e.g., Dagger, GitHub Actions), and Git-based version control.
Benefits
- Paid days off ( i. e. vacatio n, sick days, bereavement leave)
- Health and Dental plans
- Retirement plans
- Employee and Family Assistance Program (EFAP)
- Employee referral program
We welcome applicants from all backgrounds, regardless of race, color, religion, sex, veteran sta tus, sexual orientation, gender identity, national origin, age, or disabilit y or any other protected characteristics in accordance with applicable federal, state/provincial, and local laws . We 're committed to creating a workplace where everyone feels valued and respected.
We appreciate all responses and will acknowledge only those being considered for an interview .
We respectfully request no calls or unsolicited resumes from Agencies .