Site Reliability Engineer

impact.com

Victoria

On-site

CAD 110,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health & prescription coverage
Flexible working with responsible PTO policy
Free Coursera subscription

Job summary

impact.com in Victoria, Canada, is seeking a Site Reliability Engineer to ensure the performance and stability of its core application ecosystem. This role will involve close collaboration with Java and C# teams to manage high-frequency data ingestion pipelines and APIs. The ideal candidate will have strong Java/C# skills, experience with OpenTelemetry, and a deep understanding of the Grafana ecosystem. The position offers a competitive salary range of $110,000 - $130,000 annually, along with benefits such as health coverage, flexible working conditions, and educational support.

Qualifications

  • Strong proficiency in Java or C#.
  • Hands-on experience with OpenTelemetry.
  • Deep experience with the Grafana ecosystem.
  • Understanding of high-volume REST/Graph APIs.
  • Solid understanding of application infrastructure.
  • Ability to prioritize tasks in high-velocity environments.

Responsibilities

  • Implement and maintain OpenTelemetry instrumentation.
  • Build integration tests with third-party social APIs.
  • Enhance Grafana dashboards tracking performance metrics.
  • Drive root-cause analysis for distributed system failures.
  • Leverage tracing data to identify bottlenecks.
  • Debug issues across the entire stack.
  • Analyze application usage patterns for scaling decisions.

Skills

Java or C#
OpenTelemetry
Grafana
API integrations
Systems understanding
High-velocity task prioritization

Education

B.S. in Computer Science or equivalent experience

Tools

Grafana ecosystem (Prometheus, Tempo, Loki)
REST/Graph APIs

Job description

impact.com is the world's leading commerce partnership marketing platform, transforming the way businesses grow by enabling them to discover, manage, and scale partnerships across the entire customer journey. From affiliates and influencers to content publishers, brand ambassadors, and customer advocates, impact.com empowers brands to drive trusted, performance-based growth through authentic relationships. Its award-winning products—Performance (affiliate), Creator (influencer), and Advocate (customer referral)—unify every type of partner into one integrated platform.

Your Role at impact.com

As a Site Reliability Engineer, you'll be the champion of performance and stability for our core application ecosystem. Working closely with our Java and C# engineering squads, you'll ensure that our high-frequency data ingestion pipelines and customer-facing applications meet strict performance and error rate benchmarks.

Your mission is to bridge the gap between code and infrastructure, implementing OpenTelemetry (OTel) standards across the stack to provide deep visibility into how we interact with external social APIs and how our internal services communicate.

What You'll Do
  • OTel Orchestration: Become the architect of our observability pipeline. Implement and maintain OpenTelemetry instrumentation across Java and C# services to ensure high-fidelity traces, metrics, and logs.
  • API Reliability: Build integration tests with third-party social APIs and set up the appropriate monitoring and alerting systems to ensure high availability and reliability.
  • Health & Performance: Build and enhance Grafana dashboards and alerting systems that track the Golden Signals (Latency, Traffic, Errors, Saturation) specifically tailored for JVM and .NET environments.
  • Infrastructure as Code: Drive root-cause analysis (RCA) for complex distributed system failures and contribute to remediations through code optimizations or infrastructure adjustments.
  • Distributed Tracing: Leverage tracing data to identify bottlenecks in cross-service communication and optimize the path of data from social APIs to our internal stores.
  • Full-Stack Troubleshooting: Debug issues across the entire stack, from containerized application code (Java/C#) down to network calls and cloud resource utilization.
  • Capacity Planning: Analyze application usage patterns to inform scaling decisions, ensuring we handle social data bursts without compromising stability or overspending on cloud costs.
What You Bring
  • Software Pedigree: Strong proficiency in Java or C#. Comfortable reading, debugging, and instrumenting application code.
  • Observability Expert: Hands-on experience with OpenTelemetry, including auto-instrumentation, manual spans, and collector configuration.
  • Modern Tooling: Deep experience with the Grafana ecosystem (Prometheus, Tempo, Loki) or similar distributed tracing platforms (Jaeger, Honeycomb, Datadog).
  • API Savvy: Experience working with high-volume REST/Graph APIs and an understanding of OAuth flows, rate-limiting, and webhooks.
  • Systems Mindset: Solid understanding of how Java/C# applications interact with the underlying infrastructure.
  • Pragmatism: Ability to prioritize tasks in a high-velocity environment and a focus on building self-healing systems rather than manual fixes.
  • Education: B.S. in Computer Science or equivalent practical experience in a high-scale production environment.
Salary Range

$110,000 - $130,000 per year, plus an additional 5% variable annual bonus contingent on company performance and eligibility for a Restricted Stock Unit (RSU) grant.

Benefits
  • Strong extended health benefits: health & prescription coverage, vision and dental care, virtual health care, out-of-country medical coverage.
  • Insurance coverage: life, short-term disability, long-term disability, and more.
  • Health Care Spending Account.
  • Two different Employee Assistance Programs.
  • Flexible working: Responsible PTO policy and a flexible environment that supports work-life balance.
  • Health and Wellness: up to 12 fully covered therapy/coaching sessions per year, plus a monthly gym reimbursement policy.
  • Restricted Stock Units (RSUs) as part of total compensation with a 3-year vesting schedule.
  • Investment in learning: free Coursera subscription and PXA courses.
  • Parental support: generous parental leave policy (26 weeks fully paid for primary caregiver, 13 weeks for secondary caregiver).
  • Technology stipend and monthly internet allowance.

We are an equal-opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Silversmith Capital Partners • Victoria

On-site
CAD 110,000 - 130,000
Strong extended health benefits
Flexible Working
Mental health and wellness support
+2
Senior Full Stack Engineer II
Senior Full Stack Engineer II

Silversmith Capital Partners • Victoria

On-site
CAD 135,000 - 150,000
Strong extended health benefits
Flexible working hours
Health and wellness support
+4
Senior Software Engineer
Senior Software Engineer

Silversmith Capital Partners • Victoria

On-site
CAD 135,000 - 170,000
Health benefits
Flexible PTO
RSUs
+4
Senior Full Stack Engineer II
Senior Full Stack Engineer II

impact.com • Victoria

On-site
CAD 110,000 - 130,000
Health benefits
Life insurance
Health Care Spending Account
+6
Senior Software Engineer
Senior Software Engineer

Impact.com • Victoria

On-site
CAD 135,000 - 170,000
RSUs
Health benefits
Parental leave
+3
Sr. Site Reliability Administrator
Sr. Site Reliability Administrator

OpenText • Mississauga

On-site
CAD 129,000 - 194,000
Staff Software Reliability Engineer - Data Platform
Staff Software Reliability Engineer - Data Platform

Okta • Toronto

Hybrid
CAD 160,000 - 220,000
Health insurance
Dental insurance
Vision insurance
+6
Senior Site Reliability Developer
Senior Site Reliability Developer

United States Digital Space LLC • Toronto

On-site
CAD 107,000 - 157,000
Salary transparency
In-person onboarding
Staff Engineer - Analytics & Data Architecture
Staff Engineer - Analytics & Data Architecture

Oscilar • Canada

On-site
CAD 120,000 - 180,000
Stock options
Medical/Dental benefits for you and/de
Caju Card for meal allowance
+3
Sr. Site Reliability Administrator
Sr. Site Reliability Administrator

OpenText • Richmond Hill

On-site
CAD 80,000 - 131,000
Comprehensive benefits package
Variable compensation opportunities