Senior Site Reliability Developer

United States Digital Space LLC

Toronto

On-site

CAD 107,000 - 157,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Salary transparency
In-person onboarding

Job summary

The Senior Site Reliability Developer will lead the design and operations of AWS-hosted cloud infrastructure for a high-scale data platform. You will drive architecture, automation, and reliability across microservices, with strong emphasis on security, observability, and incident response.

You will collaborate with product teams to ensure rapid delivery, implement DR strategies, and continuously improve CI/CD pipelines.

Qualifications

  • Bachelor’s degree or higher in Computer Science, Engineering, or a related field.
  • 5+ years of progressive experience in Site Reliability Engineering, DevOps, or a similar field.
  • Proficiency with managing AWS resources and understanding of networking and security protocols.
  • Expertise in infrastructure as code (IaC) and cloud automation tools such as Terraform, Serverless, and CloudFormation.
  • Expertise in defining and building CI/CD processes with tools like Jenkins, GitHub, and Artifactory.
  • Experience with container-based technologies like Docker, Kubernetes and AWS ECS.
  • Experience with monitoring and logging tools such as Dynatrace, Grafana, DataDog, ELK Stack, and CloudWatch.

Responsibilities

  • Lead architecture, solution design, development and maintenance of cloud infrastructure for microservices architecture
  • Independently manage requirement analysis, solution design, implementation, and release planning
  • Ensure high adherence to trust and security compliance, guidelines and standards
  • Streamline CI/CD processes, improve system reliability, and ensure infrastructure scalability and security
  • Automate infrastructure deployment, scaling, and management using modern DevOps tools and practices
  • Implement and maintain configuration management and infrastructure as code (IaC) using Terraform
  • Lead Disaster Recovery (DR) strategies, failover exercises, gamedays, and period maintenance activities
  • Contribute to critical vulnerability (CVEs) remediation efforts
  • Promote and document security and best practices across all pillars of DevOps/SRE throughout system design
  • Provide real-time operational support and collaborate across functions to resolve system, infrastructure, and CI/CD issues
  • Participate in on-call rotations, providing critical 24x7 support for production systems

Education

Bachelor’s degree or higher in Computer Science, Engineering, or a related field

Tools

Terraform
Serverless
CloudFormation
Jenkins
GitHub
Artifactory
Docker
Kubernetes
AWS ECS
Dynatrace
Grafana
DataDog
ELK Stack
CloudWatch
Kibana
Open Search
Kafka
Flink
Jira
Google Apigee
ServiceNow
Splunk

Job description

Position Overview

We are seeking a highly motivated and experienced Senior Site Reliability Developer (SRE) to manage critical cloud infrastructure and site reliability operations for the the company Platform Services and Emerging Technologies organization. The team delivers high-value, exabyte-scale and cloud data platform components powering desktop, mobile, and web products. This enables our product teams to build cohesive in-product data experiences, our partners to integrate and expand our data, and our end-users to work with their data across all the company products. This pivotal role focuses on ensuring the highest reliability, availability, and performance of our AWS-hosted cloud infrastructure. Reporting to the Engineering Manager, you will be leading design and development of resilient and scalable architecture and innovative solutions for the platform. You will independently manage and deliver end-to-end solutions while engaging with key stakeholders and partners.

Responsibilities
  • Lead architecture, solution design, development and maintenance of cloud infrastructure for microservices architecture
  • Independently manage requirement analysis, solution design, implementation, and release planning
  • Ensure high adherence to trust and security compliance, guidelines and standards
  • Streamline CI/CD processes, improve system reliability, and ensure infrastructure scalability and security
  • Automate infrastructure deployment, scaling, and management using modern DevOps tools and practices
  • Implement and maintain configuration management and infrastructure as code (IaC) using Terraform
  • Lead Disaster Recovery (DR) strategies, failover exercises, gamedays, and period maintenance activities
  • Contribute to critical vulnerability (CVEs) remediation efforts
  • Promote and document security and best practices across all pillars of DevOps/SRE throughout system design
  • Provide real-time operational support and collaborate across functions to resolve system, infrastructure, and CI/CD issues
  • Participate in on-call rotations, providing critical 24x7 support for production systems
Minimum Qualifications
  • Bachelor’s degree or higher in Computer Science, Engineering, or a related field
  • 5+ years of progressive experience in Site Reliability Engineering, DevOps, or a similar field
  • Proficiency with managing AWS resources and understanding of networking and security protocols
  • Expertise in infrastructure as code (IaC) and cloud automation tools such as Terraform, Serverless, and CloudFormation
  • Expertise in defining and building CI/CD processes with tools like Jenkins, GitHub, and Artifactory
  • Experience with container-based technologies like Docker, Kubernetes and AWS ECS
  • Experience with monitoring and logging tools such as Dynatrace, Grafana, DataDog, ELK Stack, and CloudWatch
  • Technology Stack: Java/SpringBoot, AWS (ECS Fargate, Elastic Cache, Lambda, Kinesis, DynamoDB, VPC, IAM policies, API Gateway, NLB/ALB, Route 53, CloudWatch, Kibana, Open Search), Kafka, Flink, Jenkins, GitHub, Jira, Google Apigee, ServiceNow, and Splunk
Preferred Qualifications
  • Knowledge in applying AI and ML solutions for engineering processes and/or DevOps automation
  • Knowledge of standardized observability frameworks such as OpenTelemetry
  • Relevant certifications (e.g., AWS Certified DevOps Engineer, AWS Site Reliability Engineer)
  • Broad knowledge of AWS, Redis, server programming, databases, and cloud architectures
  • Broad knowledge with data streaming pipelines like Kinesis, Firehose, and Kafka
  • Knowledge on core Java and SpringBoot concepts in JVM optimization
  • Knowledge on build tools, e.g. Gradle
  • Strong interpersonal and communication skills to effectively collaborate in an Agile/Scrum-oriented environment
  • Self-directed team player and independent contributor, demonstrating accountability and end-to-end ownership
  • Experience in Linux Systems Administration, scripting, and troubleshooting in a production environment
  • Proficiency in programming languages such as UNIX, Python, Go, Bash, Groovy, and Node.js
About the company

Welcome to the company! Amazing things are created every day with our software – from the greenest buildings and cleanest cars to the smartest factories and biggest hit movies. We help innovators turn their ideas into reality, transforming not only how things are made, but what can be made.

We take great pride in our culture here at the company – it’s at the core of everything we do. Our culture guides the way we work and treat each other, informs how we connect with customers and partners, and defines how we show up in the world.

When you’re an Autodesker, you can do meaningful work that helps build a better world designed and made for all. Ready to shape the world and your future? Join us!

Salary transparency

Salary is one part of the company’s competitive compensation package. For Canada based roles, we expect a starting base salary between $107,000 and $157,300. Offers are based on the candidate’s experience and geographic location, and may exceed this range. In addition to base salaries, our compensation package may include annual cash bonuses, commissions for sales roles, stock grants, and a comprehensive benefits package. Belonging

Belonging

We take pride in cultivating a culture of belonging where everyone can thrive. Learn more here: https://www.the company.com/company/global-belonging

In-Person Onboarding and Identity Verification

This role may require in-person onboarding and/or in-person ID verification.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Developer
Senior Site Reliability Developer

Autodesk • Toronto

On-site
CAD 107,000 - 157,000
Senior Principal Engineer
Senior Principal Engineer

United States Digital Space LLC • Toronto

On-site
CAD 157,000 - 230,000
Senior Site Reliability Developer
Senior Site Reliability Developer

Autodesk, Inc. • Toronto

On-site
CAD 107,000 - 157,000
Senior DevOps Engineer
Senior DevOps Engineer

Quest Global • Vancouver

On-site
CAD 100,000 - 120,000
401(k) matching
Health insurance
Dental insurance
+5
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Morningstar • Toronto

On-site
CAD 90,000 - 133,000
Hybrid work model
Senior DevOps
Senior DevOps

Quartermaster inc. • Toronto

Hybrid
CAD 160,000 - 215,000
30 days of PTO annually
Health, dental, and wellness benefits
Tech allowance benefit
+1
Senior Software Engineer (DevOps) - BC, Canada Onsite (Vancouver, Canada)
Senior Software Engineer (DevOps) - BC, Canada Onsite (Vancouver, Canada)

S27a • Vancouver

Hybrid
CAD 110,000 - 125,000
Internet allowance
Laptop
Annual Bonus
+2
Senior Software Engineer, AI Search
Senior Software Engineer, AI Search

Autodesk • Toronto

On-site
CAD 107,000 - 157,000
Software Developer - Infrastructure
Software Developer - Infrastructure

United States Digital Space LLC • Oakville

Hybrid
CAD 85,000 - 106,000
Flex working arrangements
Home office reimbursement program
Baby bonus & parental leave top up program
+4
Senior Software Developer, Backend
Senior Software Developer, Backend

United States Digital Space LLC • Toronto

On-site
CAD 107,000 - 157,000
Comprehensive benefits package
Bonuses and stock grants