SME - Kubernetes, Terraform

HCL Technologies Limited

Mississauga

On-site

CAD 100,000 - 150,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

HCLTech is seeking an AWS SRE Administrator to ensure reliability and performance of applications and infrastructure on AWS. You will monitor services, manage incidents, and drive automation and operational excellence.

Collaborate with application, cloud, infrastructure, security, and operations teams to optimize platform health and implement runbooks, dashboards, and capacity planning. This role offers growth within a global tech environment.

Qualifications

  • Experience administering AWS-based platforms and services.
  • Strong knowledge of monitoring, incident response, and runbooks.
  • Ability to work across security, infrastructure, and application teams.

Responsibilities

  • Monitor AWS-hosted apps and services to ensure high availability.
  • Perform health checks, alerting, and incident triage.
  • Manage AWS services including EC2, EKS, ECS, Lambda, RDS, S3, Route 53, CloudWatch.
  • Investigate incidents and coordinate resolutions with teams.
  • Develop and maintain runbooks and SOPs.

Skills

Monitoring
Incident triage
Automation
Observability
SRE practices

Tools

EC2
EKS
ECS
Lambda
RDS
S3
Route 53
CloudWatch

Job description

AWS SRE Admin: The SRE Administrator is responsible for maintaining the reliability, availability, performance, and operational stability of applications and infrastructure hosted on AWS. The role focuses on monitoring, incident management, platform administration, observability, automation, and operational excellence to ensure seamless business services across the technology estate. The AWS SRE Administrator works closely with application, cloud, infrastructure, security, and operations teams to proactively identify issues, optimize platform performance, and drive continuous service improvements. Key Responsibilities: Monitor AWS-hosted applications, infrastructure, and services to ensure high availability and performance. Perform proactive health checks, alert monitoring, incident triage, and operational support activities. Manage AWS services including EC2, EKS, ECS, Lambda, RDS, S3, Route 53, CloudWatch, and related platform components. Investigate production incidents, perform root cause analysis, and coordinate resolution with support and engineering teams. Administer observability platforms and monitoring tools, including dashboard maintenance, alert tuning, and reporting. Support production releases by performing deployment validation, smoke testing, and post-change monitoring activities. Develop and maintain operational runbooks, standard operating procedures, and knowledge documentation. Automate repetitive operational tasks using AWS native services, scripting, and infrastructure-as-code tools. Monitor platform capacity, utilization trends, and system reliability metrics to support capacity planning. Collaborate with security, infrastructure, and application teams to ensure compliance with operational and security standards. Generate operational reports, SLA/KPI dashboards, and service health updates for stakeholders. Drive continuous improvement initiatives focused on reliability, automation, operational efficiency, and reduction of manual effort.

Key Responsibilities

AWS SRE Admin: The SRE Administrator is responsible for maintaining the reliability, availability, performance, and operational stability of applications and infrastructure hosted on AWS. The role focuses on monitoring, incident management, platform administration, observability, automation, and operational excellence to ensure seamless business services across the technology estate. The AWS SRE Administrator works closely with application, cloud, infrastructure, security, and operations teams to proactively identify issues, optimize platform performance, and drive continuous service improvements. Key Responsibilities: Monitor AWS-hosted applications, infrastructure, and services to ensure high availability and performance. Perform proactive health checks, alert monitoring, incident triage, and operational support activities. Manage AWS services including EC2, EKS, ECS, Lambda, RDS, S3, Route 53, CloudWatch, and related platform components. Investigate production incidents, perform root cause analysis, and coordinate resolution with support and engineering teams. Administer observability platforms and monitoring tools, including dashboard maintenance, alert tuning, and reporting. Support production releases by performing deployment validation, smoke testing, and post-change monitoring activities. Develop and maintain operational runbooks, standard operating procedures, and knowledge documentation. Automate repetitive operational tasks using AWS native services, scripting, and infrastructure-as-code tools. Monitor platform capacity, utilization trends, and system reliability metrics to support capacity planning. Collaborate with security, infrastructure, and application teams to ensure compliance with operational and security standards. Generate operational reports, SLA/KPI dashboards, and service health updates for stakeholders. Drive continuous improvement initiatives focused on reliability, automation, operational efficiency, and reduction of manual effort.

Skill Requirements

AWS SRE Admin: The SRE Administrator is responsible for maintaining the reliability, availability, performance, and operational stability of applications and infrastructure hosted on AWS. The role focuses on monitoring, incident management, platform administration, observability, automation, and operational excellence to ensure seamless business services across the technology estate. The AWS SRE Administrator works closely with application, cloud, infrastructure, security, and operations teams to proactively identify issues, optimize platform performance, and drive continuous service improvements. Key Responsibilities: Monitor AWS-hosted applications, infrastructure, and services to ensure high availability and performance. Perform proactive health checks, alert monitoring, incident triage, and operational support activities. Manage AWS services including EC2, EKS, ECS, Lambda, RDS, S3, Route 53, CloudWatch, and related platform components. Investigate production incidents, perform root cause analysis, and coordinate resolution with support and engineering teams. Administer observability platforms and monitoring tools, including dashboard maintenance, alert tuning, and reporting. Support production releases by performing deployment validation, smoke testing, and post-change monitoring activities. Develop and maintain operational runbooks, standard operating procedures, and knowledge documentation. Automate repetitive operational tasks using AWS native services, scripting, and infrastructure-as-code tools. Monitor platform capacity, utilization trends, and system reliability metrics to support capacity planning. Collaborate with security, infrastructure, and application teams to ensure compliance with operational and security standards. Generate operational reports, SLA/KPI dashboards, and service health updates for stakeholders. Drive continuous improvement initiatives focused on reliability, automation, operational efficiency, and reduction of manual effort.

Other Requirements

AWS SRE Admin: The SRE Administrator is responsible for maintaining the reliability, availability, performance, and operational stability of applications and infrastructure hosted on AWS. The role focuses on monitoring, incident management, platform administration, observability, automation, and operational excellence to ensure seamless business services across the technology estate. The AWS SRE Administrator works closely with application, cloud, infrastructure, security, and operations teams to proactively identify issues, optimize platform performance, and drive continuous service improvements. Key Responsibilities: Monitor AWS-hosted applications, infrastructure, and services to ensure high availability and performance. Perform proactive health checks, alert monitoring, incident triage, and operational support activities. Manage AWS services including EC2, EKS, ECS, Lambda, RDS, S3, Route 53, CloudWatch, and related platform components. Investigate production incidents, perform root cause analysis, and coordinate resolution with support and engineering teams. Administer observability platforms and monitoring tools, including dashboard maintenance, alert tuning, and reporting. Support production releases by performing deployment validation, smoke testing, and post-change monitoring activities. Develop and maintain operational runbooks, standard operating procedures, and knowledge documentation. Automate repetitive operational tasks using AWS native services, scripting, and infrastructure-as-code tools. Monitor platform capacity, utilization trends, and system reliability metrics to support capacity planning. Collaborate with security, infrastructure, and application teams to ensure compliance with operational and security standards. Generate operational reports, SLA/KPI dashboards, and service health updates for stakeholders. Drive continuous improvement initiatives focused on reliability, automation, operational efficiency, and reduction of manual effort.

At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.

HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026totaled $14.8billion.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SME - AzureMSSQL DBA, Microsoft Azure
SME - AzureMSSQL DBA, Microsoft Azure

HCL Technologies Limited • Mississauga

On-site
CAD 90,000 - 140,000
SME - DynaTrace, Windows PowerShell
SME - DynaTrace, Windows PowerShell

HCL Technologies Limited • Mississauga

On-site
CAD 90,000 - 130,000
Senior Program Manager
Senior Program Manager

HCL Technologies Limited • Toronto

On-site
CAD 90,000 - 120,000
Cloud Reliability Engineer - AWS SRE & Terraform
Cloud Reliability Engineer - AWS SRE & Terraform

HCL Technologies Limited • Mississauga

On-site
CAD 100,000 - 150,000
Senior Infrastructure SRE
Senior Infrastructure SRE

PointClickCare • Mississauga

On-site
CAD 110,000 - 150,000
Subject Matter Expert (Support&Ops)
Subject Matter Expert (Support&Ops)

HCL Technologies Limited • Mississauga

On-site
CAD 17,000 - 26,000
Azure DevOps Technical Specialist
Azure DevOps Technical Specialist

HCL Technologies Limited • Airdrie

On-site
CAD 110,000 - 150,000
Senior Developer
Senior Developer

HCL Technologies Limited • Vancouver

On-site
CAD 70,000 - 110,000
Senior Administrator - Desk Side Services, AMT Asset Management Software
Senior Administrator - Desk Side Services, AMT Asset Management Software

HCL Technologies Limited • Mississauga

On-site
CAD 60,000 - 90,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

iManage • Toronto

On-site
CAD 90,000 - 120,000
Market-competitive salary
Annual performance-based bonus
Comprehensive Health, Vision, Dental, and Life insurance
+4