A leading technology firm is seeking an experienced Site Reliability Engineer to transform production systems using AI tools and maintain cloud infrastructure on AWS. Candidates should have over four years of SRE experience, strong skills in Kubernetes, and a genuine excitement for AI tooling. Join a hybrid work environment with a focus on collaboration and enjoy competitive compensation, benefits, and opportunities for advancement.
Qualifications
4+ years of experience in SRE or infrastructure roles.
Hands-on Kubernetes experience: deploying, scaling, debugging, and securing clusters.
Genuine excitement about AI tooling.
Responsibilities
Leverage AI-powered tools to monitor and maintain production systems.
Build and operate cloud infrastructure on AWS using Terraform.
Manage and scale Kubernetes clusters.
Skills
SRE experience
AWS expertise
Kubernetes management
Terraform skills
Observability knowledge
Strong debugging skills
Clear communication
Automation mindset
Tools
AWS
Terraform
Kubernetes
Datadog
Prometheus
OpenTelemetry
Job description
A leading technology firm is seeking an experienced Site Reliability Engineer to transform production systems using AI tools and maintain cloud infrastructure on AWS. Candidates should have over four years of SRE experience, strong skills in Kubernetes, and a genuine excitement for AI tooling. Join a hybrid work environment with a focus on collaboration and enjoy competitive compensation, benefits, and opportunities for advancement.