Site Reliability Engineer

Hazeltree

Hong Kong

On-site

HKD 420,000 - 640,000

Full time

26 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, vision insurance
Retirement plan with company match
Professional development opportunities
Collaborative work environment

Job summary

Hazeltree is seeking a mid-level Site Reliability Engineer to design, scale, and maintain highly available systems on AWS. You will ensure operational reliability, monitoring, containerized workloads, and automation across production environments.

The role collaborates with development teams to improve system reliability, scalability, and performance while maintaining clear runbooks and architecture documentation.

Qualifications

  • 3-5 years hands-on experience in SRE/DevOps or infrastructure engineering.

Responsibilities

  • Design, deploy, and manage AWS infrastructure and cost optimization.

Skills

SRE/DevOps experience
Automation
Networking fundamentals
Troubleshooting under pressure
Independent worker

Tools

Docker
Kubernetes
Terraform
CloudFormation
Ansible
Helm
Linux
Windows Server

Job description

Hazeltree is the leading provider of financial treasury and data solutions to the alternative asset management industry, including hedge funds, private equity, pension and endowment funds, and traditional asset management firms. Hazeltree's unique data networking value proposition delivers significant performance improvements and operational efficiency for its treasury clients. The firm's global operations are based in New York, London, and Hong Kong.

About the Role

We are seeking a mid-level Site Reliability Engineer to support the design, scalability, and ongoing reliability of highly available systems on AWS. This role will be responsible for infrastructure reliability, monitoring, containerized workloads, and automation, working closely with development teams to ensure stable and efficient production operations.

Key Responsibilities
  • Design, deploy, and manage AWS infrastructure, including EC2, VPC, S3, RDS, IAM, Security Groups, Load Balancers, Lambda, and cost optimisation.
  • Administer and troubleshoot Linux systems, including Ubuntu, RHEL, CentOS, and Amazon Linux, across production and staging environments.
  • Deploy, manage, and troubleshoot containerized applications using Docker and Kubernetes, including deployments, services, ingress, Helm charts, and cluster upgrades.
  • Build and maintain infrastructure monitoring solutions for CPU, memory, disk, network, and availability metrics using tools such as CloudWatch, Prometheus, Grafana, LogicMonitor, or Datadog.
  • Implement application performance monitoring (APM) to track latency, errors, and throughput (e.g., Logic Monitor, ELK/EFK stack)
  • Establish alerting and on-call escalation workflows to identify and address issues before they affect customers.
  • Lead automation initiatives by replacing manual processes with scripts and Infrastructure as Code tools such as Terraform, Ansible, CloudFormation, Bash, or Python.
  • Participate in incident response, root cause analysis, and post-incident reviews.
  • Define and track SLIs/SLOs/error budgets
  • Collaborate with development teams to improve system reliability, scalability, and performance.
  • Maintain clear documentation for runbooks, architecture, and operational procedures.
Required Skills & Experience
  • 3-5 years of hands-on experience in an SRE, DevOps, or infrastructure engineering role
  • Strong experience with AWS services (EC2, VPC, S3, IAM, RDS, ELB, Auto Scaling, Lambda)
  • Practical experience with containers and Kubernetes, including building images, writing Dockerfiles, deploying and scaling workloads, managing Kubernetes clusters, preferably EKS, and using Helm.
  • Experience with infrastructure and application monitoring tools (CloudWatch, Prometheus, Grafana, LogicMonitor, New Relic, ELK/EFK, etc.)
  • Proficiency in automation and scripting.
  • Experience with Infrastructure as Code (Terraform, CloudFormation, or Ansible)
  • Strong understanding of networking fundamentals, including DNS, load balancing, TCP/IP, and firewalls.
  • Reasonable working knowledge of Windows Server administration, sufficient to support existing on premises/hybrid Windows infrastructure and assist internal IT operations as needed.
  • Strong understanding of relational and non-relational databases (e.g., PostgreSQL, MySQL, MongoDB, DynamoDB), including query performance tuning, indexing, backups, and replication.
  • Experience with incident management and on-call practices
  • Strong troubleshooting and problem-solving skills under pressure, with the ability to work independently on moderately complex issues.
Nice to Have
  • AWS certifications (Solutions Architect, SysOps Administrator) or CKA/CKAD
  • Experience with service mesh technologies such as AWS Lattice, Istio or Linkerd.
  • Exposure to security best practices and compliance frameworks
  • Experience working in a 24/7 production environment
  • Basic awareness of data warehousing concepts (e.g., ETL/ELT pipelines, data modeling) is a plus.
What We Offer
  • Competitive salary and performance-based bonuses.
  • Comprehensive health, dental, and vision insurance plans.
  • Retirement savings plan with company match.
  • Opportunities for professional development and career advancement.
  • A collaborative and supportive work environment.

We are unable to consider candidates who require sponsorship or visa-related support, now or in the future. Hazeltree Fund Services Inc. is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer - Cloud, Kubernetes & Automation
Site Reliability Engineer - Cloud, Kubernetes & Automation

Hazeltree • Hong Kong

On-site
HKD 420,000 - 640,000
Health, dental, vision insurance
Retirement plan with company match
Professional development opportunities
+1
Site Reliability Engineer (Software)
Site Reliability Engineer (Software)

Leadingnation • Hong Kong

On-site
HKD 12,677,000 - 17,151,000
Cutting-edge tech
Reports to technologists
Top asset for techies
+4
Core Site Reliability Engineer
Core Site Reliability Engineer

Selby Jennings • Hong Kong

On-site
HKD 900,000 - 1,200,000
Infrastructure Engineer - Eclipse Trading
Infrastructure Engineer - Eclipse Trading

Leadingnation • Hong Kong

On-site
HKD 500,000 - 700,000
Dynamic work environment
Collaborative team structure
Work-life balance
+2
Site Reliability Engineer (SRE) / DevOps Engineer
Site Reliability Engineer (SRE) / DevOps Engineer

Tek Systems • Hong Kong

On-site
HKD 420,000 - 660,000
Site Reliability Engineer - Global Investment Bank - Hong Kong
Site Reliability Engineer - Global Investment Bank - Hong Kong

NLS Executive Search • Hong Kong

On-site
HKD 80,000 - 120,000
Site Reliability Engineer (SRE) / DevOps Engineer
Site Reliability Engineer (SRE) / DevOps Engineer

TEKsystems • Hong Kong

On-site
HKD 500,000 - 900,000
Site Reliability Engineer
Site Reliability Engineer

Tribus • Hong Kong

On-site
HKD 900,000 - 1,500,000
Core DevOps Engineer
Core DevOps Engineer

Selby Jennings • Hong Kong

On-site
HKD 900,000 - 1,300,000
Relocation support
Sr. Manager, Site Reliability & Innovation, IT
Sr. Manager, Site Reliability & Innovation, IT

CLSA • Hong Kong

On-site
HKD 900,000 - 1,200,000