Role Overview
The firm is seeking a Mid-Level Site Reliability Engineer to support mission-critical production systems and observability platforms. This role combines SRE responsibilities with data platform engineering, focusing on system reliability, monitoring, automation, cloud infrastructure, and real-time data pipelines. You will work closely with engineering teams to improve platform stability, operational visibility, and scalability across a complex distributed environmen
Key Responsibilities
- Own and enhance end-to-end observability across production systems.
- Design actionable monitoring dashboards, alerts, and operational metrics.
- Investigate production incidents, diagnose root causes, and implement permanent fixes.
- Manage deployments through CI/CD pipelines and GitOps workflows.
- Operate and optimize Kubernetes-based workloads in cloud environments.
- Provision and maintain infrastructure using Infrastructure-as-Code practices.
- Build and support near real-time data pipelines and telemetry platforms.
- Design data quality controls and optimize analytical databases for large-scale querying.
- Develop and improve SQL-based analytics and reporting capabilities.
- Participate in a shared on-call rotation supporting critical production services.
Key Requirements
- Degree in Computer Science, Software Engineering, or equivalent practical experience.
- Experience building and maintaining observability and monitoring solutions within production environments.
- Strong understanding of incident management, troubleshooting, and root cause analysis.
- Hands-on experience with monitoring tools such as Prometheus, Grafana, ELK, Datadog, CloudWatch, or similar.
- Experience deploying software through CI/CD platforms such as GitLab CI, Jenkins, or GitHub Actions.
- Exposure to distributed stream processing technologies such as Flink, Spark Structured Streaming, or comparable frameworks.
- Experience working with OLAP or columnar databases such as ClickHouse, BigQuery, Druid, or Redshift.
- Practical Kubernetes experience, including deployments, resource management, and application configuration.
- Strong programming skills in Python, Java, Scala, Rust, or similar languages.
- Advanced SQL skills with experience building and optimising analytical workloads.
- Knowledge of messaging platforms and event-driven architectures such as Kafka or SQS
Nice to Have
- Experience with OpenTelemetry and distributed tracing.
- Knowledge of Airflow or workflow orchestration platforms.
- Advanced Terraform and Infrastructure-as-Code expertise.
- Familiarity with JVM build tooling or large-scale software delivery pipelines.
- Prior exposure to trading, market data, financial technology, or other low-latency environments.