Site Reliability Engineer - Data Platform

Unchain Data

United States

Remote

USD 140,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Kraken is seeking a Senior Site Reliability Engineer specialized in Data Infrastructure to design, build, and operate the foundational data platform. You will collaborate with cross-functional teams to ensure reliable, scalable, and cost-efficient data services across our lakehouse architecture.

You will implement IaC, manage real-time streaming with Kafka and Debezium, and provide BI tooling support while maintaining security and RBAC across multi-tenant environments.

Qualifications

  • 5+ years of experience as SRE/Infrastructure/Data Infra.
  • Experience with Kafka, Flink, and Debezium for real-time data processing.
  • Hybrid multi-tenant cloud systems on AWS.
  • IaC tools such as Terraform, Terragrunt, and Atlantis.
  • Containerization/orchestration using Kubernetes, Nomad, or Docker.
  • Scripting: Bash or Python; CI/CD familiarity.
  • Data infra tech: Airflow, Spark, databases, BI tooling.
  • Strong problem-solving and incident response.

Responsibilities

  • Design and implement data infrastructure for 10+ business units.
  • Develop IaC to provision on-prem and AWS.
  • Develop automation scripts for ops tasks.
  • Improve CI/CD pipelines for data infra.
  • Monitor, alerting, and incident response; on-call.
  • Manage real-time streaming with Kafka, Debezium.
  • Use Kubernetes for deployment and scaling.
  • Collaborate with data engineers and analysts.
  • Document architecture and best practices.
  • Support AI/ML teams infra requests.

Skills

Kafka
Flink
Debezium
AWS
Kubernetes
Terraform
Terragrunt
Atlantis
Python
Shell scripting
CI/CD
RBAC
Security

Tools

Terraform
Terragrunt
Atlantis
Kubernetes
Docker
Airflow
Spark
Debezium
Kafka

Job description

About Us

Kraken is a mission-focused company rooted in crypto values. Our Krakenites are a world-class team with crypto conviction, united by our desire to discover and unlock the potential of crypto and blockchain technology. As a Krakenite, you'll join us on our mission to accelerate the global adoption of crypto, so that everyone can achieve financial freedom and inclusion. For over a decade, Kraken's focus on our mission and crypto ethos has attracted many of the most talented crypto experts in the world.

As a fully remote company, we have Krakenites in 70+ countries who speak over 50 languages. Krakenites are industry pioneers who develop premium crypto products for experienced traders, institutions, and newcomers to the space. Kraken is committed to industry-leading security, crypto education, and world-class client support through our products like Kraken Pro, Desktop, Wallet, and Kraken Futures.

The Team

Join our Data Infrastructure team and play a pivotal role in upholding the reliability, scalability, and efficiency of our robust Data platform. As a Senior Site Reliability Engineer (SRE) specialized in Data Infrastructure, you will collaborate closely with diverse cross-functional teams to conceive, execute, and oversee the foundational data infrastructure that empowers our array of applications and services.

The Role

As a key member of our Data Infrastructure team, you will:

  • Design the data governance mechanisms that ensure our lakehouse is easy to interact with, secure and in compliance with all applicable regulations.
  • Implement the infrastructure we use to ingest our data, store it, catalog it with the right metadata and capture its lineage.
  • Provide a state-of-the-art suite of BI tools for multiple teams within the company.
  • Guarantee the availability, high performance, scalability and cost efficiency of our data platform.

Your proficiency in cloud technologies, infrastructure as code, automation, monitoring, logging, user and machine AuthNZ, and certificate management will be instrumental in upholding the exceptional operational standards we set for our services.

Responsibilities
  • Implement data infrastructure solutions (self service) that support the needs of 10+ business units and over 100 engineering and data analysts
  • Utilize Infrastructure as Code (IaC) principles to design, provision, and manage both on-premises and cloud (AWS) infrastructure components using tools such as Terraform
  • Develop and maintain automation scripts using bash/shell scripting to automate operational tasks and deployments
  • Enhance and manage CI/CD pipelines to facilitate consistent software deployments across the data infrastructure
  • Implement robust data monitoring and alerting solutions to proactively detect anomalies and performance issuesManage and implement role-based access control (RBAC) and permissions for a multitude of user groups and machine workflows across different environments
  • Manage and maintain real-time streaming data architecture using technologies like Kafka and Debezium Change Data Capture (CDC)
  • Ensure the timely and accurate processing of streaming data, enabling data analysts and engineers to gain insights from up-to-date information
  • Utilize Kubernetes to manage containerized applications within the data infrastructure, ensuring efficient deployment, scaling, and orchestration
  • Implement effective incident response procedures and participate in on-call rotations
  • Collaborate with data analysts, engineers, and cross-functional teams to understand requirements and implement appropriate solutions
  • Document architecture, processes, and best practices to enable knowledge sharing and support continuous improvement
  • Support AI/ML teams with their infra requests
Requirements
  • Proven experience (5+ years) working as a Site Reliability Engineer, Infrastructure Engineer, Data Infrastructure Engineer, or similar roles, with a focus on data infrastructure and security
  • Experience with maintaining real-time data processing technologies, such as Kafka and Flink clusters and Debezium instances
  • Working experience in managing hybrid multi-tenant cloud systems particularly on AWS
  • Infrastructure as Code tools such as Terraform, Terragrunt and Atlantis
  • Experience with containerization and orchestration tools, particularly Kubernetes, Nomad, and Docker
  • Solid understanding of bash/shell scripting and proficiency in at least one programming language (preferably Python or JVM languages)
  • Experience maintaining data-related technologies: Apache Airflow, Apache Spark, DBs, BI tooling
  • Experience solving data access management issues at large scale data-lake
  • Familiarity with CI/CD deployment pipelines and related tools
  • Strong problem-solving skills and the ability to troubleshoot complex systems
Nice to Have
  • Experience with data-related technologies (databases, data lakes, airflow, spark)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Infrastructure SRE
Senior Data Infrastructure SRE

Unchain Data • United States

Remote
USD 140,000 - 210,000
Data Platform Engineering Manager
Data Platform Engineering Manager

Unchain Data • United States

Remote
USD 180,000 - 280,000
Remote-friendly culture
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Greenwood Village (CO)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Hydrolix • United States

On-site
USD 110,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Knack Solutions • Richmond (VA)

On-site
USD 100,000 - 130,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mike Albert Fleet Solutions • Cincinnati (OH)

On-site
USD 100,000 - 135,000
Site Reliability Engineer
Site Reliability Engineer

chess • United States

Remote
USD 150,000 - 190,000
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

United States Digital Space LLC • Michigan

On-site
USD 150,000 - 190,000