VP, Site Reliability Engineer

Crypto Pro Network

New York (NY)

On-site

USD 150,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading digital assets firm in New York City is seeking a Senior Site Reliability Engineer (SRE) with extensive experience in AWS and containerized infrastructure. The ideal candidate will be responsible for architecting and maintaining scalable AWS infrastructure, optimizing EKS/Kubernetes workloads, and ensuring reliability through best practices in observability and incident response. This position requires strong communication skills and a solid background in cloud technologies to support the firm's innovative solutions in the digital economy.

Qualifications

  • 8+ years in SRE, DevOps, or Infrastructure Engineering.
  • Deep hands-on expertise in AWS, Kubernetes, and containerization.
  • Extensive experience with Infrastructure as Code (Terraform) and cloud-native automation.

Responsibilities

  • Architect, deploy, and maintain AWS-based infrastructure.
  • Drive optimization of EKS and Kubernetes.
  • Support migration from legacy VMs to containers.

Skills

AWS expertise
Containerization
Infrastructure as Code (Terraform)
Kubernetes/EKS
Observability stacks (Datadog, Prometheus)
Analytical skills
Incident management
Clear communication

Tools

Datadog
Terraform
OpenTelemetry
Grafana

Job description

Who We Are

Galaxy is a global leader in digital assets and data center infrastructure, delivering solutions that accelerate progress in finance and artificial intelligence. We believe that blockchain and digital asset innovation will transform how value moves through the world – and we’re building the products and services to make that future a reality. Our institutional digital assets platform spans trading, investment banking, asset management, staking, self-custody, and tokenization technology. We also invest in and operate cutting-edge data center infrastructure to power AI and high-performance computing, addressing the growing demand for scalable energy and compute in the U.S. We work at the intersection of finance and technology, helping institutions, startups, and developers navigate a digitally native economy. Led by CEO and Founder Michael Novogratz, our team blends deep crypto expertise with institutional experience and a shared commitment to shaping the future of Web3 and AI. Galaxy is headquartered in New York City, with offices across North America, Europe, the Middle East, and Asia. To learn more about our businesses and products, visit www.galaxy.com.

What We Value

We are a diverse team of free thinkers, and fast movers united to help investors and creators energize the global economy. We are looking for individuals who thrive in a culture of builders and overachievers and embrace high performance, transparent feedback, and a mission-first approach. Our culture shapes our way of working and gets us where we want to be.

  • Be Selective To Be Effective.
  • Be Highly Aligned, Loosely Coupled.
  • Disagree Transparently.
  • Build Dream Teams.
Who You Are

You are a Senior SRE specializing in AWS and containerized infrastructure. You thrive working hands‑on, tackling migration from legacy VMs to container ecosystems with a focus on EKS, automation, and reliability.

What You’ll Do
Reliability Engineering
  • Architect, deploy, and maintain robust, scalable, secure AWS-based infrastructure.
  • Drive adoption and optimization of EKS and Kubernetes for containerized workloads.
  • Support migration initiatives, moving workloads from legacy VMs to containers in AWS.
  • Implement and fine‑tune SLOs, SLAs, and error budgets to balance innovation and stability.
  • Collaborate on best practices with Security and Engineering teams for workload reliability.
Automation & Infrastructure as Code
  • Build Infrastructure as Code (IaC) with Terraform; maintain compliant, repeatable environments.
  • Enhance CI/CD pipelines for efficient, secure, and reliable cloud delivery.
  • Develop and refine automated solutions for autoscaling, failover, and disaster recovery.
Observability & Incident Response
  • Design and implement metrics, logging, and tracing tools (Datadog, OpenTelemetry).
  • Set up robust monitoring and alerting to proactively detect and address failures.
  • Lead incident analysis and post‑mortems; drive improvements in operational playbooks.
  • Serve as a subject matter expert for AWS, EKS, and cloud‑native tooling within the SRE team.
  • Optimize AWS resources, cost management, and resiliency best practices.
  • Ensure secure key management and regulatory compliance for decentralized workloads.
What We’re Looking For
  • 8+ years in SRE, DevOps, or Infrastructure Engineering (IC capacity preferred).
  • Deep hands‑on expertise in AWS, Kubernetes/EKS, and containerization.
  • Extensive IaC experience (Terraform) and cloud‑native automation.
  • Proven track record migrating VM‑based workloads to containers in AWS at scale.
  • Strong experience with observability stacks (Datadog, Prometheus, Grafana, OpenTelemetry).
  • Excellent analytical, problem‑solving, and incident management abilities.
  • Clear communicator who thrives in team environments, collaborating cross‑functionally.
Bonus Points
  • Experience supporting blockchain infrastructure is a strong plus.

Galaxy respects diversity and seeks to provide equal employment opportunities to all employees and job applicants for employment without regard to actual or perceived age, race, color, creed, religion, sex or gender (including pregnancy, childbirth, lactation and related medical conditions), gender identity or gender expression (including transgender status), sexual orientation, marital or partnership or caregiver status, ancestry, national origin, citizenship status, disability, military or veteran status, protected medical condition as defined by applicable state or local law, genetic information or predisposing genetic characteristic, or other characteristic protected by applicable federal, state, or local laws and ordinances.

We will endeavor to make a reasonable accommodation to the known limitations of a qualified applicant with a disability unless the accommodation would impose an undue hardship on the operation of our business. If you believe you require such assistance to complete the application process or to participate in an interview, please contact

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Vice President Site Reliability Engineering (Data Centers)
Vice President Site Reliability Engineering (Data Centers)

Crypto Pro Network • New York (NY)

On-site
USD 160,000 - 220,000
Vice President, Site Reliability Engineer (Automation & Network Focus)
Vice President, Site Reliability Engineer (Automation & Network Focus)

Crypto Pro Network • Dallas (TX)

On-site
USD 150,000 - 200,000
Competitive base salary
Company paid sick leave
Health benefits
+3
Vice President, Site Reliability Engineer (Automation & Network Focus)
Vice President, Site Reliability Engineer (Automation & Network Focus)

Crypto Pro Network • New York (NY)

On-site
USD 150,000 - 200,000
Competitive base salary
Discretionary bonus
Company paid health benefits
+4
VP, Protocol / Backend Engineer
VP, Protocol / Backend Engineer

Galaxy • New York (NY)

Hybrid
USD 100,000 - 130,000
Competitive base salary and discretionary bonus
Paid time off
Company-paid health benefits
+2
Vice President, Network Engineer
Vice President, Network Engineer

Crypto Pro Network • Dallas (TX)

On-site
USD 95,000 - 120,000
Competitive base salary and discretionary bonus
Company paid sick leave
Health benefits for employees and dependents
+6
VP, Security Engineer
VP, Security Engineer

Galaxy • New York (NY)

Hybrid
USD 100,000 - 130,000
Competitive base salary and discretionary bonus
Flexible Time Off (paid)
Company-paid health benefits
+1
Protocol Engineer - Platform
Protocol Engineer - Platform

Galaxy • New York (NY)

On-site
USD 200,000 - 250,000
Competitive base salary and discretionary bonus
Flexible Time Off
Company-paid health benefits
Electrical Engineering Lead
Electrical Engineering Lead

Crypto Pro Network • Dallas (TX)

On-site
USD 120,000 - 150,000
Competitive base salary
Flexible Time Off
Company paid health benefits
+2
Technical Project Manager (Data Centers)
Technical Project Manager (Data Centers)

Galaxy • Dallas (TX)

On-site
USD 90,000 - 120,000
Competitive base salary
Flexible Time Off
Company paid Holidays
+7
VP, Forward Deployed Lead
VP, Forward Deployed Lead

Galaxy • New York (NY)

Hybrid
USD 215,000 - 285,000
Competitive base salary and discretionary bonus
Hybrid/Flexible Working Arrangements
Flexible Time Off
+6