Senior Site Reliability Engineer

TechChain Talent

New York (NY)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading technology firm is seeking a Senior Site Reliability Engineer to enhance the reliability of their cloud infrastructure. In this role, you'll work on scaling AWS/GCP systems, manage Kubernetes clusters, and streamline deployment processes. Candidates should have at least 5 years of experience in cloud operations, proficiency in SRE methodologies, and strong skills in configuration management tools. This position offers the opportunity to work with innovative technologies and impact the future of cloud-based services.

Qualifications

  • 5+ years of experience as a SRE or DevOps engineer.
  • First-hand experience with infrastructure as code.
  • Production experience with Kubernetes clusters.

Responsibilities

  • Maintain, improve, and secure AWS/GCP infrastructure.
  • Assist teams with packaging and deploying applications.
  • Monitor and improve Kubernetes clusters.

Skills

Cloud-based systems operations
SRE methodologies
Kubernetes management
Configuration management
Infrastructure as code

Tools

Ansible
Terraform
Jenkins

Job description

About the Company

Stellar is a decentralized, public blockchain that gives developers the tools to create experiences that are more like cash than crypto. The network is faster, cheaper, and far more energy-efficient than most blockchain-based systems. It's designed so Stellar's ecosystem can make a real-world, lasting impact.

About the Role

SDF is looking for a Senior Site Reliability Engineer to help build and operate the foundation that powers our engineering teams. You'll ensure the reliability and scalability of our systems, design and improve the infrastructure behind our production environments, and automate operational work so developers can focus on building great products.

Key Responsibilities
  • Maintain, improve, scale and secure our AWS/GCP infrastructure and Linux systems.
  • Assist our development teams in running, packaging, deploying and troubleshooting applications
  • Work with developers on streamlining deployment processes with Jenkins and other CI/CD tooling.
  • Build, maintain, monitor and improve our Kubernetes clusters.
  • Work with development teams on migrating applications to Kubernetes.
  • Be responsible for maintenance and improvements to multiple internal services, for example Kubernetes, Prometheus, ELK.
  • Monitor, triage and respond to alerts in our high availability environments.
  • Participate in design and code reviews, and ensure that the foundation for our services is best in class.
  • Evaluate new technologies, design and implement as appropriate.
  • Identify automation opportunities and implement by creating custom or by using off the shelf solutions.
Requirements

5+ years of experience of working in cloud-based systems operations, as a SRE or DevOps engineer.

First-hand experience with configuration management and infrastructure as code (Ansible, Puppet, Terraform).

Proficient in utilizing SRE methodologies like capacity planning and disaster recovery testing to ensure the scalability, resilience, and availability of critical services.

Production experience building and maintaining Kubernetes clusters.

Will need to know how to code

Bonus Skills
  • Ability to understand Go, Rust, C++ and TypeScript source code
  • Experience experimenting with AI-driven approaches to operations
  • Comfortable with participating in on-call rotations and conducting thorough root cause analyses to keep systems running smoothly.
  • Experienced in managing production workloads and skilled in using monitoring tools to detect issues early.
  • A strong understanding of computer networking, TCP/UDP, load balancing, distributed computing, web services, and the fundamental protocols used by the internet (HTTP, HTTPS, DNS, etc.).
  • No blockchain needed
  • Experience using AI is a plus
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

techchaintalent • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

TechChain Talent • San Francisco (CA)

On-site
USD 120,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

techchaintalent • New York (NY)

On-site
USD 120,000 - 150,000
Senior Cloud SRE - Kubernetes, CI/CD & Automation
Senior Cloud SRE - Kubernetes, CI/CD & Automation

TechChain Talent • New York (NY)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Stellar • San Francisco (CA)

On-site
USD 180,000 - 260,000
Health coverage
Flexible time off
Parental leave
+5
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Stellar • New York (NY)

On-site
USD 180,000 - 240,000
Health, dental, vision insurance
Flexible time off
Parental leave
+5
Senior DevOps / SRE Engineer — Cloud, Kubernetes, AI Ops
Senior DevOps / SRE Engineer — Cloud, Kubernetes, AI Ops

Stellar Cyber • North Carolina

On-site
USD 165,000 - 215,000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Senior DevOps Engineer / SRE — Cloud Native Reliability
Senior DevOps Engineer / SRE — Cloud Native Reliability

Stellar Cyber • New York (NY)

Hybrid
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+1