Enable job alerts via email!

Staff Site Reliability Engineer, Platform

Gemini

Singapore

Hybrid

USD 80,000 - 150,000

Full time

3 days ago
Be an early applicant

Boost your interview chances

Create a job specific, tailored resume for higher success rate.

Job summary

An innovative firm is seeking a Staff Site Reliability Engineer to lead engineering teams towards modern DevOps practices. This role involves developing automation and operational tooling, improving system reliability, and mentoring teams on best practices. With a focus on collaboration and continuous improvement, you'll help shape the development culture at a global crypto platform. Join a company that values independence and creativity, and be part of a team that is dedicated to unlocking financial freedom through cutting-edge technology.

Benefits

Comprehensive health plans
Long-term equity incentive
Paid Parental Leave
Competitive paid time off

Qualifications

  • 7+ years of experience in monitoring, alerting, and automation tooling.
  • Expertise in infrastructure as code and containerization technologies.
  • Proficient in scripting for developer productivity.

Responsibilities

  • Provide operational support for various Gemini services.
  • Improve reliability and quality across all services.
  • Guide engineering teams on best practices and tooling.

Skills

Monitoring and Alerting
Automation Tooling
Infrastructure as Code (Terraform)
Containerization (Docker, EKS)
Configuration Management (Ansible, Chef, Puppet)
Scripting (Python, Go)
Cloud Technologies (AWS, GCP, Azure)
Technical Leadership
System Performance Analysis
Teaching and Mentoring

Job description

About the Company

Gemini is a global crypto and Web3 platform founded by Tyler Winklevoss and Cameron Winklevoss in 2014. Gemini offers a wide range of crypto products and services for individuals and institutions in over 70 countries.

Crypto is about giving you greater choice, independence, and opportunity. We are here to help you on your journey. We build crypto products that are simple, elegant, and secure. Whether you are an individual or an institution, we help you buy, sell, and store your bitcoin and cryptocurrency.

At Gemini, our mission is to unlock the next era of financial, creative, and personal freedom.

The Department: Platform

Our Platform organization’s purpose is to enable Gemini to scale effectively and empower our engineering teams to focus on building innovative financial products and experiences for individuals around the world. Platform focuses around building a scalable and secure foundations platform, enabling Engineering to deploy, validate, and operate their services in production, improve resiliency of the service and increase organizational efficiency by reducing operational toil and increase system efficiency through architectural evolution.

The Site Reliability Engineering team engages directly with our other engineering teams to onboard them onto our platform systems, reviewing and recommending design and architectural decisions, and guiding our engineering teams on how to implement the tooling provided by the larger Platform organization required to ensure systems can scale and react to changing conditions, with continuous improvement loops.

The Role: Staff Site Reliability Engineer

You will be an integral part of leading Gemini’s engineering teams towards modern DevOps practices, both by developing and providing modern automation and operational tooling, and working cross-functionally across Gemini’s engineering teams to influence and shape our development practices and culture.

Responsibilities:

  • Provide primary operational support and engineering for various Gemini services
  • Improve reliability, quality and time-to-market across all Gemini services and offerings
  • Guide engineering teams onto the various supported services provided by Platform
  • Run on-going performance evaluations and improvements for Gemini systems
  • Architecture recommendations and engagement as part of SDLC
  • Create “Production-ready Scorecards” to evaluate the health of systems pre-launch
  • Implement and teaching monitoring, alerting and automated resolution best practices
  • Define SLIs, SLOs with Engineering teams
  • Educate and guide Engineering teams on reliability and resiliency best practices, like statelessness, chaos testing, blue/green deployments, etc.
  • Design, build, and maintain operational tooling and automation that streamline processes and enhance system reliability

Qualifications:

  • 7+ years using monitoring, alerting, and automation tooling to understand and remediate performance and health issues in systems at scale
  • Good knowledge for various cloud technology providers like AWS, GCP, or Azure
  • Expert in an infrastructure as code environment (Terraform), developing automated solutions to solve support and operational issues
  • Experience as a Technical Leader within a team, helping evaluating and making tech decisions for the team
  • Expert working with containerization such as Nomad, EKS (k8s), Docker, etc.
  • Expert working with Configuration Management such as Ansible, Chef, Puppet
  • Proficient writing scripts or CLI tools that help increase Developer Productivity in high-level languages like Python, Go, etc.
  • Expert analyzing system and application performance, identifying bottlenecks, and recommending architectural or systemic improvements
  • Experience working with Engineering teams, teaching, training, and mentoring on how to implement best-practice technical solutions

It Pays to Work Here

We take a holistic approach to compensation at Gemini, which includes:

  • Comprehensive health plans covered at 100% for employees and dependents
  • Long-term incentive in the form of a new hire equity grant
  • Paid Parental Leave
  • Competitive paid time off

In Singapore, we have a hybrid work policy. Employees are expected to work from the office part of the week. We believe our hybrid approach increases productivity through more in-person collaboration where possible.

P.S. Mention Tech in Asia Jobs when you apply! Helps keep the good stuff coming

Get your free, confidential resume review.
or drag and drop a PDF, DOC, DOCX, ODT, or PAGES file up to 5MB.