Site Reliability / Infrastructure Engineer

Medal

New York (NY)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary and equity
Comprehensive medical, dental, vision coverage
401(k) plan
Wellness and fitness perks
Paid parental leave
Generous PTO policy
Daily meals at NYC HQ
Learning and development stipend

Job summary

Medal, based in New York, is seeking a Site Reliability Engineer (SRE) to enhance the reliability and scalability of its infrastructure. The ideal candidate will have hands-on experience with GCP, Kubernetes, and database management, ensuring our platform can handle billions of gaming clips seamlessly.

This role involves managing incident responses, optimizing performance, and collaborating with engineering teams to support growth. Benefits include competitive salary, comprehensive health coverage, and generous PTO.

Qualifications

  • Experience in startups with a focus on rapid growth.
  • Strong judgment in incident response and infrastructure fixes.
  • Deep experience in scaling and sharding databases in production.

Responsibilities

  • Own reliability across GCP infrastructure.
  • Lead incident response and postmortems.
  • Architect and execute database scaling strategies.
  • Manage Terraform-managed GCP environment and Kubernetes configurations.

Skills

Reliability engineering
Scalability experience
Incident response
GCP tools (Kubernetes, IAM)
Terraform
MySQL and Postgres scaling
Elasticsearch
GitHub Actions

Tools

Google Cloud Platform
Terraform
Kubernetes

Job description

The Company
Medal

Medal is the world’s largest and fastest-growing platform for gaming clips, where millions of gamers capture, share, and relive their best moments. Every year, our players record billions of clips, each representing a unique, action-packed highlight. We’re building the next generation of gaming communities: social, monetized, and creator-powered. Our mission is to design products that make sharing, discovering, and connecting around gaming moments seamless and fun.

We raised a seed round of $133M from General Catalyst and Khosla to discover the next generation of intelligence.

The Role

Medal's infrastructure handles billions of clips, video ingestion pipelines, and social features at a massive scale most engineers never get to touch. We're looking for an SRE who cares deeply about reliability and scalability.

The work centers on reliability, incident response, scaling, and making sure our infrastructure keeps up with our growth. You'll own the on-call rotation, drive postmortems, and work directly with engineering teams to meet their infra needs.

The right person probably came through startups and scale-ups. You've been in the room when things broke at 2am, you've scaled databases under pressure, and you know the difference between a durable fix and a patch that buys you a week.

Key Responsibilities
  • Own reliability across our GCP infrastructure: Kubernetes clusters, managed services, and data pipelines, driving measurable improvements to availability and latency

  • Lead incident response end-to-end: on-call rotations, runbooks, postmortems, and the follow-through that makes sure the same thing doesn't happen twice

  • Architect and execute database scaling strategies (sharding, replication, query optimization, and capacity planning) across MySQL and Postgres at meaningful scale

  • Partner with product engineering to translate feature requirements into infrastructure designs that hold up as we grow

  • Manage and evolve our Terraform-managed GCP environment and Kubernetes cluster configurations

  • Own our Elasticsearch cluster end-to-end: capacity planning, sharding strategy, index lifecycle management, version upgrades, and performance tuning at production scale

  • Build and maintain observability across the stack: metrics, dashboards, alerting, and tracing

  • Constantly improve CI/CD reliability and delivery pipelines across GitHub Actions

  • Harden IAM, secrets management, and network segmentation as part of normal infra hygiene

About You
  • You’ve worked at startups and are comfortable in an environment of rapid growth where scaling up is a priority

  • You have great judgment - you know the difference between a durable, sustainable fix vs. a patch that buys you a week

  • You have deep, hands-on experience scaling and sharding relational databases in production environments

  • You know GCP maybe a little too well: Kubernetes, VPC, IAM, Cloud Logging, and the managed services ecosystem

  • You are fluent in Terraform and have owned real infrastructure-as-code at scale

  • You've operated Elasticsearch in production and know how to keep a cluster healthy

  • You have strong incident response instincts: you can work a P0 calmly, communicate clearly under pressure, and run a postmortem that prevents recurrence.

  • You’ve worked with GitHub Actions in a production CI/CD environment.

  • You have excellent communication skills (this is crucial!) and can both flag issues clearly and rapidly during incidents, and lead / write actionable postmortems

Our Stack

Google Cloud Platform

Terraform, Salt, GitHub Actions

Java, Redis, RabbitMQ, ElasticSearch, BigQuery, Kubernetes for backend

Electron+React

C# and C++ for native windows recording & more

Swift for iOS, Kotlin for Android

Benefits
  • Competitive salary and meaningful equity

  • Comprehensive medical, dental, and vision coverage

  • 401(k)

  • Wellness and fitness perks including a Wellhub membership and mental health resources

  • Paid parental leave, fertility and maternal health benefits

  • Generous PTO policy

  • Daily meals and commuter benefits at our NYC HQ in Flatiron

  • Learning and development stipend

Benefits vary by country and employment type.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability / Infrastructure Engineer
Site Reliability / Infrastructure Engineer

General Intuition & Medal • New York (NY)

On-site
USD 120,000 - 160,000
Competitive salary and meaningful equity
Comprehensive medical, dental, and vision coverage
401(k)
+5
Site Reliability / Infrastructure Engineer
Site Reliability / Infrastructure Engineer

Gamedevsofcolorexpo • New York (NY)

On-site
USD 180,000 - 275,000
Comprehensive medical, dental, vision coverage
401(k)
Fertility and parental benefits
+4
Senior Site Reliability Engineer II
Senior Site Reliability Engineer II

Juniper Square • United States

On-site
USD 165,000 - 195,000
Health, dental, and vision care
Life insurance
Mental wellness coverage
+3
Senior Platform Development Engineer
Senior Platform Development Engineer

SentiLink • United States

On-site
USD 140,000 - 210,000
Employer paid health insurance
401(k) plan with employer match
Flexible paid time off
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Juniper Square • United States

Remote
USD 165,000 - 195,000
Health, dental, and vision care
Life insurance
Mental wellness coverage
+4
Site Reliability Engineer - NYC
Site Reliability Engineer - NYC

Mistral • New York (NY)

Hybrid
USD 140,000 - 190,000
Competitive salary and equity
Healthcare: Medical/Dental/Vision for你
401K with match
+6
Senior Systems Software Engineer
Senior Systems Software Engineer

SentiLink • United States

On-site
USD 160,000 - 210,000
Health insurance
401(k) plan with employer match
Flexible paid time off
+2
Software engineer (Full-stack)
Software engineer (Full-stack)

Untitledinbrackets • New York (NY)

Hybrid
USD 150,000 - 230,000
Flexible days in office
Free lunch in the office
Competitive salary and equity
+2
Senior/Lead Backend Engineer
Senior/Lead Backend Engineer

Medal, Highlight, & General Intuition • New York (NY)

On-site
USD 120,000 - 160,000
Senior Software Engineer, Performance & Reliability
Senior Software Engineer, Performance & Reliability

SentiLink • United States

Hybrid
USD 140,000 - 190,000
Employer paid group health insurance
401(k) plan with employer match
Flexible paid time off
+2