Site Reliability Engineer - NYC

Mistral

New York (NY)

Hybrid

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary and equity
Healthcare: Medical/Dental/Vision for你
401K with match
PTO 18 days
Parking/Transit stipend
Gym reimbursement
Meal stipend
Visa sponsorship
BetterUp coaching

Job summary

Mistral is seeking an experienced Site Reliability Engineer to shape the reliability, scalability, and performance of our platform and customer applications in a hybrid NYC-based setting. You will work with software engineers and research teams to meet high reliability standards.

Lead infrastructure automation, monitor systems, and optimize CI/CD, containers, and orchestration. You will collaborate with AI/ML researchers and support HPC workloads across diverse environments.

Qualifications

  • Master’s degree in Computer Science, Engineering or related field.
  • 7+ years in a DevOps/SRE role.
  • Strong experience with cloud computing and highly available distributed systems.
  • Exposure to reliability issues in critical environments (RCA, on‑call, etc.).
  • Experience with reliability KPIs (observability, alerts, SLAs).
  • Hands‑on with CI/CD, containerization, orchestration (Docker, Kubernetes).
  • Knowledge of monitoring/logging/observability tools (Prometheus, Grafana, ELK, Datadog).
  • Familiarity with IaC tools like Terraform or CloudFormation.
  • Proficiency in scripting (Python, Go, Bash).
  • Strong networking, security, and system administration concepts.
  • Excellent problem‑solving and communication skills.
  • Self‑motivated in a fast‑paced startup.

Responsibilities

  • Design, build, and maintain scalable, highly available infrastructures for web services and ML workloads.
  • Ensure high availability of platform, inference, and model training environments across HPC clusters.
  • Operate production systems, troubleshoot incidents, and perform on‑call rotations.
  • Develop and improve monitoring, alerting, and incident response.
  • Build and maintain CI/CD, containers, orchestration, logging, and dashboards for APIs and training runs.
  • Collaborate with AI/ML researchers to enable reproducible experiments and safe workflows.
  • Automate infrastructure using Kubernetes, Flux, and Terraform; contribute to tooling and docs.

Skills

DevOps/SRE experience
Cloud computing
Observability
CI/CD
Scripting (Python/Go/Bash)
Networking & security
On-call incident response

Education

Master’s degree in Computer Science or related field

Tools

Docker
Kubernetes
Flux
Terraform
Prometheus
Grafana
ELK Stack
Datadog

Job description

Role Summary

We are seeking highly experienced Site Reliability Engineers (SRE) to shape the reliability, scalability, and performance of our platform and customer‑facing applications. You will work closely with our software engineers and research teams to ensure our systems meet and exceed our internal and external customers' expectations.

What you will do
Operations
  • Design, build, and maintain scalable, highly available, and fault‑tolerant infrastructures to support our web services and ML workloads
  • Ensure our platform, inference, and model training environments are always highly available and enable seamless replication of work environments across several HPC clusters
  • Operate systems and troubleshoot issues in production environments (interrupts, on‑call responses, user admin, data extraction, infrastructure scaling, etc.)
  • Implement and improve monitoring, alerting, and incident response systems to ensure optimal system performance and minimize downtime
  • Implement and maintain workflows and tools (CI/CD, containerization, orchestration, monitoring, logging, and alerting systems) for both our client‑facing APIs and large training runs
  • Participate occasionally in on‑call rotations to respond to incidents and perform root cause analysis to prevent future occurrences
Development
  • Drive continuous improvement in infrastructure automation, deployment, and orchestration using tools like Kubernetes, Flux, Terraform
  • Collaborate with AI/ML researchers to develop and implement solutions that enable safe and reproducible model‑training experiments
  • Build a cloud‑agnostic platform offering an abstraction layer between science and infrastructure
  • Design and develop new workflows and tooling to improve the reliability, availability, and performance of our systems (automation scripts, refactoring, new API‑based features, web apps, dashboards, etc.)
  • Collaborate with the security team to ensure infrastructure adheres to best security practices and compliance requirements
  • Document processes and procedures to ensure consistency and knowledge sharing across the team
  • Contribute to open‑source projects, research publications, blog articles, and conferences
About you
  • Master’s degree in Computer Science, Engineering or a related field
  • 7+ years of experience in a DevOps/SRE role
  • Strong experience with cloud computing and highly available distributed systems
  • Exposure to site reliability issues in critical environments (issue root cause analysis, in‑production troubleshooting, on‑call rotations…)
  • Experience working against reliability KPIs (observability, alerting, SLAs)
  • Hands‑on experience with CI/CD, containerization, and orchestration tools (Docker, Kubernetes…)
  • Knowledge of monitoring, logging, alerting, and observability tools (Prometheus, Grafana, ELK Stack, Datadog…)
  • Familiarity with infrastructure‑as‑code tools like Terraform or CloudFormation
  • Proficiency in scripting languages (Python, Go, Bash…) and knowledge of software development best practices
  • Strong understanding of networking, security, and system administration concepts
  • Excellent problem‑solving and communication skills
  • Self‑motivated and able to work well in a fast‑paced startup environment

Your application will be all the more interesting if you also have:

  • Experience in an AI/ML environment
  • Experience with high‑performance computing (HPC) systems and workload managers (Slurm)
  • Worked with modern AI‑oriented solutions (Fluidstack, Coreweave, Vast…)
Location & Work Policy

This role is based in our NYC office, and we’re currently considering candidates who either already live in the area or are open to relocating. We strongly believe in the value of in‑person collaboration and we encourage going to the office as much as we can (at least 3 days per week) to create bonds and smooth communication. Our remote policy aims to provide flexibility, improve work‑life balance, and increase productivity.

What we offer
  • Competitive salary and equity
  • Healthcare: Medical/Dental/Vision covered for you and your family
  • 401K: 6% matching
  • PTO: 18 days
  • Transportation: Reimburse office parking charges or $120/month for public transport
  • Sport: $120/month reimbursement for gym membership
  • Meal stipend: $400 monthly allowance for meals
  • Visa sponsorship
  • Coaching: we offer BetterUp coaching on a voluntary basis
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Weekday (YC W21) • New York (NY)

On-site
USD 150,000 - 250,000
Health, dental, vision insurance
Generous PTO
Learning & development
+2
Senior Site Reliability Engineer II
Senior Site Reliability Engineer II

Juniper Square • United States

On-site
USD 165,000 - 195,000
Health, dental, and vision care
Life insurance
Mental wellness coverage
+3
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • North Carolina

On-site
USD 165,000 - 215,000
Pre‑IPO Stock Options
Medical, Dental & Vision care
401(k)
+2
Senior Site Reliability Engineer (SRE) - AI Inftastructure
Senior Site Reliability Engineer (SRE) - AI Inftastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Senior Site Reliability Engineer
Senior Site Reliability Engineer

BetterUp • New York (NY)

Hybrid
USD 164,000 - 205,000
Access to BetterUp coaching
Competitive compensation plan
Medical, dental, and vision insurance
+3
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

Remote
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

hardrockdigital • United States

Hybrid
USD 120,000 - 160,000
Competitive pay and benefits
Flexible vacation allowance
Startup culture with global brand support
+1
Senior DevOps Engineer/Site Reliability Engineer-East Coast
Senior DevOps Engineer/Site Reliability Engineer-East Coast

Stellar Cyber • New York (NY)

Hybrid
USD 165,000 - 215,000
Pre-IPO Stock Options
Medical, Dental & Vision care
401(k)
+1
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Raydar • New York (NY)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Inclusion Services S.A • Chicago (IL)

On-site
USD 90,000 - 130,000
100% company-covered health insurance
401k plan with 4% match
15 days paid time off
+3