Senior Site Reliability Engineer — NYC, Equity & Visa Support

Mistral

New York (NY)

Hybrid

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary and equity
Healthcare: Medical/Dental/Vision for你
401K with match
PTO 18 days
Parking/Transit stipend
Gym reimbursement
Meal stipend
Visa sponsorship
BetterUp coaching

Job summary

Mistral is seeking an experienced Site Reliability Engineer to shape the reliability, scalability, and performance of our platform and customer applications in a hybrid NYC-based setting. You will work with software engineers and research teams to meet high reliability standards.

Lead infrastructure automation, monitor systems, and optimize CI/CD, containers, and orchestration. You will collaborate with AI/ML researchers and support HPC workloads across diverse environments.

Qualifications

  • Master’s degree in Computer Science, Engineering or related field.
  • 7+ years in a DevOps/SRE role.
  • Strong experience with cloud computing and highly available distributed systems.
  • Exposure to reliability issues in critical environments (RCA, on‑call, etc.).
  • Experience with reliability KPIs (observability, alerts, SLAs).
  • Hands‑on with CI/CD, containerization, orchestration (Docker, Kubernetes).
  • Knowledge of monitoring/logging/observability tools (Prometheus, Grafana, ELK, Datadog).
  • Familiarity with IaC tools like Terraform or CloudFormation.
  • Proficiency in scripting (Python, Go, Bash).
  • Strong networking, security, and system administration concepts.
  • Excellent problem‑solving and communication skills.
  • Self‑motivated in a fast‑paced startup.

Responsibilities

  • Design, build, and maintain scalable, highly available infrastructures for web services and ML workloads.
  • Ensure high availability of platform, inference, and model training environments across HPC clusters.
  • Operate production systems, troubleshoot incidents, and perform on‑call rotations.
  • Develop and improve monitoring, alerting, and incident response.
  • Build and maintain CI/CD, containers, orchestration, logging, and dashboards for APIs and training runs.
  • Collaborate with AI/ML researchers to enable reproducible experiments and safe workflows.
  • Automate infrastructure using Kubernetes, Flux, and Terraform; contribute to tooling and docs.

Skills

DevOps/SRE experience
Cloud computing
Observability
CI/CD
Scripting (Python/Go/Bash)
Networking & security
On-call incident response

Education

Master’s degree in Computer Science or related field

Tools

Docker
Kubernetes
Flux
Terraform
Prometheus
Grafana
ELK Stack
Datadog

Job description

Mistral is seeking an experienced Site Reliability Engineer to shape the reliability, scalability, and performance of our platform and customer applications in a hybrid NYC-based setting. You will work with software engineers and research teams to meet high reliability standards.

Lead infrastructure automation, monitor systems, and optimize CI/CD, containers, and orchestration. You will collaborate with AI/ML researchers and support HPC workloads across diverse environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud Platform SRE for Scalable AI Systems
Senior Cloud Platform SRE for Scalable AI Systems

Mistral AI • Germany (OH)

On-site
USD 130,000 - 210,000
Healthcare coverage
Relocation support
Retirement plans
+2
Site Reliability Engineer, Mistral Cloud
Site Reliability Engineer, Mistral Cloud

Mistral AI • Germany (OH)

On-site
USD 130,000 - 210,000
Healthcare coverage
Relocation support
Retirement plans
+2
Staff Site Reliability Engineer — NYC On-site, Equity
Staff Site Reliability Engineer — NYC On-site, Equity

Menlo Ventures • New York (NY)

On-site
Confidential
Medical, Dental & Vision benefits
Generous parental leave
401(K) with generous company match
+1
Senior Site Reliability Engineer - Equity Trading Platform
Senior Site Reliability Engineer - Equity Trading Platform

Stellent IT LLC • New York (NY)

On-site
USD 130,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Axiom Pursuits • San Francisco (CA)

On-site
USD 150,000 - 180,000
Senior SRE: ML Infra at Scale, Multi-Cloud K8s
Senior SRE: ML Infra at Scale, Multi-Cloud K8s

The Consensus • New York (NY)

On-site
USD 110,000 - 140,000
Competitive compensation
100% insurance coverage
Flexible PTO policy
+3
Site Reliability Engineer - NYC
Site Reliability Engineer - NYC

Mistral • New York (NY)

Hybrid
USD 140,000 - 190,000
Competitive salary and equity
Healthcare: Medical/Dental/Vision for你
401K with match
+6
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Senior Site Reliability Engineer — NYC, Equity & Impact
Senior Site Reliability Engineer — NYC, Equity & Impact

Menlo Ventures • New York (NY)

On-site
Confidential
Medical, Dental & Vision benefits
Generous parental leave
401(K) with company match
+1
Senior Site Reliability Engineer – Remote, Impact & Automation
Senior Site Reliability Engineer – Remote, Impact & Automation

Midwest Startups • United States

On-site
USD 175,000 - 185,000
Market-leading medical, dental, and視on
Stock options
Premium-Tier Origin Financial Wellness
+6