Mistral Cloud - Software Engineer, Managed Kubernetes

United States Digital Space LLC

Paris (TX)

Hybrid

USD 104,000 - 161,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Healthcare coverage
Relocation support
Wellness programs

Job summary

Mistral is seeking a Senior Software Engineer to design, build, and operate scalable Kubernetes-based control plane services and operators. You will own end-to-end features, collaborate across teams, and drive reliability initiatives while contributing to core platform tooling.

Ideal candidates have a Master’s degree, 5+ years of experience in software/SRE roles, and strong skills in Golang, containers, and observability tools.

Qualifications

  • Master’s degree in Computer Science, Engineering or a related field.
  • 5+ years in software engineering, DevOps, or SRE roles.
  • Strong Golang coding and modern development practices.
  • Experience with containerization/orchestration and reliability KPIs.

Responsibilities

  • Design, build, and maintain scalable control plane services, operators, and custom controllers for Kubernetes.
  • Cluster lifecycle management including provisioning, upgrades, patching, decommissioning.
  • Monitoring and Observability: implement monitoring, alerting, incident response.
  • Internal tooling: create workflows, APIs, CLIs to empower users and ML/AI teams.
  • Reliability: design resilient systems for large-scale distributed environments.
  • Incident Response: participate in on-call rotations and root cause analysis.

Skills

Golang
Docker
Kubernetes
Observability
On-call
Networking
Security
Sysadmin
CI/CD

Education

Master's degree in Computer Science or related field

Tools

Prometheus
Grafana
ELK Stack
Slurm

Job description

About MistralMistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

More information on Mistral Compute here: https://the company/products/compute

Location and Office Policy

We will prioritize candidates who either reside in one of our main offices (Paris, London, NYC) or are open to relocating. We will also consider remote candidates based in the following countries:

  • EMEA: France, United Kingdom, Germany, Switzerland, Netherlands, Spain, Austria, Poland, Luxembourg
  • Americas: USA, Canada

We strongly believe in the value of in-person collaboration to foster strong relationships and seamless communication within our team. In any case, we ask all new hires to visit our Paris HQ office (accommodation and travelling covered)

  • for the first week of their onboarding
  • then at least 3 days every month
What You Will Do

Software Development: Design, build, and maintain scalable control plane services, operators, and custom controllers for Kubernetes

  • Cluster lifecycle management: Develop automation including provisioning, upgrades, patching, and decommissioning.
  • Monitoring and Observability: Implement and improve monitoring, alerting, and incident response systems to ensure optimal system performance and minimize downtime
  • Internal tooling: Create workflows, tools, APIs, and command-line interfaces (CLIs) to empower customers and ML/AI teams to deploy and monitor inference services efficiently.
  • Reliability: Design resilient systems capable of gracefully handling failures in large-scale distributed environments.
  • Investigation: identify and resolve unique cluster issues at a low level, documenting findings for future reference.
  • Incident Response: Participate in on-call rotations to respond to incidents and perform root cause analysis to prevent future occurrences
What We're Looking For

Master’s degree in Computer Science, Engineering or a related field

  • 5+ years of experience in a similar role (Software Engineer, DevOps, SRE ...)
  • Strong coding proficiency (ideally Golang) and knowledge of software development best practices
  • Mastery of containerization and orchestration tools (Docker, Kubernetes...)
  • Exposure to site reliability issues in critical environments (issue root cause analysis, in-production troubleshooting, on-call rotations...)
  • Experience working against reliability KPIs (observability, alerting, SLAs)
  • Strong problem-solving abilities and attention to detail.
  • Ability to own and deliver end-to-end features with minimal oversight.
  • Excellent communication skills and collaborative attitude.
  • Team-oriented, humble and eager to learn.
  • Strong understanding of networking, security, and system administration concepts
Now, it would be ideal if you had experience with.
  • high-performance computing (HPC) systems and workload managers (Slurm)
  • monitoring, logging, alerting and observability tools (Prometheus, Grafana, ELK Stack...)
  • networking, storage, security, and system administration concepts
What We Offer

We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page.

Privacy Policy

Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.

Find Jobs in France on Arbeitnow

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer, ML Platform
Research Engineer, ML Platform

Mistral • Palo Alto (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Healthcare coverage
Parental leave
Relocation support
+2
AI Compute Engineer
AI Compute Engineer

Mistral • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Healthcare coverage
Parental leave
Retirement plans
+3
Operations Engineer, Fleet Health & Delivery
Operations Engineer, Fleet Health & Delivery

Lindus Health • Palo Alto (CA)

On-site
USD 140,000 - 200,000
Research Engineer, ML Platform
Research Engineer, ML Platform

Socket.dev • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Healthcare coverage
Relocation support
Retirement plans
+2
Research Engineer, Data Infrastructure
Research Engineer, Data Infrastructure

Socket.dev • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Competitive salary and equity
Medical/Dental/Vision coverage
401K with 6% matching
+6
Research Platform Engineer
Research Platform Engineer

Mistral • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Healthcare coverage
Parental leave
Retirement plans
+3
Site Reliability Engineer, Mistral Cloud
Site Reliability Engineer, Mistral Cloud

Mistral AI • Germany (OH)

On-site
USD 130,000 - 210,000
Healthcare coverage
Relocation support
Retirement plans
+2
Applied AI, Forward Deployed Machine Learning Engineer
Applied AI, Forward Deployed Machine Learning Engineer

Mistral • San Francisco (CA)

On-site
USD 140,000 - 190,000
Research Engineer, Full Stack
Research Engineer, Full Stack

Socket.dev • Palo Alto (CA)

On-site
USD 120,000 - 180,000
Research Engineer, Machine Learning
Research Engineer, Machine Learning

Mistral • San Francisco (CA)

On-site
USD 150,000 - 210,000