Technical Support Engineer (L2) - Compute

Mistral

San Francisco (CA)

Hybrid

USD 140,000 - 190,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Healthcare coverage
Relocation support
Wellness programs

Job summary

Mistral is building a Compute Support team to ensure reliability, performance, and scalability of GPU clusters in a Kubernetes environment. You will be the first contact for customers and internal teams, providing guidance and troubleshooting for compute-related inquiries in a hybrid role.

As a founding member, you will shape processes, standards, and culture, with on-call rotations and collaboration with SRE teams to optimize IaC and monitoring.

Qualifications

  • 5+ years of experience in system administration, technical support, or infrastructure operations in compute-heavy environments.
  • Deep Linux/Unix expertise with kernel debugging and performance tuning.
  • Hands-on Kubernetes experience with pod/node troubleshooting and networking/storage concepts.
  • Familiarity with container runtimes (Docker, containerd) and basic IaC tooling.
  • Experience with monitoring (Prometheus, Grafana) and log analysis.

Responsibilities

  • Serve as L1/L2 support for Linux, Kubernetes, and compute infrastructure issues.
  • Diagnose performance bottlenecks, hardware failures, and resource contention in distributed systems.
  • Triage requests, provide actionable guidance, and maintain runbooks and post-mortems.
  • Collaborate with SRE teams to improve IaC and tooling (Go-based).
  • Participate in on-call rotations to ensure 24/7 coverage.

Skills

Linux/Unix
Kernel debugging
Kubernetes
Containerization
Docker
Go
Python/Bash
Monitoring tools
On-call / SLAs
System performance tuning

Tools

Prometheus
Grafana
ELK/OpenTelemetry
Terraform/Ansible

Job description

About Mistral

Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector, co-creating customized AI systems that they can run on their terms.

About Mistral

Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector, co-creating customized AI systems that they can run on their terms. We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited. Mistral AI is building a new Compute Support team to ensure the reliability, performance, and scalability of our GPU clusters in a Kubernetes environment. As one of the founding members of this team, you will play a pivotal role in shaping its processes, standards, and culture. In this hybrid L1/L2 support role, you will be the first point of contact for customers and internal teams, providing technical guidance, troubleshooting, and issue resolution for compute-related inquiries. You’ll also serve as the escalation point for complex issues, leveraging your deep systems knowledge, Kubernetes expertise, and operational debugging skills to ensure our AI workloads run smoothly at scale. This is a support-focused role that blends system administration, customer-facing communication, and compute infrastructure expertise. While coding is not the primary focus, familiarity with Go for debugging and automation is a plus.

Key Responsibilities
First Customer’s Point of Contact
  • Act as the first interlocutor for customers and internal teams, providing timely and effective responses to compute-related inquiries.
  • Triage and prioritize incoming requests, ensuring SLAs are met for acknowledgment, first response, and resolution.
  • Gather and analyze initial issue details (e.g., logs, error messages, system metrics) to diagnose problems efficiently.
  • Provide clear, actionable guidance to customers and internal users, including temporary workarounds where applicable.
Technical Support
  • Serve as the L2 escalation point for Linux, Kubernetes, and compute infrastructure issues, providing deep technical troubleshooting for:
    • GPU/TPU workloads (e.g., CUDA errors, memory leaks, job failures).
    • Kubernetes clusters (e.g., pod crashes, node failures, networking misconfigurations).
    • Bare metal and cloud environments (e.g., AWS EC2, GCP VMs, HPC clusters).
  • Diagnose and resolve performance bottlenecks, hardware failures, and resource contention in distributed systems.
  • Analyze system metrics, logs, and traces (e.g., dmesg, journalctl, nvidia-smi, Prometheus, Grafana) to identify root causes of issues.
  • Participate in on-call rotations to provide 24/7 support for critical compute systems
  • Nice to have: Optimize system configurations (e.g., kernel parameters, filesystem tuning, network settings) for high-performance computing (HPC) and AI workloads.
Kubernetes & Containerization
  • Debug Kubernetes clusters with a focus on:
    • Pod and node issues (e.g., CrashLoopBackOff, OOMKilled, ImagePullBackOff).
    • Networking and storage (e.g., CNI plugins, PersistentVolumes, StorageClasses).
    • Resource management (e.g., Requests/Limits, QOS classes, node affinity).
  • Troubleshoot container runtime issues (Docker, containerd) such as image pull failures, OCI compliance, or runtime errors.
  • Collaborate with SRE teams to understand and improve existing IaC (Terraform, Ansible) and Go-based tooling.
Documentation & Process Improvement
  • Create and maintain runbooks, playbooks, and internal documentation for common compute issues (e.g., GPU debugging, Kubernetes troubleshooting).
  • Contribute to post-mortems with actionable follow-ups to prevent recurring incidents.
  • Train internal teams on best practices for compute infrastructure and debugging.
  • Help define and refine support processes as a founding member of the new Compute Support team.
What We’re Looking For
Required Skills & Experience
  • 5+ years of experience in system administration, technical support, or infrastructure operations, with a strong focus on compute-heavy environments (bare metal, cloud, HPC, or virtualization).
  • Deep Linux/Unix expertise:
    • Debugging kernel-level issues (e.g., OOM killer, I/O bottlenecks, CPU throttling).
    • Performance tuning (e.g., sysctl, ulimit, filesystem optimizations).
    • Networking troubleshooting (e.g., iptables, tcpdump, DNS, NFS, SSH).
    • Storage management (e.g., LVM, RAID, NVMe, GPU-local storage).
  • Hands-on Kubernetes experience is mandatory:
    • Debugging pods, nodes, and clusters (e.g., kubectl describe, kubectl logs, crictl).
    • Networking and storage (e.g., CNI plugins, PersistentVolumes).
    • Resource management (e.g., Requests/Limits, QOS classes).
  • Familiarity with containerization (Docker, containerd) and basic IaC awareness Proficiency with monitoring tools (Prometheus, Grafana, ELK, OpenTelemetry).
  • Basic scripting/automation skills (preferably Go, but Bash or Python are acceptable for ad-hoc tasks).
  • Experience with bare metal, virtualization, or HPC environments (e.g., Fluidstack, Coreweave, Vast) is a strong plus.
  • Excellent problem-solving and communication skills, with the ability to explain technical issues clearly and patiently to both technical and non-technical stakeholders.
  • Customer-focused mindset with a passion for resolving issues efficiently and improving system reliability.
  • Ability to work in on-call rotations and handle high-pressure situations with a structured, calm approach.
  • Commitment to meeting tight SLAs for issue acknowledgment, triage, and first response.
What we offer
  • We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.
Privacy Policy

Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Compute Solution Architect
Compute Solution Architect

Engg • New York (NY)

On-site
USD 150,000 - 230,000
Healthcare coverage
Parental leave
Retirement plans
+4
Compute Support Engineer - Kubernetes & Linux Expert
Compute Support Engineer - Kubernetes & Linux Expert

Lindus Health • Paris (TX)

On-site
USD 120,000 - 180,000
Healthcare coverage
Relocation support
Meal and transportation allowances
Senior Compute Support Engineer: Linux & Kubernetes
Senior Compute Support Engineer: Linux & Kubernetes

Mistral • San Francisco (CA)

Hybrid
USD 140,000 - 190,000
Healthcare coverage
Relocation support
Wellness programs
Compute Solution Architect
Compute Solution Architect

Mistral • New York (NY), Northern (KY)

Hybrid
USD 140,000 - 190,000
Healthcare coverage
Relocation support
Retirement plans
+1
Research Platform Engineer
Research Platform Engineer

Mistral • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Healthcare coverage
Parental leave
Retirement plans
+3
AI Compute Engineer
AI Compute Engineer

Mistral • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Healthcare coverage
Parental leave
Retirement plans
+3
Research Engineer, ML Platform
Research Engineer, ML Platform

Mistral • Palo Alto (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Healthcare coverage
Parental leave
Relocation support
+2
Research Engineer, ML Platform
Research Engineer, ML Platform

Socket.dev • Palo Alto (CA)

On-site
USD 180,000 - 280,000
Healthcare coverage
Relocation support
Retirement plans
+2
Mistral Cloud - Software Engineer, Managed Kubernetes
Mistral Cloud - Software Engineer, Managed Kubernetes

United States Digital Space LLC • Paris (TX)

Hybrid
USD 104,000 - 161,000
Healthcare coverage
Relocation support
Wellness programs
Operations Engineer, Fleet Health & Delivery
Operations Engineer, Fleet Health & Delivery

Lindus Health • Palo Alto (CA)

On-site
USD 140,000 - 200,000