Research Engineer, Code Agents Infra

Socket.dev

Palo Alto (CA)

On-site

USD 210,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Healthcare coverage
Relocation support
Wellness programs

Job summary

Mistral is seeking a seasoned Systems/Distributed Infrastructure Engineer to lead end-to-end execution and data infrastructure powering its agentic models and coding assistants. You will help design scalable synthetic data generation, ultra-fast training environments, and high-throughput execution engines across hybrid and multi-cloud clusters.

You will tackle challenges from 1 million concurrent sandboxes to optimize training codebases, pipelines, and orchestration, working with researchers to

Qualifications

  • 4+ years in Systems Engineering/Distributed Systems/Cloud Infra or MLOps.
  • Experience building high-throughput data processing pipelines and generation pipelines for large-scale datasets.
  • Strong proficiency in Python, Go, C++ or Rust; experience profiling and optimizing high-performance ML/backend code.

Responsibilities

  • Design, deploy and operate high-throughput sandboxing for untrusted code across 1M+ environments.
  • Architect data generation pipelines for synthetic code, agent trajectories, rollouts and self-play data.
  • Optimize agent training codebases and distributed runtimes (Python, Go, C++, Rust) to improve GPU utilization and reduce rollout overhead.
  • Reduce sandbox cold-start times with container warm pools and efficient image delivery across hybrid clusters.
  • Develop Kubernetes-native controllers and queuing systems for diverse hardware fleets.
  • Ensure multi-tenant isolation and secure boundary enforcement across runtimes and networks.

Skills

Distributed Systems
Cloud Infra
MLOps
Kubernetes
Container Tech
Python
Go/C++/Rust
4+ years experience

Tools

Kubernetes
Docker
CRIU
Firecracker
gVisor

Job description

About Mistral

Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

The Role

This role focuses on building and operating the end-to-end execution, training, and data infrastructure that powers Mistral’s agentic models and coding assistants. You will be a core contributor to our agent research stack: designing scalable systems for synthetic data generation, building ultra-fast training and RL execution environments, and maintaining high-throughput execution engines.

You will tackle the engineering challenges at every step of the agent lifecycle: from orchestrating 1M+ concurrent and short-lived sandboxes for untrusted code execution to optimizing agent training codebases, distributed trajectory collection pipelines, and dataset processing workflows across massive hybrid and multi-cloud clusters.

What You Will Do
  • Large-Scale Sandboxing Infrastructure: Design, deploy, and operate our high-throughput sandboxing platform, executing LLM-generated untrusted code across over 1 million isolated environments concurrently for model evaluation and interactive RL environments.

  • Agent Data Generation Pipelines: Architect and scale high-throughput pipelines for synthetic code generation, agent trajectories, rollouts, and self-play data collection to power post-training and RL loops.

  • Training Codebase & Systems Optimization: Optimize agent training codebases and distributed execution runtimes (PyTorch, Ray, SLURM/Kubernetes) to minimize multi-step rollout overhead, improve GPU utilization, and eliminate scaling bottlenecks.

  • Low-Latency Orchestration & Warm Pooling: Reduce sandbox cold-start times to sub-second levels using container warm pools, snapshot/restore technology (e.g., CRIU, microVMs), and optimized image delivery layers across hybrid clusters.

  • Multi-Cluster Queueing & Resource Allocation: Implement Kubernetes-native custom controllers, CRDs, and queuing systems to dynamically route short-lived evaluation, synthetic data, and agent execution tasks across diverse hardware fleets.

  • Isolation, Security & Security Boundary: Ensure strict multi-tenant network and process isolation for untrusted agent code using container/sandboxing runtimes (e.g., gVisor, Firecracker) and default-deny network postures.

  • Operational Excellence: Maintain high availability, telemetry, and automated self-healing across millions of transient jobs while participating in on-call rotations for critical agent training and execution pipelines.

What We’re Looking For
  • 4+ years of experience in Systems Engineering, Distributed Systems, Cloud Infrastructure, or MLOps supporting LLM/RL workloads.

  • Data & Pipeline Engineering: Proven experience building high-throughput data processing and generation pipelines for large-scale datasets (e.g., Ray, Spark, custom distributed queues).

  • Deep experience with Kubernetes & Container Tech: Strong expertise writing custom K8s operators/controllers, managing Linux cgroups/namespaces, and optimizing Docker image layers and distribution systems.

  • High-Performance Software Engineering: Advanced proficiency in Python, Go, C++ or Rust, with a track record of profiling and optimizing high-performance ML or backend systems codebases.

  • Sandboxing & Isolation Technologies: Hands-on experience with lightweight virtualization, container runtimes, or WASM (e.g., Docker, gVisor, Firecracker).

  • Queueing & Scheduling: Deep familiarity with task queue systems, resource schedulers, and low-latency queuing architectures for high-volume, short-lived workloads.

  • Comfort with Ambiguity: Passion for working directly alongside AI researchers to rapidly turn frontier agent ideas into scalable, production-­grade infrastructure.

What We Offer

We offer a comprehensive benefits package designed to support your well-being, growth, and work life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page.

Privacy Policy

Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, Code Agents Infra
Research Engineer, Code Agents Infra

Mistral • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Applied Scientist
Applied Scientist

Headline - Asia • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Applied Scientist
Applied Scientist

Mistral • New York (NY)

On-site
USD 120,000 - 180,000
Applied Scientist
Applied Scientist

Mistral • Palo Alto (CA)

On-site
USD 180,000 - 240,000
AI Scientist
AI Scientist

Mistral AI • Warsaw (IN)

On-site
USD 150,000 - 230,000
Healthcare coverage
Relocation support
Wellness programs
+1
Research Engineer, Machine Learning
Research Engineer, Machine Learning

Mistral • San Francisco (CA)

On-site
USD 150,000 - 210,000
Applied AI, Forward Deployed Machine Learning Engineer
Applied AI, Forward Deployed Machine Learning Engineer

Mistral • San Francisco (CA)

On-site
USD 140,000 - 190,000
AI Scientist
AI Scientist

Mistral • San Francisco (CA)

On-site
USD 150,000 - 260,000
Healthcare coverage
Parental leave
Relocation support
+2
AI Scientist
AI Scientist

Mistral • Palo Alto (CA)

On-site
USD 190,000 - 260,000
Healthcare coverage
Relocation support
Wellness programs
+1
Research Engineer, Machine Learning
Research Engineer, Machine Learning

Mistral • Palo Alto (CA)

On-site
USD 180,000 - 230,000
Salary & equity
Healthcare coverage
401K matching
+6