Senior Research Engineer - ML Systems

PERMUTE

Chicago (IL)

On-site

USD 150,000 - 210,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Permute is seeking a Senior Research Engineer to productionize, optimize, and extend model systems that power AI reasoning over structured data. The role centers on building reliable production systems from research ideas, with profiling, testing, and failure handling as core priorities.

The ideal candidate can implement research, diagnose performance, write clean production code, and make sound architectural decisions in a fast-moving startup environment.

Qualifications

  • Strong background in machine learning research and ML systems.
  • Experience building and training models with PyTorch.
  • Strong foundations in algorithms, statistics, optimization, and experimental design.
  • Strong software engineering and system architecture skills.
  • 5+ years building ML or performance-sensitive software systems.

Responsibilities

  • Productionize and optimize learned evidence architecture for structured data.
  • Improve training and inference performance (throughput, latency, memory, reliability, cost).
  • Port and optimize model training and inference workloads from CPU to GPU.
  • Build production systems for training, evaluation, deployment, and inference.
  • Develop tooling for experimentation, reproducibility, monitoring, observability.
  • Write clean Python and PyTorch code integrated with Permute's platform.
  • Design and evaluate new heads, layers, objectives, and fine-tuning methods.
  • Explore transformer-based architectures and reinforcement learning.
  • Collaborate with engineering and product teams to deliver model capabilities for production features.

Skills

ML research
ML systems
Software engineering
System architecture
Performance optimization

Education

Bachelor's or higher in Mathematics/Physics/CS

Tools

PyTorch

Job description

Senior Research Engineer – ML Systems
Employment Type: Full-time
Company: Permute (www.permute.ai)

Overview

Permute is seeking a Senior Research Engineer to productionize, optimize, and extend the model systems that power AI reasoning over structured data. This role is for builders who can move from research ideas to reliable production systems, including the profiling, testing, and failure handling that prototypes often skip.

We care as much about how you think and build as we do about your background. The ideal candidate can implement research, diagnose model and systems performance, write clean production code, and make sound architectural decisions in a fast-moving startup environment.

Responsibilities
  • Productionize and optimize our existing learned evidence architecture for structured data
  • Improve training and inference performance, including throughput, latency, memory use, reliability, and cost
  • Port and optimize model training and inference workloads from CPU to GPU
  • Build production systems supporting model training, evaluation, deployment, and inference
  • Develop tooling for experimentation, reproducibility, monitoring, and observability
  • Write clean, maintainable Python and PyTorch systems that integrate with Permute's broader platform
  • Design and evaluate new heads, layers, objectives, and fine-tuning methods
  • Explore new model variants, including transformer-based architectures and reinforcement learning
  • Collaborate with engineering and product teams to deliver model capabilities that power production AI features
Required Qualifications
  • Strong background in machine learning research and ML systems
  • Experience building and training models with PyTorch
  • Strong foundation in algorithms, statistics, optimization, and experimental design
  • Strong software engineering and system architecture skills
  • 5+ years building ML or performance-sensitive software systems
Preferred Background
  • Degree in Mathematics, Physics, Computer Science, or a related technical field

Experience with:

  • End-to-end production ML systems
  • Model training, MLOps, evaluation, and deployment
  • Performance engineering, including CUDA, Triton, quantization, or model compilation
  • Transformers, fine-tuning, post-training, or reinforcement learning
  • Meaningful contributions to open-source ML frameworks or model implementations
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Systems Engineer: Productionize & Optimize AI
Senior ML Systems Engineer: Productionize & Optimize AI

PERMUTE • Chicago (IL)

On-site
USD 150,000 - 210,000
Senior Data Engineer
Senior Data Engineer

Permute • Chicago (IL)

On-site
USD 100,000 - 130,000
Senior Software Engineer
Senior Software Engineer

Permute • Chicago (IL)

On-site
USD 120,000 - 150,000
Member of Technical Staff, MLSys
Member of Technical Staff, MLSys

Bake AI • San Mateo (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Member of the Technical Staff - Systems ML Engineer
Member of the Technical Staff - Systems ML Engineer

Breakout Ventures • Cambridge (MA)

On-site
USD 180,000 - 270,000
Equity
Lunch subsidy
Health insurance
+1
Model Systems Engineer
Model Systems Engineer

Mind Robotics • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Member of the Technical Staff - Systems ML Engineer
Member of the Technical Staff - Systems ML Engineer

Transfyr Bio • Cambridge (MA)

On-site
USD 150,000 - 230,000
Low-cost health insurance
HSA
401K with matching
+1
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • San Francisco (CA)

On-site
USD 140,000 - 210,000
Lead Engineer, Machine Learning
Lead Engineer, Machine Learning

Salt Digital Recruitment • United States

On-site
USD 180,000 - 260,000
ML Engineer: Production Inference & Deployment
ML Engineer: Production Inference & Deployment

HiringCafe • Cupertino (CA)

On-site
USD 250,000 - 310,000
Generous health, dental, and vision coverage
Paid parental leave
Relocation support