Machine Learning Systems Research Engineer (Agent Post-training, Enterprise GenAI)

Scale AI

New York (NY)

On-site

USD 180,000 - 250,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Dental & Vision insurance
Mental health support
Paid time off
Learning stipend
Office events

Job summary

Scale AI in New York is seeking an ML System Research Engineer responsible for building and optimizing the next-generation agent training platform. You will support large-scale LLM training and inference, post-train state-of-the-art models, and collaborate with ML teams to accelerate research and development for enterprise AI workloads.

The role requires strong software engineering skills, experience with CUDA, PyTorch, transformers, and post-training methods such as RLHF/RLVR, PPO/GRPO, and a

Qualifications

  • Strong written and verbal communication skills to operate in a cross functional team environment.
  • Experience with multi-node LLM training and inference.
  • At least 1-3 years of LLM training in a production environment.
  • PhD or Masters in Computer Science or a related field.
  • Ability to demonstrate know-how on how to operate the architecture of the modern GPU cluster.
  • Strong software engineering skills, proficient in frameworks and tools such as CUDA, Pytorch, transformers, flash attention, etc
  • Experience with post-training methods like RLHF/RLVR and related algorithms like PPO/GRPO etc
  • Passionate about system optimization.

Responsibilities

  • Build, profile and optimize our training and inference framework.
  • Post-train state of the art models, developed both internally and from the community, to define stable post-training recipes for our enterprise engagements.
  • Collaborate with ML teams to accelerate their research and development, and enable them to develop the next generation of models and data curation.
  • Create a next-gen agent training algorithm for multi-agent/multi-tool rollouts

Skills

Communication
LLM training
Production experience
GPU cluster
CUDA
PyTorch
Transformers
RLHF/RLVR
PPO/GRPO
System optimization

Education

PhD or Masters in CS

Tools

CUDA
PyTorch
Transformers
Flash attention

Job description

  • The Enterprise ML Research Lab works on the front lines of this AI revolution
  • We are working on an arsenal of proprietary research and resources that serve all of our enterprise clients
  • As an ML Sys Research Engineer, you’ll work on building out the algorithms for our next-gen Agent RL training platform, support large scale training, and research and integrate state-of-the-art technologies to optimize our ML system
  • Your customer will be other MLREs and AAIs on the Enterprise AI team who are taking the training algorithms and applying them to client use-cases ranging from next-generation AI cybersecurity firewall LLMs to training foundation healthtech search models
  • Build, profile and optimize our training and inference framework
  • Post-train state of the art models, developed both internally and from the community, to define stable post-training recipes for our enterprise engagements
  • Collaborate with ML teams to accelerate their research and development, and enable them to develop the next generation of models and data curation.
  • Create a next-gen agent training algorithm for multi-agent/multi-tool rollouts
Benefits
  • Health & Wellbeing: Our holistic approach to supporting Scaliens includes comprehensive health coverage, dental and vision insurance, mental healthcare services, and more. PTO policies and accommodating schedules ensure you’ll get time off when you need it to relax and recharge. Note that our offerings may vary by region as we strive to respond to the unique needs of Scaliens around the globe.
  • Personal & Career Growth: Continuously learn and grow through annual learning & development stipend, attending leadership breakfasts, manager training, speaker series, and joining an ERG.
  • Building Scale Community: We welcome guests to our offices, and you can expect to see Scalien families and friends around. Join local happy hours, and accept invites to game nights, book clubs, and many other employee-led community events.
  • Parental Support: Balancing work and family is essential, and Scale understands the importance of having adequate leave policies in place to promote a healthy home and work life.

If you are excited about shaping the future of the modern AI movement, we would love to hear from you!

  • Strong written and verbal communication skills to operate in a cross functional team environment
  • Experience with multi-node LLM training and inference
  • At least 1-3 years of LLM training in a production environment
  • PhD or Masters in Computer Science or a related field
  • Ability to demonstrate know-how on how to operate the architecture of the modern GPU cluster
  • Strong software engineering skills, proficient in frameworks and tools such as CUDA, Pytorch, transformers, flash attention, etc
  • Experience with post-training methods like RLHF/RLVR and related algorithms like PPO/GRPO etc
  • Passionate about system optimization
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Machine Learning Research Engineer (Agent Post-training, Enterprise GenAI)
Staff Machine Learning Research Engineer (Agent Post-training, Enterprise GenAI)

Scale AI • New York (NY)

On-site
USD 180,000 - 240,000
Health & wellbeing
Career development stipend
Community events
+1
Machine Learning Research Engineer (Agents, Enterprise GenAI)
Machine Learning Research Engineer (Agents, Enterprise GenAI)

Scale AI • New York (NY)

On-site
USD 150,000 - 210,000
Health & wellbeing benefits
Career growth stipend
Community events
+1
Machine Learning Research Engineer (Agent Data Foundation, Enterprise GenAI)
Machine Learning Research Engineer (Agent Data Foundation, Enterprise GenAI)

Scale AI • New York (NY)

On-site
USD 180,000 - 280,000
Health & Wellbeing
Personal & Career Growth
Building Scale Community
+1
Senior/Staff Machine Learning Research Engineer (General Agents, Enterprise GenAI)
Senior/Staff Machine Learning Research Engineer (General Agents, Enterprise GenAI)

Scale AI • New York (NY)

On-site
USD 180,000 - 260,000
Health & vision insurance
Dental insurance
Mental healthcare
+1
Machine Learning Research Engineer (ML Systems)
Machine Learning Research Engineer (ML Systems)

Scale AI • New York (NY)

On-site
USD 150,000 - 200,000
Health coverage
Dental & vision insurance
Learning stipend
+1
Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI
Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI

United States Digital Space LLC • New York (NY), San Francisco (CA)

On-site
USD 265,000 - 331,000
Machine Learning Research Scientist (Post-Training)
Machine Learning Research Scientist (Post-Training)

Scale AI • New York (NY)

On-site
USD 150,000 - 200,000
Health & wellbeing benefits
Learning & development stipend
Community & ERG events
+1
Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI New York, NY[...]
Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI New York, NY[...]

Scale AI, Inc. • New York (NY)

On-site
USD 250,000 - 350,000
Comprehensive health, dental, and vision coverage
Retirement benefits
Learning and development stipend
+2
Senior / Staff Machine Learning Research Scientist (Agents)
Senior / Staff Machine Learning Research Scientist (Agents)

Scale AI • New York (NY)

On-site
USD 150,000 - 210,000
Health coverage
Growth stipend
Community events
+1
Tech Lead Manager (MLRE, ML Systems)
Tech Lead Manager (MLRE, ML Systems)

Scale AI • New York (NY)

On-site
USD 160,000 - 210,000
Health coverage
Dental and vision insurance
Mental healthcare services
+2