Machine Learning Research Engineer

Career Techniques

New York (NY)

Hybrid

USD 200,000 - 300,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Career Techniques in New York is seeking an experienced ML researcher to expand our distributed ML stack, benchmark workloads, and stress-test both software and hardware across the platform. You will build high-level abstractions, integrate tooling with data frameworks, and enable rapid experimentation using Ray.

The role requires hands-on ML experience with PyTorch/TensorFlow, scalable ML deployment, and familiarity with AI agent workflows.

Qualifications

  • Strong software engineering foundation in Python and API design.
  • Hands-on experience with modern frameworks (PyTorch, TensorFlow) and large-scale deployment.
  • Experience scaling ML workloads across GPUs and multi-node clusters.
  • Familiarity with AI agent workflows and LLM tooling.
  • Ability to identify bottlenecks across hardware and software (memory, GPU utilization, data pipelines).

Responsibilities

  • Platform Validation & Infrastructure Benchmarking: test training and inference environments.
  • Streamline Rapid Prototyping for ML Research: build abstractions and integrate core ML tooling.
  • Agentic Workflows for ML Research: leverage AI agents to generate experiments and stress-test clusters.
  • Research Platform Feedback & Insights Sharing: publish empirical findings and advise engineering teams.

Skills

Python
APIs
PyTorch
TensorFlow
Distributed Compute
Ray
Dask

Tools

Ray
Dask
PyTorch Distributed

Job description

You will focus on expanding our ML research platform to benchmark, rapidly prototype, and stress-test both software and hardware layers across our entire distributed ML stack. By leveraging AI agents and auto-research capabilities, you will push our systems to their limits, identify bottlenecks, and create a frictionless environment to test novel machine learning models on realistic, large-scale data.

Responsibilities
  • Platform Validation & Infrastructure Benchmarking:
    • Serve as the primary feedback loop for the entire ML stack.
    • Actively run complex models through our full ML pipeline to comprehensively test both the training and inference environments.
    • Validate the central infrastructure in practice, seeing exactly how new research ideas fare and identifying system bottlenecks before broader rollout to research teams.
  • Streamline Rapid Prototyping for ML Research:
    • Build high-level abstractions that allow users to bypass setup friction.
    • Integrate our core ML tooling directly with our underlying simulation and data frameworks, providing a unified entry point to access our full tech stack.
    • Enable rapid iteration on real-world data and seamless distributed training via Ray.
  • Agentic Workflows for ML Research:
    • Leverage AI agents and auto-research workflows to autonomously generate experiments, stress-test our distributed clusters, and provide data-driven, actionable feedback on what infrastructure needs to be optimized or built next.
  • Research Platform Feedback & Insights Sharing:
    • Act as the critical bridge between infrastructure builders and ML researchers.
    • Be the first to exhaustively test new models and push the platform's limits.
    • Document and publish empirical findings on system capabilities and hardware performance.
    • Take your validated insights to assist engineering teams with platform improvements and advise researchers on how to best leverage the stack.
Qualifications
  • Strong Software Engineering Foundation:
    • Deep proficiency in Python and software design principles.
    • Ability to build clean, scalable APIs and abstractions that other developers and researchers are enthusiastic about using.
  • Applied Machine Learning:
    • Hands-on experience with modern frameworks (PyTorch, TensorFlow, etc.)
    • Strong practical understanding of how to train, evaluate, and deploy models at scale.
  • Distributed Compute:
    • Experience scaling ML workloads across GPUs and multi-node clusters using frameworks like Ray, Dask, or PyTorch Distributed.
  • AI Agent Workflows:
    • Familiarity with LLM tooling, agentic frameworks, and using AI to automate coding, research, or testing tasks.
  • System Profiling & Optimization:
    • Ability to debug and identify bottlenecks across hardware and software layers (e.g., memory limits, GPU utilization, data pipeline latency).

Comp: $200-300K + Bonus

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer
Senior Machine Learning Engineer

Sierracorp • San Francisco (CA)

On-site
USD 150,000 - 200,000
Machine Learning Lead Engineer
Machine Learning Lead Engineer

Cox • Atlanta (GA)

On-site
USD 134,900 - 224,900
Flexible vacation policy
Paid wellness hours
Paid holidays
Applied ML Systems Engineer – Finance
Applied ML Systems Engineer – Finance

Park Lane Recruitment • New York (NY)

On-site
USD 250,000 - 350,000
401(k) matching
Medical coverage
Wellness reimbursement
+2
Software Engineer, Core Machine Learning
Software Engineer, Core Machine Learning

Meta • Sunnyvale (CA)

On-site
USD 184,000 - 257,000
Bonus
Equity
Benefits
Senior ML Engineer
Senior ML Engineer

Next Ventures • New York (NY)

On-site
USD 130,000 - 160,000
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 300,000
Cash and equity
Ownership & autonomy
Visa sponsorship
+5
Machine Learning Systems Engineer
Machine Learning Systems Engineer

Motional AD Inc. • Boston (MA)

Hybrid
USD 144,000 - 192,000
Medical, dental, vision insurance
401k with company match
Health saving accounts
+2
Machine Learning Research Engineer
Machine Learning Research Engineer

Iceberg • New York (NY)

On-site
USD 200,000 - 300,000
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3
Software Engineer, Research Acceleration
Software Engineer, Research Acceleration

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Unlimited PTO
Paid parental leave
Generous health, dental, and vision benefits