Machine Learning Infrastructure Engineer, GenAI Technology

Point72

New York (NY)

On-site

USD 180,000 - 300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Fully-paid health care benefits
Parental and family leave policies
Volunteer opportunities
Support for employee-led affinityroups
Mental and physical wellness programs
Tuition assistance
401(k) with employer match

Job summary

Point72’s Technology team is building scalable GenAI infrastructure to accelerate model development and deployment. You will design high-performance compute environments, and partner with ML researchers to deliver end-to-end ML workflows in cloud and on‑prem settings.

This role emphasizes CI/CD for models, observability, cost management, and strong collaboration. A degree in a technical field and 3–7 years of infra experience are expected.

Qualifications

  • Bachelor's or master's degree in CS/EE or related field.
  • 3–7 years of experience building scalable compute or ML infra.
  • Experience with Kubernetes and cloud platforms (AWS/GCP/Azure).
  • Proficiency in Python and systems-level programming (Go/C++/Rust).
  • Hands-on with MLops tools: MLflow, Ray, Airflow, Kubeflow, Terraform.

Responsibilities

  • Design and implement infrastructure for high-performance generative AI and ML workloads.
  • Operate distributed systems for training, tuning, inference, and data preprocessing pipelines.
  • Collaborate with ML researchers to optimize compute, throughput, and latency.
  • Develop and automate deployment, orchestration, and CI/CD for models and data workflows.
  • Implement observability and cost-management for GPU/accelerator environments.
  • Benchmark emerging hardware/software to improve scalability and throughput.
  • Drive security, compliance, and runbooks for GenAI infrastructure.
  • Troubleshoot and optimize GPU/CPU compute stacks for reliability.
  • Document architecture and mentor engineers to expand team capability.

Skills

Python
Systems programming
Distributed thinking

Education

Bachelor's or Master's in CS/EE or related

Tools

MLflow
Ray
Airflow
Kubeflow
Terraform
Kubernetes

Job description

A Career with Point72's Technology Team

As Point72 reimagines the future of investing, our Technology team is constantly evolving our firm’s IT infrastructure and engineering capabilities, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts who experiment and work to discover new ways to harness open-source solutions, modern cloud architectures, and sophisticated Artificial Intelligence (AI) solutions, while embracing enterprise agile methodologies. Our commitment to building and innovating in the AI space provides the framework intended to drive smarter decision making and enhance how we build and operate our platforms and applications.

As a member of Point72’s Technology team, we encourage and support your professional development from day one—helping you advance your technical skills, contribute innovative ideas, and satisfy your own intellectual curiosity—all while delivering real business impact for our multi-billion-dollar global business.

WHAT YOU'LL DO
  • Design and implement high-performance infrastructure to support large-scale generative AI and machine learning workloads, enabling faster model iteration and real business impact
  • Design and operate distributed systems for model training, hyperparameter tuning, inference, and data preprocessing pipelines to deliver reliable end-to-end machine learning (ML) workflows
  • Collaborate with ML researchers and engineers to produce models, optimizing compute utilization, training throughput, and inference latency
  • Develop and automate deployment, orchestration, and CI/CD pipelines for models and data workflows using container orchestration and infrastructure-as-code (IaC)
  • Implement observability, monitoring, and cost-management strategies for GPU and accelerator compute environments to maintain predictable performance and spend
  • Evaluate, integrate, and benchmark emerging hardware and software technologies across cloud and on-prem environments to improve scalability and throughput
  • Drive security, compliance, and operational runbooks for GenAI infrastructure including access controls, secrets management, and incident response procedures
  • Troubleshoot, profile, and optimize performance across GPU and CPU compute stacks to remove bottlenecks and increase reliability
  • Document architecture, operational practices, and mentor engineers to expand team capability and accelerate adoption of production-ready GenAI infrastructure
WHAT'S REQUIRED
  • Bachelor's or master's degree in computer science, electrical engineering, or a related technical field
  • 3–7 years of experience building and maintaining scalable compute or machine learning infrastructure systems
  • Deep understanding of distributed systems, container orchestration (Kubernetes), and public cloud platforms such as AWS, Google Cloud Platform, or Azure
  • Hands-on experience with machine learning operations and infrastructure tools such as MLflow, Ray, Airflow, Kubeflow, and Terraform
  • Strong understanding of reinforcement learning concepts and their infrastructure implications
  • Proficiency in Python and systems-level programming in one or more languages such as Go, C++, or Rust
  • Strong debugging, performance profiling, and optimization skills across GPU and CPU compute stacks
  • Experience implementing monitoring, observability, and cost-optimization for GPU/accelerator-based compute environments
  • Excellent collaboration and communication skills with a systems-thinking mindset
  • Commitment to the highest ethical standards
WE TAKE CARE OF OUR PEOPLE

We invest in our people, their careers, their health, and their well-being. When you work here, we provide:

  • Fully-paid health care benefits
  • Generous parental and family leave policies
  • Volunteer opportunities
  • Support for employee-led affinity groups representing women, people of color and the LGBT+ community
  • Mental and physical wellness programs
  • Tuition assistance
  • A 401(k) savings program with an employer match and more
ABOUT POINT72

Point72 is a leading global alternative investment firm led by Steven A. Cohen. Building on more than 30 years of investing experience, Point72 seeks to deliver superior returns for its investors through fundamental and systematic investing strategies across asset classes and geographies. We aim to attract and retain the industry's brightest talent by cultivating an investor-led culture and committing to our people's long‑term growth. For more information, visit https://point72.com/.

The annual base salary range for this role is $180,000-$300,000 (USD) , which does not include discretionary bonus compensation or our comprehensive benefits package. Actual compensation offered to the successful candidate may vary from posted hiring range based upon geographic location, work experience, education, and/or skill level, among other things.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Infrastructure Engineer, Technology
Machine Learning Infrastructure Engineer, Technology

Point72 • New York (NY)

On-site
USD 185,000 - 300,000
Fully paid health care
Parental and family leave
Wellness programs
+1
Machine Learning Engineer, Knowledge Graph Intelligence
Machine Learning Engineer, Knowledge Graph Intelligence

Linuxconfig • United States

On-site
USD 175,000 - 250,000
Fully-paid health care benefits
Generous parental and family leave policies
Tuition assistance
+1
GenAI Security Engineer
GenAI Security Engineer

Point72 Asset Management, L.P • New York (NY)

On-site
USD 225,000 - 300,000
Fully-paid health care benefits
Generous parental and family leave
Mental and physical wellness programs
+2
AI Engineer, L/S Equity
AI Engineer, L/S Equity

NLP PEOPLE • San Francisco (CA)

On-site
USD 150,000 - 200,000
Fully-paid health care benefits
Generous parental and family leave policies
Tuition assistance
+1
Machine Learning Engineer, Knowledge Graph Intelligence
Machine Learning Engineer, Knowledge Graph Intelligence

Point72 • New York (NY)

On-site
USD 175,000 - 250,000
Software Engineer, EPIC Technology Services
Software Engineer, EPIC Technology Services

Point72 • New York (NY)

On-site
USD 200,000
Fully-paid health care benefits
Generous parental leave policies
Mental wellness programs
+2
Software Engineer - Data
Software Engineer - Data

Point72 Asset Management, L.P • New York (NY)

On-site
USD 200,000 - 300,000
Fully-paid health care benefits
Generous parental and family leave policies
Tuition assistance
+1
Software Engineer - Data
Software Engineer - Data

Socket.dev • New York (NY)

On-site
USD 200,000 - 300,000
Fully-paid health care benefits
Generous parental and family leave policies
Volunteer opportunities
+4
GenAI ML Infra Engineer
GenAI ML Infra Engineer

Point72 • New York (NY)

On-site
USD 185,000 - 300,000
Fully paid health care
Parental and family leave
Wellness programs
+1
Business Manager, Technology
Business Manager, Technology

Point72 • New York (NY)

On-site
USD 200,000 - 275,000
Fully-paid health care benefits
Parental and family leave policies
Volunteer opportunities
+1