Machine Learning Engineer - Inference

Fintal Partners

New York (NY)

On-site

USD 140,000 - 220,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Fintal Partners in New York seeks a Machine Learning Engineer focused on inference to join a technically driven AI research team building and deploying large-scale models for financial markets. You will design high-performance inference systems, optimize GPU kernels with CUDA, Triton and related tools, and collaborate with researchers to deliver low-latency, reliable model serving for live trading.

This role blends ML, systems engineering and high-performance computing in a fast-paced, global

Qualifications

  • 2+ years of professional experience building deep learning or machine learning systems.
  • Strong software engineering and systems fundamentals.
  • Experience building deep learning systems in computationally intensive domains.
  • Ability to translate techniques and approaches across different machine learning domains.

Responsibilities

  • Design and optimize high-performance inference systems for large-scale deep learning models.
  • Develop and optimize GPU kernels using CUDA, Triton, Pallas, CuTe DSL, and related technologies
  • Improve lower-level performance across PyTorch, JAX, XLA, and CUDA Graphs
  • Explore and develop inference solutions across GPUs, ASICs, and FPGAs
  • Optimize data streaming and model-serving infrastructure for demanding real-time environments
  • Work closely with ML researchers to co-design efficient inference architectures
  • Improve latency, throughput, hardware utilization, and overall inference efficiency
  • Help shape the team’s broader machine learning and systems research agenda

Job description

A leading quantitative trading firm is seeking a Machine Learning Engineer specializing in inference to join a highly technical AI research team developing and deploying large-scale machine learning models for financial markets.

The team builds powerful foundation models for markets, trained on vast quantities of market and alternative data to predict future market behavior. These models are deployed directly into live trading environments, making inference speed, efficiency, and reliability critical to the business.

As a Machine Learning Engineer, you will work at the intersection of machine learning, high-performance computing, and systems engineering, with a broad mandate to improve large-scale model inference.

What You’ll Work On

  • Design and optimize high-performance inference systems for large-scale deep learning models
  • Develop and optimize GPU kernels using CUDA, Triton, Pallas, CuTe DSL, and related technologies
  • Improve lower-level performance across PyTorch, JAX, XLA, and CUDA Graphs
  • Explore and develop inference solutions across GPUs, ASICs, and FPGAs
  • Optimize data streaming and model-serving infrastructure for demanding real-time environments
  • Work closely with ML researchers to co-design efficient inference architectures
  • Improve latency, throughput, hardware utilization, and overall inference efficiency
  • Help shape the team’s broader machine learning and systems research agenda

The inference environment spans multiple platforms deployed globally and supports a variety of model architectures and trading strategies. The work is highly performance-sensitive, technically challenging, and has a direct impact on live trading performance.

Qualifications

  • 2+ years of professional experience building deep learning or machine learning systems
  • Strong software engineering and systems fundamentals
  • Experience building deep learning systems in computationally intensive domains such as robotics, recommendation systems, biology, chemistry, physics, audio, video, or similar areas
  • Ability to translate techniques and approaches across different machine learning domains

Plus experience with one or more of the following:

  • Lower-level PyTorch, JAX, or XLA development
  • CUDA Graphs
  • GPU performance optimization
  • FPGA or ASIC development
  • High-performance ML inference systems

Nice to Have

  • Experience with large-scale or low-latency model inference
  • LLM or foundation model experience
  • Experience optimizing GPU kernels or distributed ML workloads

Prior finance or trading experience is not required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning Performance Engineer (Inference)
Machine Learning Performance Engineer (Inference)

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Long Ridge Partners • New York (NY)

On-site
USD 600,000 - 1,500,000
Generous PTO
Hybrid work options
Free breakfast, lunch, snacks
High-Performance ML Inference Engineer for Markets
High-Performance ML Inference Engineer for Markets

Fintal Partners • New York (NY)

On-site
USD 140,000 - 220,000
Machine Learning Performance Engineer - Quant Research & Trading
Machine Learning Performance Engineer - Quant Research & Trading

Acquire Me • United States

On-site
USD 200,000 - 350,000
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Fintal Partners • Chicago (IL)

On-site
USD 150,000 - 230,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Rapidtrade • New York (NY)

On-site
USD 180,000 - 240,000
Machine Learning Engineer, Inference Infrastructure
Machine Learning Engineer, Inference Infrastructure

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Equity
Medical benefits (full coverage)
PTO & Hybrid work policy
+1
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
Inference Engineer
Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 220,000