High-Performance ML Inference Engineer for Markets

Fintal Partners

New York (NY)

On-site

USD 140,000 - 220,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Fintal Partners in New York seeks a Machine Learning Engineer focused on inference to join a technically driven AI research team building and deploying large-scale models for financial markets. You will design high-performance inference systems, optimize GPU kernels with CUDA, Triton and related tools, and collaborate with researchers to deliver low-latency, reliable model serving for live trading.

This role blends ML, systems engineering and high-performance computing in a fast-paced, global

Qualifications

  • 2+ years of professional experience building deep learning or machine learning systems.
  • Strong software engineering and systems fundamentals.
  • Experience building deep learning systems in computationally intensive domains.
  • Ability to translate techniques and approaches across different machine learning domains.

Responsibilities

  • Design and optimize high-performance inference systems for large-scale deep learning models.
  • Develop and optimize GPU kernels using CUDA, Triton, Pallas, CuTe DSL, and related technologies
  • Improve lower-level performance across PyTorch, JAX, XLA, and CUDA Graphs
  • Explore and develop inference solutions across GPUs, ASICs, and FPGAs
  • Optimize data streaming and model-serving infrastructure for demanding real-time environments
  • Work closely with ML researchers to co-design efficient inference architectures
  • Improve latency, throughput, hardware utilization, and overall inference efficiency
  • Help shape the team’s broader machine learning and systems research agenda

Job description

Fintal Partners in New York seeks a Machine Learning Engineer focused on inference to join a technically driven AI research team building and deploying large-scale models for financial markets. You will design high-performance inference systems, optimize GPU kernels with CUDA, Triton and related tools, and collaborate with researchers to deliver low-latency, reliable model serving for live trading.

This role blends ML, systems engineering and high-performance computing in a fast-paced, global

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning Engineer - Inference
Machine Learning Engineer - Inference

Fintal Partners • New York (NY)

On-site
USD 140,000 - 220,000
Microsecond ML Inference Architect
Microsecond ML Inference Architect

Long Ridge Partners • New York (NY)

On-site
USD 600,000 - 1,500,000
Generous PTO
Hybrid work options
Free breakfast, lunch, snacks
Low-Latency ML Inference Engineer
Low-Latency ML Inference Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
ML Hardware Engineer: Custom Hardware Inference
ML Hardware Engineer: Custom Hardware Inference

Fintal Partners • New York (NY)

On-site
USD 140,000 - 190,000
Quant ML Researcher: Build Market-Predictive Models
Quant ML Researcher: Build Market-Predictive Models

Fintal Partners • New York (NY)

On-site
USD 150,000 - 210,000
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Long Ridge Partners • New York (NY)

On-site
USD 600,000 - 1,500,000
Generous PTO
Hybrid work options
Free breakfast, lunch, snacks
Machine Learning Performance Engineer (Inference)
Machine Learning Performance Engineer (Inference)

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Quant ML Researcher: Market Alpha & Modeling
Quant ML Researcher: Market Alpha & Modeling

Fintal Partners • New York (NY)

On-site
USD 180,000 - 320,000
AI Research Engineer Intern — Scalable Finance ML
AI Research Engineer Intern — Scalable Finance ML

Trading Interview • New York (NY)

On-site
USD 255,000 - 345,000
Kernel Engineer for High-Performance ML Compute (CUDA/Triton)
Kernel Engineer for High-Performance ML Compute (CUDA/Triton)

Inception • San Francisco (CA)

On-site
USD 180,000 - 260,000