Senior Model Optimization Engineer - Remote AI Performance

Bright Vision Technologies

United States

Remote

USD 150,000 - 175,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Bright Vision Technologies is seeking a Model Optimization Engineer to maximize throughput, minimize latency, and reduce cost across training and inference for large neural network systems. You will optimize from kernel-level code to distributed systems, working with GPU architectures, model parallelism, memory management, and compiler-level techniques.

The role requires 6+ years in performance engineering, deep learning workloads on GPUs, and strong Python/C++.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
  • Six+ years of experience in performance engineering, ML systems, or HPC.
  • Strong proficiency in Python and C++.
  • Hands-on experience optimizing deep learning workloads on modern GPUs.
  • Deep understanding of distributed training and inference techniques.
  • Experience with profiling tools across CPU, GPU, and distributed systems.
  • Familiarity with model compression techniques and their accuracy implications.
  • Strong grasp of memory hierarchies, communication primitives, and parallelism strategies.
  • Excellent measurement, debugging, and analytical reasoning skills.
  • Strong communication and collaboration skills.

Responsibilities

  • Optimize throughput, latency, and cost across training and inference workloads for large neural networks.
  • Span the full stack from low-level kernel optimization to distributed system tuning.
  • Work with cross-functional teams to translate requirements into engineered solutions.
  • Mentor junior engineers and contribute to design and code reviews.

Skills

Python
C++
Performance engineering
GPU optimization
Distributed systems
Profiling tools

Education

Bachelor’s or Master’s degree in Computer Science or Computer Engineering

Tools

Triton
CUTLASS
TensorRT

Job description

Bright Vision Technologies is seeking a Model Optimization Engineer to maximize throughput, minimize latency, and reduce cost across training and inference for large neural network systems. You will optimize from kernel-level code to distributed systems, working with GPU architectures, model parallelism, memory management, and compiler-level techniques.

The role requires 6+ years in performance engineering, deep learning workloads on GPUs, and strong Python/C++.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Performance Engineer — Remote
Senior AI Performance Engineer — Remote

Socket.dev • Ann Arbor (MI)

On-site
USD 75,000 - 100,000
Model Optimization Engineer
Model Optimization Engineer

Bright Vision Technologies • United States

Remote
USD 150,000 - 175,000
Remote AI Performance Engineer — Optimize ML Pipelines
Remote AI Performance Engineer — Optimize ML Pipelines

Bright Vision Technologies • Cranberry Township

Remote
USD 100,000 - 150,000
Senior Model Serving Engineer – Remote AI Infra
Senior Model Serving Engineer – Remote AI Infra

United States Digital Space LLC • United States

Remote
USD 74,000 - 98,000
Senior ML Model Serving Engineer (Remote)
Senior ML Model Serving Engineer (Remote)

NEPSE Trading • Northern (KY)

Hybrid
USD 74,000 - 98,000
Senior AI Training Performance Engineer (GPU & Scale)
Senior AI Training Performance Engineer (GPU & Scale)

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Remote AI Performance Engineer - Scale ML Inference
Remote AI Performance Engineer - Scale ML Inference

United States Digital Space LLC • United States

Remote
USD 75,000 - 100,000
Member of Technical Staff (Performance Optimization)
Member of Technical Staff (Performance Optimization)

Fireworks AI • United States

On-site
USD 150,000 - 260,000
Remote Model Serving Engineer: Scalable AI Inference
Remote Model Serving Engineer: Scalable AI Inference

Socket.dev • Ann Arbor (MI)

On-site
USD 74,000 - 98,000
Remote CUDA GPU Software Engineer — HPC & AI Performance
Remote CUDA GPU Software Engineer — HPC & AI Performance

RiseMe • Monroeville

On-site
USD 100,000 - 175,000