AI Performance Engineer

Bright Vision Technologies

Cranberry Township (Butler County)

Remote

USD 100,000 - 150,000

Full time

10 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Bright Vision Technologies is seeking an AI Performance Engineer for a 100% remote U.S. role. You will profile and optimize end-to-end AI training and inference pipelines, focusing on throughput, latency, and cost reductions across data loading, compute, and memory.

The role requires 6+ years in performance engineering and expertise in Python/C++, distributed training, and GPU optimizations. Collaboration with ML and platform teams is essential to deploy scalable improvements.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
  • Six or more years of experience in performance engineering, ML systems, or HPC.
  • Strong proficiency in Python and C++.
  • Hands-on experience optimizing deep learning workloads on modern GPUs.
  • Deep understanding of distributed training and inference techniques.
  • Experience with profiling tools across CPU, GPU, and distributed systems.
  • Familiarity with model compression techniques and their accuracy implications.
  • Strong grasp of memory hierarchies, communication primitives, and parallelism strategies.
  • Excellent measurement, debugging, and analytical reasoning skills.
  • Strong communication and collaboration skills.

Responsibilities

  • Profile and optimize end-to-end AI training and inference pipelines for throughput, latency, and cost.
  • Identify and eliminate bottlenecks across data loading, model compute, communication, and memory.
  • Implement and tune quantization, sparsity, and pruning strategies to reduce model footprint and accelerate inference.
  • Optimize distributed training using tensor parallelism, pipeline parallelism, FSDP, and ZeRO-style sharding.
  • Tune attention implementations using FlashAttention, paged attention, and related techniques.
  • Implement KV cache optimization, continuous batching, and speculative decoding for LLM serving.
  • Drive compiler-level optimizations using Triton, XLA, TorchInductor, or TVM, working with the broader ML framework community to land improvements that translate into measurable end-to-end performance gains.
  • Optimize data pipelines, sharding strategies, and storage access patterns for high-throughput training.
  • Build and maintain rigorous benchmark suites and regression frameworks across workloads.
  • Collaborate with ML and platform engineering teams to embed best practices in standard pipelines.
  • Drive cost-efficiency improvements through model architecture, hardware selection, and scheduling strategies.
  • Evaluate new hardware and software offerings, and advise on adoption.
  • Document performance tuning playbooks and share findings broadly across engineering teams.
  • Stay current with AI systems research and translate advances into production improvements.

Skills

Python
C++
Performance engineering
ML systems
HPC
Profiling tools
Distributed training
GPU optimization
Memory hierarchies
Debugging
Collaboration

Education

Bachelor’s or Master’s degree in Computer Science or Computer Engineering

Tools

Triton
XLA
TorchInductor
TVM
CUTLASS
TensorRT-LLM
DeepSpeed

Job description

AI Performance Engineer – Remote

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.

Job Title: AI Performance Engineer

Location: 100% Remote (U.S.)

Position Type: Full-time, Direct W2

Salary Range: $100,000–$150,000 Annually

Experience Required: 6+ years

Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.

Key Responsibilities
  • Profile and optimize end-to-end AI training and inference pipelines for throughput, latency, and cost.
  • Identify and eliminate bottlenecks across data loading, model compute, communication, and memory.
  • Implement and tune quantization, sparsity, and pruning strategies to reduce model footprint and accelerate inference.
  • Optimize distributed training using tensor parallelism, pipeline parallelism, FSDP, and ZeRO-style sharding.
  • Tune attention implementations using FlashAttention, paged attention, and related techniques.
  • Implement KV cache optimization, continuous batching, and speculative decoding for LLM serving.
  • Drive compiler-level optimizations using Triton, XLA, TorchInductor, or TVM, working with the broader ML framework community to land improvements that translate into measurable end-to-end performance gains.
  • Optimize data pipelines, sharding strategies, and storage access patterns for high-throughput training.
  • Build and maintain rigorous benchmark suites and regression frameworks across workloads.
  • Collaborate with ML and platform engineering teams to embed best practices in standard pipelines.
  • Drive cost-efficiency improvements through model architecture, hardware selection, and scheduling strategies.
  • Evaluate new hardware and software offerings, and advise on adoption.
  • Document performance tuning playbooks and share findings broadly across engineering teams.
  • Stay current with AI systems research and translate advances into production improvements.
Required Qualifications
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
  • Six or more years of experience in performance engineering, ML systems, or HPC.
  • Strong proficiency in Python and C++.
  • Hands-on experience optimizing deep learning workloads on modern GPUs.
  • Deep understanding of distributed training and inference techniques.
  • Experience with profiling tools across CPU, GPU, and distributed systems.
  • Familiarity with model compression techniques and their accuracy implications.
  • Strong grasp of memory hierarchies, communication primitives, and parallelism strategies.
  • Excellent measurement, debugging, and analytical reasoning skills.
  • Strong communication and collaboration skills.
Preferred Qualifications
  • Experience optimizing LLM inference at production scale.
  • Contributions to vLLM, TensorRT-LLM, DeepSpeed, or similar projects.
  • Familiarity with custom kernel authoring in Triton or CUTLASS.
  • Experience with FinOps for AI workloads.
  • Publications or talks on AI systems performance.

Bright Vision Technologies is an Equal Opportunity Employer.

Equal Employment Opportunity (EEO) Statement

Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.

BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees’ ability to perform their job duties may result in disciplinary action up to and including termination of employment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Model Optimization Engineer
Model Optimization Engineer

Bright Vision Technologies • United States

Remote
USD 150,000 - 175,000
Data Platform Engineer
Data Platform Engineer

Bright Vision Technologies • Reston (VA)

On-site
USD 100,000 - 150,000
AI Platform Engineer
AI Platform Engineer

Bright Vision Technologies • United States

Remote
USD 130,000 - 180,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Bright Vision Technologies • Kirkland (WA)

Remote
USD 100,000 - 150,000
ML Platform Engineer
ML Platform Engineer

Bright Vision Technologies • Cranberry Township

Remote
USD 100,000 - 160,000
AI Pipeline Engineer
AI Pipeline Engineer

Bright Vision Technologies • Kirkland (WA)

Remote
USD 100,000 - 150,000
AI Innovation Engineer
AI Innovation Engineer

Bright Vision Technologies • United States

Remote
USD 135,000 - 210,000
Equal Opportunity Employer
Embedded AI Engineer
Embedded AI Engineer

Bright Vision Technologies • Kirkland (WA)

Remote
USD 100,000 - 150,000
Data Engineering Specialist – AI
Data Engineering Specialist – AI

Bright Vision Technologies • Columbus (OH), Dublin (OH)

Remote
USD 100,000 - 150,000
Applied AI Engineer
Applied AI Engineer

Bright Vision Technologies • Columbus (OH), Dublin (OH)

Remote
USD 130,000 - 180,000
Fully remote