AI/ML Infrastructure Engineer Software

Front Door Defense

San Francisco, Northern (CA, KY)

Hybrid

USD 140,000 - 210,000

Full time

5 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Restaurant d'entreprise
Indemnités de stage/alternance

Job summary

Zensors in San Francisco seeks an AI/ML Infrastructure Engineer to build and optimize scalable AI inference infrastructure for real-time video analytics.

You will evolve the engine powering our visual sensing platform, accelerate training and inference for CV models, and collaborate across research and platform teams to push throughput while reducing latency.

Applicants should have strong systems programming skills and a track record of delivering performant ML pipelines in production.

Qualifications

  • BS/MS or Ph.D. in Computer Science, Electrical Engineering, or a related discipline.
  • Strong programming skills in C/C++ and Python.
  • Experience with model optimization, quantization, and efficient deep learning techniques.
  • Deep understanding of GPU hardware performance, including memory/cache management.
  • Experience with profiling/benchmarking tools (Nsight) to validate performance.
  • Experience identifying and resolving bottlenecks in high-bandwidth video pipelines.
  • Strong communication and cross-functional collaboration skills.

Responsibilities

  • Optimizing Core ML Pipelines: identify bottlenecks and optimize for server and edge compute.
  • Cross-Stack Collaboration: work with research and platform teams to optimize inference infrastructure.
  • Model Acceleration: apply quantization, pruning, and layer fusion to CV models.
  • Building Efficient Operators: develop optimized ML operators across PyTorch/CUDA/TensorRT.
  • Resource Efficiency: reduce compute cost per video stream for scalability.
  • Data Management: support collection and labeling for ML training.

Skills

C/C++
Python
Model optimization
GPU performance
Profiling tools
Video processing bottlenecks
Communication

Education

CS/EE degree

Tools

CUDA
TensorRT
NVIDIA DeepStream
PyTorch

Job description

AI/ML Infrastructure Engineer
Build and optimize scalable AI inference infrastructure for real-time video analytics

Location: San Francisco, California

About The Role
Machine Learning Engineer In ML Runtime & Optimization

The AI Infrastructure team at Zensors builds the engine that powers our visual sensing platform. We provide the tools to automate the lifecycle of our AI workflow, including model development, evaluation, optimization, deployment, and monitoring across thousands of video streams.

As a Machine Learning Engineer in ML Runtime & Optimization, you will develop technologies to accelerate the training and inference of computer vision models that power smart spaces and cities.

Your responsibilities will include:

  • Optimizing Core ML Pipelines: Identifying key bottlenecks in our current video analytics pipeline and performing in-depth analysis to ensure the best possible performance on current server and edge compute architectures.
  • Cross-Stack Collaboration: Collaborating closely with AI research and platform engineering teams to optimize core parallel algorithms and influence the design of our next-generation inference infrastructure.
  • Model Acceleration: Applying advanced model optimization techniques—such as quantization (Int8/FP16), pruning, and layer fusion—to our Vision Transformers (ViTs) and CNNs to maximize throughput and minimize latency.
  • Building Efficient Operators: Working across the entire ML framework/compiler stack (e.g., PyTorch, CUDA, TensorRT, and NVIDIA DeepStream) to write custom optimized ML operator libraries.
  • Resource Efficiency: Reducing the compute cost per video stream to enable massive scalability of our SaaS product.
  • Data Management: Building, improving, maintaining, and operating systems to facilitate the collection, labeling, and use of visual data for ML training.
Requirements
  • BS/MS or Ph.D. in Computer Science, Electrical Engineering, or a related discipline.
  • Strong programming skills in C/C++ and Python.
  • Experience with model optimization, quantization, and efficient deep learning techniques (e.g., knowledge distillation, pruning).
  • Deep understanding of GPU hardware performance, including execution models, thread hierarchy, memory/cache management, and the cost/performance trade-offs of video processing.
  • Experience with profiling and benchmarking tools (e.g., Nsight Systems, Nsight Compute) to validate performance on complex architectures.
  • Experience identifying and resolving compute and data flow bottlenecks, particularly in high-bandwidth video processing pipelines.
  • Strong communication skills and the ability to work cross-functionally between research and infrastructure teams.
Preferred Qualifications
  • Familiarity with database systems (e.g., SQL, Neo4j).
  • Work in Computer Vision, Deep Learning, and Vision Transformers.
  • Experience with video processing frameworks such as NVIDIA DeepStream, DALI, or FFmpeg.
  • Familiarity with ML compilers (e.g., TVM, MLIR) or inference engines like TensorRT or ONNX Runtime.
  • Knowledge of distributed training systems or cloud-scale inference serving (e.g., Triton Inference Server).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/ML Inference & Vision Infrastructure Engineer
AI/ML Inference & Vision Infrastructure Engineer

Front Door Defense • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Restaurant d'entreprise
Indemnités de stage/alternance
Dev / ML Ops Engineer
Dev / ML Ops Engineer

Zensors Inc • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive base salary + equity options
Comprehensive health, dental, and vision benefits
Founding DevOps / Systems / ML Ops Engineer - Physical AI Startup
Founding DevOps / Systems / ML Ops Engineer - Physical AI Startup

Skyrocket Ventures • San Francisco (CA)

Hybrid
USD 170,000 - 200,000
Senior Software Engineer - ML Infrastructure
Senior Software Engineer - ML Infrastructure

Claryo • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Medical/Dental/Vision
401k with employer matching
Parental leave
+1
Principal ML Infrastructure Engineer (Relocation Available)
Principal ML Infrastructure Engineer (Relocation Available)

Franklin Fitch • Dallas (TX)

On-site
USD 100,000 - 140,000
Software Engineer, AI Infrastructure – LVM Inference & Evaluation
Software Engineer, AI Infrastructure – LVM Inference & Evaluation

Jobtailor • Redwood City (CA)

On-site
USD 180,000 - 240,000
AI/ML Engineer (Computer Vision)
AI/ML Engineer (Computer Vision)

Blue-Signal-Search • San Francisco (CA)

On-site
USD 120,000 - 170,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Ultra • New York (NY)

Hybrid
USD 180,000 - 240,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Research Engineer, Infrastructure, Inference
Research Engineer, Infrastructure, Inference

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1