AI/ML Engineer – GPU-Accelerated Real-Time Inference

Benz Technology

Fort Meade (MD)

On-site

USD 150,000 - 190,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Benz Technology is seeking an AI/ML inference engineer to join a team focused on high-performance GPU-accelerated applications in a cloud environment. You will develop AI/ML models, deploy cloud infrastructure with Terraform on AWS, and work with large streaming datasets for real-time processing.

The role requires proficiency in C++, Python, or Golang and experience building latency-sensitive systems, with a focus on scalable streaming analytics.

Qualifications

  • AI/ML development or integration experience.
  • Terraform on AWS.
  • Proficiency in C++, Python, and/or Golang.
  • GPU-enabled / high-performance application development.
  • Large-scale streaming data analytics.

Responsibilities

  • Develop and integrate AI/ML models into production inference pipelines.
  • Build and deploy cloud infrastructure using Terraform on AWS.
  • Write high-performance, GPU-enabled applications for real-time processing.
  • Design and maintain large-scale streaming data analytic systems.
  • Contribute across the stack using C++, Python, and/or Golang.
  • Optimize application performance for latency-sensitive workloads.

Skills

AI/ML development
GPU programming
Latency optimization
Streaming analytics

Tools

Terraform on AWS
Docker
Kubernetes

Job description

Benz Technology is seeking an AI/ML inference engineer to join a team focused on high-performance GPU-accelerated applications in a cloud environment. You will develop AI/ML models, deploy cloud infrastructure with Terraform on AWS, and work with large streaming datasets for real-time processing.

The role requires proficiency in C++, Python, or Golang and experience building latency-sensitive systems, with a focus on scalable streaming analytics.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer – AI/ML & High Performance
Software Engineer – AI/ML & High Performance

Benz Technology • Fort Meade (MD)

On-site
USD 150,000 - 190,000
AI Inference Engineer – High-Performance GPU Systems
AI Inference Engineer – High-Performance GPU Systems

Perplexity • California (MO)

On-site
USD 120,000 - 170,000
Senior ML Engineer: Real-Time Inference & Streaming (AWS)
Senior ML Engineer: Real-Time Inference & Streaming (AWS)

TWG Global AI • New York (NY)

On-site
USD 190,000 - 290,000
Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH
Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH

EngRadar • New York (NY)

On-site
USD 150,000 - 230,000
Equity grants
Medical plan
Vision plan
+5
Senior Real-Time ML Engineer (AWS, Low-Latency Inference)
Senior Real-Time ML Engineer (AWS, Low-Latency Inference)

TWG Global AI • Santa Monica (CA)

On-site
USD 190,000 - 290,000
Senior ML Engineer - Real-Time Inference & Streaming (AWS)
Senior ML Engineer - Real-Time Inference & Streaming (AWS)

TWG AI • New York (NY)

On-site
USD 190,000 - 290,000
Bonus
Medical benefits
Remote ML Infrastructure Engineer — GPU Clusters & AI Platform
Remote ML Infrastructure Engineer — GPU Clusters & AI Platform

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Real-Time ML Engineer – AWS, Low-Latency Inference
Real-Time ML Engineer – AWS, Low-Latency Inference

TWG AI • Santa Monica (CA)

On-site
USD 190,000 - 290,000
Senior DL Inference Engineer — GPU-Accelerated AI, Equity
Senior DL Inference Engineer — GPU-Accelerated AI, Equity

NVIDIA Gruppe • California (MO)

On-site
USD 152,000 - 288,000