Remote AI Inference Engineer: Edge & Performance

ITRex Group

United States

Remote

USD 120,000 - 170,000

Full time

10 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Remote flexibility
Competitive salary + medical benefits
Learning opportunities

Job summary

ITRex is seeking a strong C++ Engineer to deploy and optimize modern AI models for production, focusing on edge devices and high-performance inference.

You will integrate, profile, and optimize AI inference pipelines using llama.cpp, ggml, and related frameworks, collaborating with researchers to move models from research to production.

Remote-friendly role across the US; you’ll contribute to edge deployments, performance tuning, and memory management in distributed systems.

Qualifications

  • 4+ years of professional experience with Modern C++ (C++17/20).
  • Strong memory management, multithreading, profiling and perf optimization skills.
  • Experience debugging memory leaks, fragmentation, OOM and concurrency issues.
  • Experience Linux development environments and production deployments.

Responsibilities

  • Deploy ML models to edge devices and optimize inference pipelines.
  • Collaborate with researchers to move models from research to production.
  • Integrate AI features into existing products and evolve AI capabilities in production environment.

Skills

Modern C++
Multithreading
Profiling
Linux
Memory management
OOP/low-level debugging

Tools

llama.cpp
ggml
ONNX Runtime
TensorRT
OpenVINO
MLC LLM
ExecuTorch
TVM

Job description

ITRex is seeking a strong C++ Engineer to deploy and optimize modern AI models for production, focusing on edge devices and high-performance inference.

You will integrate, profile, and optimize AI inference pipelines using llama.cpp, ggml, and related frameworks, collaborating with researchers to move models from research to production.

Remote-friendly role across the US; you’ll contribute to edge deployments, performance tuning, and memory management in distributed systems.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Edge AI Inference Engineer (C++/Llama ggml)
Edge AI Inference Engineer (C++/Llama ggml)

ITRex Group • United States

Remote
USD 120,000 - 180,000
Remote flexibility
Medical benefits
Learning budget
Edge AI Inference Engineer — C++ & Production ML
Edge AI Inference Engineer — C++ & Production ML

ITRex Group • Town of Poland (NY)

Hybrid
USD 100,000 - 130,000
Remote flexibility
Competitive salary and medical benefits
Career progression opportunities
AI Inference Engineer
AI Inference Engineer

ITRex Group • United States

Remote
USD 120,000 - 180,000
Remote flexibility
Medical benefits
Learning budget
AI Inference Engineer (f/m/d)
AI Inference Engineer (f/m/d)

ITRex Group • Town of Poland (NY)

On-site
USD 100,000 - 130,000
Remote flexibility
Competitive salary and medical benefits
Career progression opportunities
AI Inference Engineer QVAC
AI Inference Engineer QVAC

ITRex Group • United States

On-site
USD 120,000 - 170,000
Remote flexibility
Competitive salary + medical benefits
Learning opportunities
Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH
Inference Runtime Engineer - On-Device & Cloud AI, Flexible WFH

EngRadar • New York (NY)

On-site
USD 150,000 - 230,000
Equity grants
Medical plan
Vision plan
+5
Remote AI Performance Engineer - Scale ML Inference
Remote AI Performance Engineer - Scale ML Inference

United States Digital Space LLC • United States

Remote
USD 75,000 - 100,000
Senior GPU AI Platforms Engineer - Edge LLM Inference
Senior GPU AI Platforms Engineer - Edge LLM Inference

NVIDIA • Durham (NC)

On-site
USD 224,000 - 356,500
Equity
Benefits
AI Inference Engineer - Kernel & Edge Optimization
AI Inference Engineer - Kernel & Edge Optimization

Jobgether SRL • United States

Remote
USD 82,000 - 177,000
Remote-first team
Exposure to cutting-edge AI research
Collaborative engineering environment
Edge AI Inference Engineer: Model Optimization & Deployment
Edge AI Inference Engineer: Model Optimization & Deployment

Zoox • Foster City (CA)

On-site
USD 180,000 - 260,000
Health insurance
Stock options
RSUs
+1