Inference Lead

Stellent IT LLC

Charlotte (NC)

Hybrid

USD 96,000 - 138,000

Full time

9 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Stellent IT LLC in Charlotte, North Carolina is seeking an Inference Lead (Machine Learning Platform Engineer Lead Real-Time Inference) for a six-month contract. The role focuses on designing low-latency, resilient model-serving services for predictive AI use cases, with cloud and on-prem deployment.

You will define deployment patterns, implement Kubernetes-based operations, develop scalable APIs, and lead a team of inference and platform engineers while ensuring performance, reliability, and

Qualifications

  • Design and implement low-latency online inference and real-time model-serving systems.
  • Experience building scalable APIs and microservices for predictive models.
  • Proficiency with Kubernetes, autoscaling, and traffic management.
  • Ability to perform benchmarking, latency optimization, and load testing.
  • Strong focus on monitoring, SLOs, capacity planning, and production operations.
  • Experience with CI/CD and progressive deployment approaches.
  • Design for resilience, failover, and disaster recovery.

Responsibilities

  • Architect low-latency online inference and real-time model-serving solutions.
  • Develop scalable APIs, microservices, and deployment patterns for predictive models.
  • Implement Kubernetes-based deployment, autoscaling, load balancing, and traffic-management strategies.
  • Conduct benchmarking, performance tuning, capacity planning, and load testing.
  • Optimize latency, throughput, resource consumption, availability, and cost.
  • Define monitoring, alerting, SLOs, runbooks, and incident-response practices.
  • Build CI/CD pipelines for repeatable model and service releases.
  • Design resilience, failover, disaster recovery, and graceful-degradation patterns.
  • Lead technical reviews and mentor inference and platform engineers.

Skills

Low-latency inference
Kubernetes
REST/gRPC APIs
CI/CD pipelines
Performance tuning
Monitoring & SLOs
Cloud and on-prem deployments
Resilience & high-availability
Distributed microservices

Education

Degree in CS/Engineering or related field

Tools

Kubernetes
Containers

Job description

Job Title:- Inference Lead (Machine Learning Platform Engineer Lead Real-Time Inference)
Location:- Charlotte, North Carolina (Hybrid Onsite - local candidates preferred. )
Duration:- 6 months
All visa except OPT/CPT
Contract
Real-Time Services Real-Time Inference Engineering Lead
Role Summary

The Real-Time Inference Engineering Lead will design and industrialize low-latency, resilient model-serving services for predictive AI use cases. The role will define deployment patterns, capacity controls, monitoring, performance standards, and operational practices across cloud and on-premises environments.

Key Responsibilities
  • Architect low-latency online inference and real-time model-serving solutions.
  • Develop scalable APIs, microservices, and deployment patterns for predictive models.
  • Implement Kubernetes-based deployment, autoscaling, load balancing, and traffic-management strategies.
  • Conduct benchmarking, performance tuning, capacity planning, and load testing.
  • Optimize latency, throughput, resource consumption, availability, and cost.
  • Define monitoring, alerting, SLOs, runbooks, and incident-response practices.
  • Build CI/CD pipelines for repeatable model and service releases.
  • Design resilience, failover, rollback, disaster recovery, and graceful-degradation patterns.
  • Lead technical reviews and mentor inference and platform engineers.
Required Skills
  • Online inference and real-time model-serving architecture.
  • REST/gRPC APIs and distributed microservices.
  • Kubernetes, containers, autoscaling, and traffic management.
  • Performance engineering, latency optimization, and load testing.
  • Monitoring, SLOs, capacity planning, and production operations.
  • CI/CD and progressive-deployment approaches.
  • Resilience and high-availability engineering.
  • Cloud and on-premises deployment experience.
Preferred Qualifications
  • Degree in computer science, engineering, or a related discipline.
  • Experience with enterprise model-serving platforms and inference runtimes.
  • Cloud, Kubernetes, SRE, or ML engineering certification.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Lead
Inference Lead

VDart Inc • Charlotte (NC)

Hybrid
USD 140,000 - 190,000
Inference Lead
Inference Lead

Compunnel, Inc. • Charlotte (NC), Northern (KY)

Hybrid
USD 170,000 - 210,000
Lead ML Platform Engineer
Lead ML Platform Engineer

Creative Solutions Services, LLC • Charlotte (NC)

On-site
USD 80,000 - 89,000
Medical insurance
Dental insurance
Vision insurance
+3
Real-Time Inference Platform Lead
Real-Time Inference Platform Lead

NTT DATA North America • Charlotte (NC)

Hybrid
USD 84,000 - 125,000
Medical, dental, vision insurance
401k with company match
Paid time off
Real-Time Inference Platform Lead
Real-Time Inference Platform Lead

Creative Solutions Services, LLC • Charlotte (NC)

On-site
USD 80,000 - 89,000
Medical insurance
Dental insurance
Vision insurance
+3
Real-Time Inference Platform Lead
Real-Time Inference Platform Lead

NTT DATA, Inc. • Charlotte (NC)

Hybrid
USD 84,000 - 125,000
Medical, dental, and vision insurance
401k with company match
Paid time off
+2
Real-Time ML Inference Lead — Latency & Platform Architect
Real-Time ML Inference Lead — Latency & Platform Architect

Stellent IT LLC • Charlotte (NC)

Hybrid
USD 96,000 - 138,000
Real-Time Inference Architect & Platform Lead
Real-Time Inference Architect & Platform Lead

Compunnel, Inc. • Charlotte (NC), Northern (KY)

Hybrid
USD 170,000 - 210,000
Real-Time Inference Engineering Lead (FTE / Hybrid)
Real-Time Inference Engineering Lead (FTE / Hybrid)

NTT DATA, Inc. • Charlotte (NC)

Hybrid
USD 84,000 - 125,000
Medical, dental, and vision insurance
401k with company match
Paid time off
+2
Real-Time AI Inference Engineering Lead
Real-Time AI Inference Engineering Lead

VDart Inc • Charlotte (NC)

Hybrid
USD 140,000 - 190,000