Inference Lead

VDart Inc

Charlotte (NC)

Hybrid

USD 140,000 - 190,000

Full time

9 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

VDart Inc. is seeking an Inference Lead in Charlotte, NC. The role focuses on designing and industrializing low-latency model-serving services for predictive AI, with hybrid work options and cross-cloud/on-prem deployment considerations.

The candidate will define deployment patterns, capacity controls, monitoring standards, and operational practices to ensure reliable, scalable inference across environments. Experience leading inference and platform teams is preferred.

Qualifications

  • Experience with real-time model serving
  • Proficient in designing scalable inference platforms
  • Familiarity with cloud and on-prem deployments

Responsibilities

  • Architect low-latency online inference and real-time model-serving solutions.
  • Develop scalable APIs, microservices, and deployment patterns for predictive models.
  • Implement Kubernetes-based deployment, autoscaling, load balancing, and traffic-management strategies.
  • Conduct benchmarking, performance tuning, capacity planning, and load testing.
  • Design monitoring, SLOs, runbooks, and incident-response practices.

Skills

Online inference architecture
REST/gRPC APIs
Kubernetes
Performance optimization
Monitoring & SLOs
CI/CD pipelines
Resilience & HA
Cloud & on-prem deployment

Education

Degree in CS/Engineering

Job description

Role: Inference Lead
Location: Charlotte, NC (Hybrid)
Type: Contract
Role Summary:
  • The Real-Time Inference Engineering Lead will design and industrialize low-latency, resilient model-serving services for predictive AI use cases.
  • The role will define deployment patterns, capacity controls, monitoring, performance standards, and operational practices across cloud and on-premises environments.
Key Responsibilities:
  • Architect low-latency online inference and real-time model-serving solutions.
  • Develop scalable APIs, microservices, and deployment patterns for predictive models.
  • Implement Kubernetes-based deployment, autoscaling, load balancing, and traffic-management strategies.
  • Conduct benchmarking, performance tuning, capacity planning, and load testing.
  • Optimize latency, throughput, resource consumption, availability, and cost.Define monitoring, alerting, SLOs, runbooks, and incident-response practices.
  • Build CI/CD pipelines for repeatable model and service releases.
  • Design resilience, failover, rollback, disaster recovery, and graceful-degradation patterns.
  • Lead technical reviews and mentor inference and platform engineers.
Required Skills:
  • Online inference and real-time model-serving architecture.
  • REST/gRPC APIs and distributed microservices.
  • Kubernetes, containers, autoscaling, and traffic management.
  • Performance engineering, latency optimization, and load testing.
  • Monitoring, SLOs, capacity planning, and production operations.
  • CI/CD and progressive-deployment approaches.
  • Resilience and high-availability engineering.
  • Cloud and on-premises deployment experience.
Preferred Qualifications
  • Degree in computer science, engineering, or a related discipline.
  • Experience with enterprise model-serving platforms and inference runtimes.
  • Cloud, Kubernetes, SRE, or ML engineering certification.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Lead
Inference Lead

Compunnel, Inc. • Charlotte (NC), Northern (KY)

Hybrid
USD 170,000 - 210,000
Inference Lead
Inference Lead

Stellent IT LLC • Charlotte (NC)

Hybrid
USD 96,000 - 138,000
Real-Time Inference Platform Lead
Real-Time Inference Platform Lead

NTT DATA North America • Charlotte (NC)

Hybrid
USD 84,000 - 125,000
Medical, dental, vision insurance
401k with company match
Paid time off
Lead ML Platform Engineer
Lead ML Platform Engineer

Creative Solutions Services, LLC • Charlotte (NC)

On-site
USD 80,000 - 89,000
Medical insurance
Dental insurance
Vision insurance
+3
Real-Time Inference Platform Lead
Real-Time Inference Platform Lead

Creative Solutions Services, LLC • Charlotte (NC)

On-site
USD 80,000 - 89,000
Medical insurance
Dental insurance
Vision insurance
+3
Real-Time Inference Architect & Platform Lead
Real-Time Inference Architect & Platform Lead

Compunnel, Inc. • Charlotte (NC), Northern (KY)

Hybrid
USD 170,000 - 210,000
Real-Time Inference Platform Lead
Real-Time Inference Platform Lead

NTT DATA, Inc. • Charlotte (NC)

Hybrid
USD 84,000 - 125,000
Medical, dental, and vision insurance
401k with company match
Paid time off
+2
Real-Time AI Inference Engineering Lead
Real-Time AI Inference Engineering Lead

VDart Inc • Charlotte (NC)

Hybrid
USD 140,000 - 190,000
Real-Time Inference Engineering Lead (FTE / Hybrid)
Real-Time Inference Engineering Lead (FTE / Hybrid)

NTT DATA, Inc. • Charlotte (NC)

Hybrid
USD 84,000 - 125,000
Medical, dental, and vision insurance
401k with company match
Paid time off
+2
Real-Time ML Inference Lead — Latency & Platform Architect
Real-Time ML Inference Lead — Latency & Platform Architect

Stellent IT LLC • Charlotte (NC)

Hybrid
USD 96,000 - 138,000