Inference Lead

Compunnel, Inc.

Charlotte, Northern (NC, KY)

Hybrid

USD 170,000 - 210,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Compunnel, Inc. is seeking a Real-Time Inference Engineering Lead to design and industrialize low-latency model-serving services for predictive AI use cases.

You will define deployment patterns, capacity controls, monitoring, and high-availability practices across cloud and on-premises environments. The role requires strong expertise in real-time inference architecture, distributed services, Kubernetes, performance engineering, CI/CD, and technical leadership to mentor inference and platform

Qualifications

  • 8–12 years of professional experience in software engineering, inference engineering, platform engineering, ML engineering, SRE, or related technical disciplines.
  • Strong experience with online inference and real-time model-serving architecture.
  • Hands-on experience designing REST and/or gRPC APIs and distributed microservices.
  • Strong experience with Kubernetes, containers, autoscaling, and traffic management.
  • Experience with performance engineering, latency optimization, benchmarking, and load testing.
  • Experience with monitoring, SLOs, capacity planning, and production operations.
  • Strong experience building CI/CD pipelines and implementing progressive-deployment approaches.
  • Experience designing resilient, highly available systems, including failover, rollback, disaster recovery, and graceful degradation.
  • Experience deploying and operating solutions across both cloud and on-premises environments.
  • Strong technical leadership skills with experience conducting technical reviews and mentoring engineering teams.

Responsibilities

  • Architect low-latency online inference and real-time model-serving solutions.
  • Develop scalable APIs, microservices, and deployment patterns for predictive models.
  • Implement Kubernetes-based deployment, autoscaling, load balancing, and traffic-management strategies.
  • Conduct benchmarking, performance tuning, capacity planning, and load testing.
  • Optimize latency, throughput, resource consumption, availability, and operational cost.
  • Define monitoring, alerting, SLOs, runbooks, and incident-response practices.
  • Build CI/CD pipelines to support repeatable model and service releases.
  • Design resilience, failover, disaster recovery, and graceful-degradation patterns.
  • Lead technical reviews and provide mentorship to inference and platform engineers.
  • Establish operational and performance standards for real-time model-serving services across cloud and on-premises environments.

Skills

Real-time inference
REST/gRPC APIs
Performance engineering
SRE/production ops
Technical leadership

Education

Bachelor's degree in Computer Science or related

Tools

Kubernetes
Containers (Docker)
Autoscaling
Load balancing
CI/CD pipelines

Job description

The Real-Time Inference Engineering Lead will design and industrialize low-latency, resilient model-serving services for predictive AI use cases. The role will define deployment patterns, capacity controls, monitoring, performance standards, and operational practices across cloud and on-premises environments. This position requires strong expertise in real-time inference architecture, distributed services, Kubernetes, performance engineering, production operations, CI/CD, and high-availability engineering, along with the ability to provide technical leadership and mentorship to inference and platform engineering teams.

Key Responsibilities
  • Architect low-latency online inference and real-time model-serving solutions.
  • Develop scalable APIs, microservices, and deployment patterns for predictive models.
  • Implement Kubernetes-based deployment, autoscaling, load balancing, and traffic-management strategies.
  • Conduct benchmarking, performance tuning, capacity planning, and load testing.
  • Optimize latency, throughput, resource consumption, availability, and operational cost.
  • Define monitoring, alerting, Service Level Objectives (SLOs), runbooks, and incident-response practices.
  • Build CI/CD pipelines to support repeatable model and service releases.
  • Design resilience, failover, rollback, disaster recovery, and graceful-degradation patterns.
  • Lead technical reviews and provide mentorship to inference and platform engineers.
  • Establish operational and performance standards for real-time model-serving services across cloud and on-premises environments.
Required Qualifications
  • 8–12 years of professional experience in software engineering, inference engineering, platform engineering, ML engineering, SRE, or related technical disciplines.
  • Strong experience with online inference and real-time model-serving architecture.
  • Hands-on experience designing REST and/or gRPC APIs and distributed microservices.
  • Strong experience with Kubernetes, containers, autoscaling, and traffic management.
  • Experience with performance engineering, latency optimization, benchmarking, and load testing.
  • Experience with monitoring, SLOs, capacity planning, and production operations.
  • Strong experience building CI/CD pipelines and implementing progressive-deployment approaches.
  • Experience designing resilient, highly available systems, including failover, rollback, disaster recovery, and graceful degradation.
  • Experience deploying and operating solutions across both cloud and on-premises environments.
  • Strong technical leadership skills with experience conducting technical reviews and mentoring engineering teams.
Preferred Qualifications
  • Degree in Computer Science, Engineering, or a related discipline.
  • Experience with enterprise model-serving platforms and inference runtimes.
  • Cloud, Kubernetes, SRE, or ML engineering certification.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Lead
Inference Lead

VDart Inc • Charlotte (NC)

Hybrid
USD 140,000 - 190,000
Inference Lead
Inference Lead

Stellent IT LLC • Charlotte (NC)

Hybrid
USD 96,000 - 138,000
Real-Time Inference Architect & Platform Lead
Real-Time Inference Architect & Platform Lead

Compunnel, Inc. • Charlotte (NC), Northern (KY)

Hybrid
USD 170,000 - 210,000
Lead ML Platform Engineer
Lead ML Platform Engineer

Creative Solutions Services, LLC • Charlotte (NC)

On-site
USD 80,000 - 89,000
Medical insurance
Dental insurance
Vision insurance
+3
Real-Time Inference Platform Lead
Real-Time Inference Platform Lead

NTT DATA North America • Charlotte (NC)

Hybrid
USD 84,000 - 125,000
Medical, dental, vision insurance
401k with company match
Paid time off
Staff Software Engineer, AI Inference
Staff Software Engineer, AI Inference

ChatGPT Jobs • New York (NY)

On-site
USD 180,000 - 240,000
Health Insurance
Equity
Real-Time Inference Platform Lead
Real-Time Inference Platform Lead

Creative Solutions Services, LLC • Charlotte (NC)

On-site
USD 80,000 - 89,000
Medical insurance
Dental insurance
Vision insurance
+3
INFERENCE ENGINEER
INFERENCE ENGINEER

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff + Senior Software Engineer, Inference
Staff + Senior Software Engineer, Inference

Anthropic • New York (NY)

Hybrid
USD 320,000 - 485,000
Real-Time Inference Platform Lead
Real-Time Inference Platform Lead

NTT DATA, Inc. • Charlotte (NC)

Hybrid
USD 84,000 - 125,000
Medical, dental, and vision insurance
401k with company match
Paid time off
+2