Real-Time Inference Architect & Platform Lead

Compunnel, Inc.

Charlotte, Northern (NC, KY)

Hybrid

USD 170,000 - 210,000

Full time

7 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Compunnel, Inc. is seeking a Real-Time Inference Engineering Lead to design and industrialize low-latency model-serving services for predictive AI use cases.

You will define deployment patterns, capacity controls, monitoring, and high-availability practices across cloud and on-premises environments. The role requires strong expertise in real-time inference architecture, distributed services, Kubernetes, performance engineering, CI/CD, and technical leadership to mentor inference and platform

Qualifications

  • 8–12 years of professional experience in software engineering, inference engineering, platform engineering, ML engineering, SRE, or related technical disciplines.
  • Strong experience with online inference and real-time model-serving architecture.
  • Hands-on experience designing REST and/or gRPC APIs and distributed microservices.
  • Strong experience with Kubernetes, containers, autoscaling, and traffic management.
  • Experience with performance engineering, latency optimization, benchmarking, and load testing.
  • Experience with monitoring, SLOs, capacity planning, and production operations.
  • Strong experience building CI/CD pipelines and implementing progressive-deployment approaches.
  • Experience designing resilient, highly available systems, including failover, rollback, disaster recovery, and graceful degradation.
  • Experience deploying and operating solutions across both cloud and on-premises environments.
  • Strong technical leadership skills with experience conducting technical reviews and mentoring engineering teams.

Responsibilities

  • Architect low-latency online inference and real-time model-serving solutions.
  • Develop scalable APIs, microservices, and deployment patterns for predictive models.
  • Implement Kubernetes-based deployment, autoscaling, load balancing, and traffic-management strategies.
  • Conduct benchmarking, performance tuning, capacity planning, and load testing.
  • Optimize latency, throughput, resource consumption, availability, and operational cost.
  • Define monitoring, alerting, SLOs, runbooks, and incident-response practices.
  • Build CI/CD pipelines to support repeatable model and service releases.
  • Design resilience, failover, disaster recovery, and graceful-degradation patterns.
  • Lead technical reviews and provide mentorship to inference and platform engineers.
  • Establish operational and performance standards for real-time model-serving services across cloud and on-premises environments.

Skills

Real-time inference
REST/gRPC APIs
Performance engineering
SRE/production ops
Technical leadership

Education

Bachelor's degree in Computer Science or related

Tools

Kubernetes
Containers (Docker)
Autoscaling
Load balancing
CI/CD pipelines

Job description

Compunnel, Inc. is seeking a Real-Time Inference Engineering Lead to design and industrialize low-latency model-serving services for predictive AI use cases.

You will define deployment patterns, capacity controls, monitoring, and high-availability practices across cloud and on-premises environments. The role requires strong expertise in real-time inference architecture, distributed services, Kubernetes, performance engineering, CI/CD, and technical leadership to mentor inference and platform

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Lead
Inference Lead

Compunnel, Inc. • Charlotte (NC), Northern (KY)

Hybrid
USD 170,000 - 210,000
Inference Lead
Inference Lead

VDart Inc • Charlotte (NC)

Hybrid
USD 140,000 - 190,000
Real-Time Inference Platform Lead
Real-Time Inference Platform Lead

Creative Solutions Services, LLC • Charlotte (NC)

On-site
USD 80,000 - 89,000
Medical insurance
Dental insurance
Vision insurance
+3
Real-Time Inference Platform Lead
Real-Time Inference Platform Lead

NTT DATA North America • Charlotte (NC)

Hybrid
USD 84,000 - 125,000
Medical, dental, vision insurance
401k with company match
Paid time off
Inference Lead
Inference Lead

Stellent IT LLC • Charlotte (NC)

Hybrid
USD 96,000 - 138,000
Real-Time Inference Platform Lead
Real-Time Inference Platform Lead

NTT DATA, Inc. • Charlotte (NC)

Hybrid
USD 84,000 - 125,000
Medical, dental, and vision insurance
401k with company match
Paid time off
+2
Real-Time ML Inference Lead — Latency & Platform Architect
Real-Time ML Inference Lead — Latency & Platform Architect

Stellent IT LLC • Charlotte (NC)

Hybrid
USD 96,000 - 138,000
Real-Time AI Inference Engineering Lead
Real-Time AI Inference Engineering Lead

VDart Inc • Charlotte (NC)

Hybrid
USD 140,000 - 190,000
Senior Engineering Manager – AI Inference Platform
Senior Engineering Manager – AI Inference Platform

Neura Market • Bellevue (WA), Northern (KY)

Hybrid
USD 188,000 - 303,000
Medical, dental, vision
Equity awards
ESPP
+2
Senior AI Inference Platform Engineer
Senior AI Inference Platform Engineer

Cloudflare • Austin (TX)

Hybrid
USD 180,000 - 240,000