AI Infrastructure Engineer

Propio

Overland Park (KS)

Hybrid

USD 140,000 - 200,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Propio is seeking an AI Infrastructure Engineer in Overland Park, hybrid. You will design, build, and operate AWS-based, GPU-accelerated inference and streaming serving platforms for LLMs, ASR, and multimodal models, while supporting research training environments and ML/LLMOps capabilities.

You'll collaborate with researchers and edge teams to enable scalable, low-latency workloads and implement secure, observable infrastructure.

Qualifications

  • 3+ years building AI/ML infrastructure, inference platforms, or distributed systems in production.
  • Hands-on AWS experience with EKS, EC2 GPU workloads, ECR, S3, IAM/KMS, VPC networking, CloudWatch and/or OpenTelemetry.
  • Hands-on with at least one inference stack (vLLM, SGLang, TensorRT-LLM, Triton, KServe, or Ray Serve).
  • Experience operating ML/LLM systems in production, including model serving, autoscaling, monitoring, incident response, and performance benchmarking.
  • Familiarity with GPU infra and at least one serving stack (vLLM, SGLang, TensorRT-LLM, or Triton).
  • Working understanding of LLM Ops practices, including evaluation, observability and tracing, cost control, and versioning.

Responsibilities

  • Real-time inference systems: design, deploy, and operate low-latency streaming inference on AWS for LLMs and multimodal models.
  • Support research training environments: reproducible containers, GPU job scheduling, distributed execution, experiment tracking.
  • Build ML/LLMOps pipelines: registries, versioning, automated evaluation gates, deployments, canary and rollback workflows.
  • Develop Edge AI workflows: model optimization, packaging, validation, deployment for edge targets.
  • Ensure reliability and security: capacity planning, incident response, disaster recovery, IAM/KMS, private networking.

Skills

AI/ML infrastructure experience
AWS production experience
Inference stack knowledge
LLMOps practices understanding
GPU infrastructure familiarity
Production systems monitoring

Tools

vLLM
SGLang
TensorRT-LLM
Triton
KServe
Ray Serve
Megatron-LM
Slurm
SageMaker

Job description

Overland Park, KS (Hybrid) AI Engineering

Job Type

Full-time

Description

Own Your Impact. At Propio, we don't believe careers happen to people. We believe people create them.

Here, you're trusted to make decisions, challenge assumptions, drive innovation, and shape outcomes. Your success is not limited by hierarchy or tenure. It's fueled by your ambition, your curiosity, and your willingness to own your impact. If you're looking for a role where you can simply maintain the status quo, this probably isn't it, but if you're looking for a place where your ideas matter, your growth is accelerated, and your work creates meaningful impact across the world, we'd love to talk.

Why Propio?

Every day, communication changes lives. A patient receives care they otherwise couldn't access. A family gains critical information. A business connects with a customer. A community becomes more inclusive. These moments happen because barriers are removed. And behind those moments are Propio team members who show up every day to solve problems, innovate, and build the future. This isn't just work. This is world impact.

As an AI Infrastructure Engineer, you will primarily design, build, and operate our AWS-based, GPU-accelerated inference and streaming serving platform, optimizing it for low latency, high concurrency, reliability, and cost. You will also support the research team’s training environment, further the AI team’s ML/LLMOps capabilities, and enable the development of edge AI deployments.

You’ll be empowered to:
  • Take ownership of important initiatives and outcomes.
  • Drive meaningful business results.
  • Influence decisions and contribute new ideas.
  • Partner with talented, high-performing team members.
  • Challenge yourself through continuous learning and growth.
  • Help shape the future of a rapidly growing organization
What You'll Own
  • Real-time inference and serving. Design, deploy, and operate low-latency streaming inference systems on AWS for LLM, ASR, TTS, and multimodal models. Optimize time to first token/audio, p95/p99 end-to-end latency, real-time factor, throughput, concurrency, GPU utilization, and cost per stream using technologies such as vLLM, SGLang, TensorRT-LLM, Triton, TensorRT.
  • Training environment. Support and evolve the research team’s AWS-based training environment, including reproducible containers, GPU job scheduling, distributed job execution, checkpoint and resume capabilities, experiment tracking, model and data artifact access, and researcher self-service. Enable workloads using FSDP, DeepSpeed, Megatron-LM, Slurm, EKS, or SageMaker.
  • ML/LLMOps. Build model and artifact registries, lineage and versioning, automated evaluation gates, deployment pipelines, shadow and canary releases, rollback workflows, runtime and configuration management, and production observability
  • Edge AI infrastructure. Partner with researchers and device and embedded teams to build repeatable model optimization, packaging, validation, and deployment workflows for resource-constrained edge targets. Support model export and compilation, post-training quantization, runtime integration.
  • Reliability and security. Own capacity planning, production readiness, incident response, disaster recovery, and the secure operation of the AI platform. Apply least-privilege IAM, KMS encryption, private networking, secrets management.
What Makes Someone Successful Here

The most successful people at Propio aren't necessarily the ones with the longest resumes. They're the people who:

  • Take ownership instead of waiting for direction.
  • Embrace challenges as opportunities to grow.
  • Continuously seek better ways of working.
  • Turn ideas into action.
  • Hold themselves and others accountable to high standards.
  • Are driven by making a measureable impact.
Requirements
What You'll Bring

Required Qualifications:

  • 3+ years of experience building and operating AI/ML infrastructure, inference platforms, or distributed systems in production.
  • Hands-on production experience with AWS, particularly EKS, EC2 GPU workloads, ECR, S3, IAM/KMS, VPC networking, and CloudWatch and/or OpenTelemetry.
  • Hands‑on experience with at least one inference stack, such as vLLM, SGLang, TensorRT-LLM, Triton, KServe, or Ray Serve.
  • Experience operating ML/LLM systems in production, including model serving, autoscaling, monitoring, incident response, and performance benchmarking.
  • Familiarity with GPU infrastructure and at least one serving stack, such as vLLM, SGLang, TensorRT-LLM, or Triton.
  • A working understanding of LLM Ops practices, including evaluation, observability and tracing, cost control, and versioning.
Preferred Qualifications
  • Low‑latency, real‑time, or streaming inference experience, especially for audio/speech—directly relevant to interpretation.
  • Experience with real‑time audio pipelines and streaming protocols, including WebRTC, WebSocket, and gRPC streaming.
  • Experience supporting distributed training environments using FSDP, DeepSpeed, Megatron-LM, Slurm, EKS, or SageMaker HyperPod.
  • Inference optimization experience, including quantization, KV‑cache optimization, speculative decoding, and continuous batching.
  • Experience building Edge AI deployment toolchains using ONNX Runtime, TensorRT/Jetson, ExecuTorch, llama.cpp, MLC‑LLM, or similar runtimes.
  • Experience with SRE practices, capacity planning, production incident response, and secure infrastructure for PHI/PII.
  • Contributions to relevant open‑source infrastructure, serving, observability, or edge‑runtime projects.

Even if your experience doesn't perfectly match every qualification, we encourage you to apply. We’re looking for potential, drive, and a commitment to growth as much as experience.

What You'll Gain
  • Own Your Growth: We invest in people who invest in themselves. You’ll have opportunities to learn, develop, and expand your capabilities while building a meaningful career.
  • Own Your Impact: You’ll see the connection between your work and our success. We believe great people deserve the opportunity to make a real difference.
  • Own Your Innovation: The best ideas can come from anywhere. We encourage curiosity, creativity, and challenging the way things have always been done.
  • Own Your Success: Whether you’re building expertise, pursuing leadership opportunities, or expanding your career path, we’ll give you room to grow and the support to get there.

At Propio, your work doesn't just move a company forward. It helps connect people, communities, and opportunities across the world.

Notice of AI Use in Job Application Review

As part of our commitment in creating a fair, efficient, and consistent hiring process we may use artificial intelligence (AI) to help our recruiting teams organize, summarize, and analyze information provided by candidates, including resumes, application responses, and other materials submitted during the application process.AI may be used to identify patterns, highlight relevant skills, and experience, and assist in comparing a candidate’s qualifications with the requirement of a specific role. These tools are to improve efficiency and consistency while supporting more informed hiring decisions, which will ultimately be made by the hiring team.

#LI-JS1

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Applied Scientist/Research Engineer, LLM Training Data
Applied Scientist/Research Engineer, LLM Training Data

Propio Language Services • Overland Park (KS)

On-site
USD 120,000 - 180,000
Senior AI Engineer
Senior AI Engineer

Propio • Overland Park (KS)

On-site
USD 120,000 - 180,000
Senior AI Engineer
Senior AI Engineer

Propio Language Services • Overland Park (KS)

On-site
USD 120,000 - 210,000
Client Success Administrator
Client Success Administrator

Propio Language Services • Overland Park (KS)

On-site
USD 55,000 - 75,000
Principal Engineer, UAIS
Principal Engineer, UAIS

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 302,000 - 336,000
Principal AI Engineer
Principal AI Engineer

Phizenix • New York (NY)

On-site
USD 180,000 - 260,000
AI Principal Technical Consultant, AI Services
AI Principal Technical Consultant, AI Services

AHEAD • Chicago (IL)

On-site
USD 230,000 - 300,000
Medical, Dental, and Vision Insurance
401(k)
Paid company holidays
+3
Senior AI Services Architect & Delivery Lead
Senior AI Services Architect & Delivery Lead

Medium • Chicago (IL)

On-site
USD 230,000 - 300,000
DevOps & AI / ML Infrastructure Engineer CreatorIQ · Remote · US · ML Platform & Ops $121,000–$155,000 5d ago
DevOps & AI / ML Infrastructure Engineer CreatorIQ · Remote · US · ML Platform & Ops $121,000–$155,000 5d ago

Aimlroles • New York (NY)

Hybrid
USD 120,000 - 170,000
Vacation & holidays
Wellness benefits
401(k) plan
+2
Principal Engineer
Principal Engineer

your Jared • Northern (KY), San Diego (CA)

Hybrid
USD 180,000 - 240,000
Fully remote, work from home
Employee Share Option Plan
Flexible working hours
+4