Principal Software Platform Engineer | AI Infrastructure & MLOps

Evolution USA

United States

On-site

USD 180,000 - 240,000

Full time

18 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Evolution USA is seeking a Lead Software Platform Engineer to own the architecture, scalability, and operational excellence of a next-generation AI and ML platform. You will lead the design of a cloud-native system that safely deploys LLMs, agents, and AI workloads in production for external users.

You will influence technical direction, standards, and engineering practices while mentoring other engineers across the organization.

Qualifications

  • 10+ years of software engineering and cloud infra experience.
  • Proven ability to design and scale distributed, cloud-native systems in production.
  • Experience as a technical lead, architect, or principal engineer responsible for major architectural decisions.
  • Hands-on experience building and operating AI/ML infrastructure at scale.
  • Strong knowledge of modern LLM architectures and deployment patterns (RAG, embeddings, prompt management).
  • Backend focus with Python and TypeScript, API development.

Responsibilities

  • Own architecture and roadmap for a cloud-native AI and ML platform.
  • Design scalable infra for model deployment, lifecycle, and multi-model serving.
  • Build reliable real-time and batch inference systems.
  • Develop frameworks for LLM apps, retrieval systems, and agent architectures in production.
  • Define platform security, governance, tenant isolation, and data protection standards.
  • Lead design reviews and mentor engineers on architecture and scalability.
  • Drive observability, monitoring, and incident response for production workloads.

Skills

Python
TypeScript
Cloud-native
Distributed systems
LLM deployment
AI infra
CI/CD
Security/compliance
Observability
Leadership

Tools

AWS
Kubernetes
Terraform
Docker
CI/CD pipelines

Job description

A high-growth technology company is seeking a Lead Software Platform Engineer to own the architecture, scalability, and operational excellence of a next-generation AI and machine learning platform. This is a rare opportunity to shape the foundation that enables complex AI workloads to run securely, reliably, and at scale in highly regulated environments.

This role sits at the intersection of distributed systems, cloud infrastructure, AI/ML operations, and software platform engineering. You will serve as a technical leader responsible for defining how AI models, large language models, and intelligent agents are deployed, governed, monitored, and operated in production.

The platform you help build will be a customer-facing product rather than an internal tool. Your work will directly influence how sophisticated organizations apply AI to mission-critical workflows, requiring a balance of innovation, reliability, security, and cost efficiency.

The Opportunity

As a senior technical leader, you will own the strategy and architecture for AI infrastructure that supports production-grade machine learning and generative AI systems. You will collaborate closely with software engineers, data engineers, AI practitioners, and platform teams to design systems that are scalable, observable, secure, and maintainable.

This position is ideal for someone who enjoys solving large-scale technical challenges and has experience building platforms that support external users, stringent compliance requirements, and demanding service-level expectations.

You will have significant influence over technical direction, platform standards, architectural decisions, and engineering best practices while helping mentor and elevate other engineers across the organization.

What You’ll Be Responsible For
  • Owning the architecture and technical roadmap for a cloud-native AI and machine learning platform.
  • Designing and scaling infrastructure that supports model deployment, model lifecycle management, prompt management, and multi-model serving.
  • Building reliable systems for both real-time and batch inference workloads.
  • Developing frameworks for running and operating LLM-based applications, retrieval systems, and intelligent agent architectures in production.
  • Establishing platform-wide standards for security, governance, tenant isolation, and data protection.
  • Creating robust evaluation frameworks that measure model quality, detect regressions, and support safe releases.
  • Defining observability practices including monitoring, alerting, logging, tracing, and performance management.
  • Driving reproducibility, lineage, auditability, and compliance requirements across AI workflows.
  • Contributing to infrastructure automation and deployment processes using modern cloud engineering practices.
  • Leading technical design reviews and acting as a trusted advisor on architecture, scalability, reliability, and platform strategy.
  • Evaluating emerging AI technologies, frameworks, and tooling while making informed build-versus-buy decision.
  • Supporting operational excellence through incident response, production readiness reviews, and continuous improvement initiatives.
What We’re Looking For

You are an experienced platform engineer or technical architect who has spent your career designing and operating large-scale cloud systems. You understand what it takes to move AI and machine learning workloads from experimentation into production environments where reliability, cost control, security, and performance matter.

We're particularly interested in engineers who have built customer-facing AI infrastructure rather than internal-only tools.

Key Requirements
  • 10+ years of software engineering and cloud infrastructure experience.
  • Proven success designing and scaling distributed, cloud-native systems in production.
  • Experience serving as a technical lead, architect, or principal-level engineer responsible for major architectural decisions.
  • Deep hands-on experience building and operating AI/ML infrastructure at scale.
  • Strong knowledge of modern LLM architectures and production deployment patterns, including retrieval-augmented generation (RAG), prompt lifecycle management, embeddings, and tool integration.
  • Advanced coding skills in Python and TypeScript, with a strong focus on backend services and API development.
  • Experience with model lifecycle management, deployment workflows, and model-serving infrastructure.
  • Strong understanding of software quality, automated testing, release management, and CI/CD practices.
  • Familiarity with cloud-native technologies including AWS, containers, and infrastructure-as-code tooling.
  • Experience designing secure multi-tenant systems and handling sensitive or regulated data.
  • Expertise in monitoring, observability, reliability engineering, and service-level management.
  • Excellent communication skills with the ability to influence technical direction across multiple teams.
Highly Desirable Experience
  • Production experience with agentic AI systems and orchestration frameworks.
  • Knowledge of advanced model serving, latency optimization, and cost management techniques.
  • Experience supporting multimodal AI workloads involving images or other complex data types.
  • Exposure to model optimization techniques such as fine-tuning, distillation, or performance tuning.
  • Background operating AI systems within regulated industries where auditability and compliance are critical.
  • Experience working in data-intensive, scientific, healthcare, pharmaceutical, or similarly complex environments.
Why This Role Stands Out

This is a high-impact leadership position where you'll help define the future architecture of a rapidly evolving AI platform. You'll work on meaningful technical challenges involving scalability, reliability, security, governance, and production AI operations while partnering with a team that views engineering excellence as a strategic advantage.

If you're excited by the challenge of building the platforms that power next-generation AI applications and enjoy solving difficult infrastructure problems at scale, we'd love to hear from you.

Applicants must be authorized to work in the United States. Sponsorship is not available for this position.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Platform Engineer (Cloud & AI Platform)
Senior Platform Engineer (Cloud & AI Platform)

OEC • Alpharetta (GA)

On-site
USD 140,000 - 190,000
AI Platform Lead
AI Platform Lead

SeekUp • New York (NY)

On-site
USD 130,000 - 170,000
Senior AI Software Engineer
Senior AI Software Engineer

Harnham • San Francisco (CA)

On-site
USD 180,000 - 260,000
Enterprise Platform Architect
Enterprise Platform Architect

evolver • Palo Alto (CA)

On-site
USD 190,000 - 270,000
Competitive compensation
Hybrid work
Career growth
+1
Senior AI Engineer
Senior AI Engineer

Harnham • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Software Engineer – AI Platform
Senior Software Engineer – AI Platform

PRI Technology • New York (NY)

On-site
USD 160,000 - 240,000
Lead Machine Learning Engineer
Lead Machine Learning Engineer

Motion Recruitment Partners LLC • Raleigh (NC)

On-site
USD 180,000 - 240,000
Lead Engineer - Data Engg & AI
Lead Engineer - Data Engg & AI

Anblicks • Dallas (TX)

On-site
USD 150,000 - 190,000
Manager, Platform Engineering
Manager, Platform Engineering

Johnson Controls • Milwaukee (WI)

On-site
USD 180,000 - 240,000
Competitive salary
Paid vacation/holidays/sick time
Comprehensive benefits (401K, medical,
+3
AI Infrastructure Engineer
AI Infrastructure Engineer

dicedemo • Boston (CT)

On-site
USD 130,000 - 170,000