AI/ML Infrastructure Engineer - USA

Atlassian

Palo Alto, Northern (CA, KY)

Hybrid

USD 150,000 - 210,000

Full time

45 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Predii is seeking an AI/ML Infrastructure Engineer to own multi‑cloud infrastructure across Azure, GCP, and AWS, leveraging Kubernetes/AKS, Terraform, Ansible, and Helm. You will drive DevOps, CI/CD pipelines, and reliability initiatives while ensuring security and SOC 2 compliance.

You’ll work across a distributed US–India team, exploring new tools and techniques, and mentoring others as needed. Hybrid-friendly, CA-preferred with remote options on the West Coast.

Qualifications

  • 3–5+ years hands‑on AI/ML Infrastructure / DevOps / SRE / Platform Engineering, in production — not just labs.
  • Strong cloud + Kubernetes chops — Azure/AKS preferred; GCP/AWS/Docker is a big plus.
  • Solid networking and security fundamentals.
  • Comfortable with Terraform, Ansible, Helm, Bash, Python (or similar).
  • Real Git-based CI/CD experience — automated deploys, security scanning included.
  • Battle‑tested on observability & prod ops — monitoring, incident response, RCA, runbooks, backup, DR.
  • Working knowledge of security/compliance across Linux, Windows Server, macOS.
  • A self-starter mindset — comfortable working independently across a distributed US–India team.

Responsibilities

  • Cloud & Platform — build + run deployments on Azure, GCP, AWS with Kubernetes/AKS, Terraform, Ansible, Helm.
  • DevOps & CI/CD — ship pipelines that are reliable and repeatable.
  • Reliability & Observability — build monitoring, logging, alerting, tracing; hunt down root causes, not just symptoms.
  • Security & Compliance — RBAC, auth, network security, vuln management, SOC 2 Type 2 controls.
  • Resilience & Ops — own backups, DR, capacity planning, cloud costs, and production support.
  • Keep Leveling Up — evaluate new tools across DevOps, infra, and DevSecOps; you're not stuck with 2019's stack.
  • IT Support — jump in on Windows Server / macOS support when needed.

Skills

Azure
AKS
GCP
AWS
Kubernetes
Terraform
Ansible
Helm
Bash
Python
Git CI/CD
Observability
SRE
Security
Windows Server
macOS
Distributed teams

Tools

Docker

Job description

Aug 20, 2026

Mid-Level to Senior | Engineering & Platform Ops

US-based — CA preferred, open to West Coast + remote. Hybrid-friendly.

ABOUT PREDII

Predii builds the intelligence layer that runs the automotive service and parts industry. Our platform, Predii 360, turns messy repair-order, DMS, and parts data into real-time intelligence — powering parts lookup, diagnostics, and repair search for dealership and aftermarket customers at scale, processing billions of repair orders and serving live search at sub-second latency.

We're small, fast, and allergic to red tape. No 12-layer approval chains, no work that disappears into a backlog forever. If you build something here, it ships — and real customers use it. Learn more at www.predii.com .

PREDII RESEARCH

We do real research, not just integration. We continue to submit state-of-the-art research on topics including: engineering-diagram and technical-document understanding, domain-calibrated evaluation frameworks for technical content, multi-agent architectures that optimize for correctness and honesty, detecting "confident-but-wrong" failures that standard monitoring misses, moving from reactive detection to causal, explainable prognosis, and multilingual evaluation of technical and repair content. We've found that multi-agent systems that just concatenate outputs get less trustworthy as they get more capable, so we design ours to contest and qualify each other's findings instead. And we run open-weight models in production at enterprise scale, because repair-grade accuracy shouldn't cost frontier-model money. All of it is deliberately vertical: deep automotive domain expertise applied to automotive problems, not a general-purpose model with an automotive skin.

THE VIBE

We need a AI/ML Infrastructure Engineer who wants more than tickets — someone ready to actually own infrastructure across multiple clouds and help shape how we build. This is real ownership, not busywork. You'll touch:

Multi-cloud infra (Azure, GCP, AWS)

Kubernetes, CI/CD, automation-everything

Security, compliance, access — keeping the house locked

Monitoring & reliability — catching problems before customers do

Cloud cost optimization (yes, we care about the bill too)

Senior folks: expect to shape architecture and mentor the team, not just execute someone else's roadmap.

WHAT YOU'LL ACTUALLY DO

Cloud & Platform — build + run deployments on Azure, GCP, AWS with Kubernetes/AKS, Terraform, Ansible, Helm.

DevOps & CI/CD — ship pipelines that are reliable and repeatable, not held together with duct tape.

Reliability & Observability — build monitoring, logging, alerting, tracing; hunt down root causes, not just symptoms.

Security & Compliance — RBAC, auth, network security, vuln management, SOC 2 Type 2 controls.

Resilience & Ops — own backups, DR, capacity planning, cloud costs, and production support.

Keep Leveling Up — evaluate new tools across DevOps, infra, and DevSecOps; you're not stuck with 2019's stack.

IT Support — jump in on Windows Server / macOS support when needed.

YOUR TOOLKIT

Security & Compliance: SOC 2 Type 2, access/change controls, business continuity + DR

WHAT YOU BRING

3–5+ years hands‑on AI/ML Infrastructure / DevOps / SRE / Platform Engineering, in production — not just labs.

Strong cloud + Kubernetes chops — Azure/AKS preferred; GCP/AWS/Docker is a big plus.

Solid networking and security fundamentals.

Comfortable with Terraform, Ansible, Helm, Bash, Python (or similar).

Real Git-based CI/CD experience — automated deploys, security scanning included.

Battle‑tested on observability & prod ops — monitoring, incident response, RCA, runbooks, backup, DR.

Working knowledge of security/compliance across Linux, Windows Server, macOS.

A self-starter mindset — comfortable working independently across a distributed US–India team.

BONUS POINTS

Auth/IAM experience, especially multi‑tenant setups.

Background in automotive or data‑heavy platforms.

Been through a SOC 2 Type 2 audit before.

Startup or small‑team energy — you've worn more than one hat.

Cloud, Kubernetes, or Terraform certs.

Senior folks: mentoring or technical leadership experience.

HOW WE ROLL

Own it — flag issues early, make the call, follow through.

Be proactive — don't wait to be told; spot the risk, bring the fix.

Share the load — security and reliability are everyone's job, not just yours.

Make it count — your work ships to production and touches real customers, real fast.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Engineer / Architect - USA
AI Engineer / Architect - USA

Atlassian • Palo Alto (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior DevOps Engineer
Senior DevOps Engineer

TTEC Digital • United States

Remote
USD 130,000 - 180,000
Senior Platform Software Engineer
Senior Platform Software Engineer

Cello • United States

Remote
USD 120,000 - 190,000
Senior Site Reliability Engineer, AI Agents & Automation
Senior Site Reliability Engineer, AI Agents & Automation

ServiceTitan • United States

On-site
USD 140,000 - 190,000
Flexible time off
Fully paid medical, dental, and vision
HSA/FSA programs
+7
AI Engineer
AI Engineer

Valsoft Corporation • United States

On-site
USD 140,000 - 230,000
DevOps Engineer, Infrastructure & Platforms
DevOps Engineer, Infrastructure & Platforms

Ricursive Intelligence • Palo Alto (CA)

On-site
USD 150,000 - 210,000
Senior Software Engineer – Data & ML Platform
Senior Software Engineer – Data & ML Platform

GoMaterials • United States

Remote
USD 130,000 - 185,000
AI Solutions Engineer (Delivery Lead)
AI Solutions Engineer (Delivery Lead)

UNKNOWN • Raleigh (NC), Northern (KY)

On-site
USD 140,000 - 200,000
Principal Product Engineer, Cloud Platform
Principal Product Engineer, Cloud Platform

Verdigris Technologies • Palo Alto (CA)

On-site
USD 190,000 - 270,000
Site Reliability Engineer II (SREII)
Site Reliability Engineer II (SREII)

Prodege, LLC • United States

Remote
USD 130,000 - 180,000