A complete application in a minute — tailored resume and cover letter, ready to send.
Predii is seeking an AI/ML Infrastructure Engineer to own multi‑cloud infrastructure across Azure, GCP, and AWS, leveraging Kubernetes/AKS, Terraform, Ansible, and Helm. You will drive DevOps, CI/CD pipelines, and reliability initiatives while ensuring security and SOC 2 compliance.
You’ll work across a distributed US–India team, exploring new tools and techniques, and mentoring others as needed. Hybrid-friendly, CA-preferred with remote options on the West Coast.
Aug 20, 2026
Mid-Level to Senior | Engineering & Platform Ops
US-based — CA preferred, open to West Coast + remote. Hybrid-friendly.
Predii builds the intelligence layer that runs the automotive service and parts industry. Our platform, Predii 360, turns messy repair-order, DMS, and parts data into real-time intelligence — powering parts lookup, diagnostics, and repair search for dealership and aftermarket customers at scale, processing billions of repair orders and serving live search at sub-second latency.
We're small, fast, and allergic to red tape. No 12-layer approval chains, no work that disappears into a backlog forever. If you build something here, it ships — and real customers use it. Learn more at www.predii.com .
We do real research, not just integration. We continue to submit state-of-the-art research on topics including: engineering-diagram and technical-document understanding, domain-calibrated evaluation frameworks for technical content, multi-agent architectures that optimize for correctness and honesty, detecting "confident-but-wrong" failures that standard monitoring misses, moving from reactive detection to causal, explainable prognosis, and multilingual evaluation of technical and repair content. We've found that multi-agent systems that just concatenate outputs get less trustworthy as they get more capable, so we design ours to contest and qualify each other's findings instead. And we run open-weight models in production at enterprise scale, because repair-grade accuracy shouldn't cost frontier-model money. All of it is deliberately vertical: deep automotive domain expertise applied to automotive problems, not a general-purpose model with an automotive skin.
We need a AI/ML Infrastructure Engineer who wants more than tickets — someone ready to actually own infrastructure across multiple clouds and help shape how we build. This is real ownership, not busywork. You'll touch:
Multi-cloud infra (Azure, GCP, AWS)
Kubernetes, CI/CD, automation-everything
Security, compliance, access — keeping the house locked
Monitoring & reliability — catching problems before customers do
Cloud cost optimization (yes, we care about the bill too)
Senior folks: expect to shape architecture and mentor the team, not just execute someone else's roadmap.
Cloud & Platform — build + run deployments on Azure, GCP, AWS with Kubernetes/AKS, Terraform, Ansible, Helm.
DevOps & CI/CD — ship pipelines that are reliable and repeatable, not held together with duct tape.
Reliability & Observability — build monitoring, logging, alerting, tracing; hunt down root causes, not just symptoms.
Security & Compliance — RBAC, auth, network security, vuln management, SOC 2 Type 2 controls.
Resilience & Ops — own backups, DR, capacity planning, cloud costs, and production support.
Keep Leveling Up — evaluate new tools across DevOps, infra, and DevSecOps; you're not stuck with 2019's stack.
IT Support — jump in on Windows Server / macOS support when needed.
Security & Compliance: SOC 2 Type 2, access/change controls, business continuity + DR
3–5+ years hands‑on AI/ML Infrastructure / DevOps / SRE / Platform Engineering, in production — not just labs.
Strong cloud + Kubernetes chops — Azure/AKS preferred; GCP/AWS/Docker is a big plus.
Solid networking and security fundamentals.
Comfortable with Terraform, Ansible, Helm, Bash, Python (or similar).
Real Git-based CI/CD experience — automated deploys, security scanning included.
Battle‑tested on observability & prod ops — monitoring, incident response, RCA, runbooks, backup, DR.
Working knowledge of security/compliance across Linux, Windows Server, macOS.
A self-starter mindset — comfortable working independently across a distributed US–India team.
Auth/IAM experience, especially multi‑tenant setups.
Background in automotive or data‑heavy platforms.
Been through a SOC 2 Type 2 audit before.
Startup or small‑team energy — you've worn more than one hat.
Cloud, Kubernetes, or Terraform certs.
Senior folks: mentoring or technical leadership experience.
Own it — flag issues early, make the call, follow through.
Be proactive — don't wait to be told; spot the risk, bring the fix.
Share the load — security and reliability are everyone's job, not just yours.
Make it count — your work ships to production and touches real customers, real fast.