Senior Cloud AI Inference Engineer

Arm Limited

Seattle (WA)

Hybrid

USD 209,000 - 283,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Relocation package
Visa sponsorship

Job summary

Arm seeks an engineer for its AI Inference Cloud team to define the technical direction and develop highly available, scalable services for running AI inference workloads.

You will guide architecture and actively contribute across Kubernetes orchestration, workload management, service delivery, and observability, partnering with AI compute, Inference Runtime, and product teams to enhance performance and usability of Arm's AI platform.

Qualifications

  • 5+ years building distributed systems, cloud platforms, or production infrastructure.
  • Deep production experience with Kubernetes, including controllers, operators, scheduling, resource management, networking, and workload lifecycle management.
  • Strong software and production engineering skills in Go, C++, Rust, Python or similar, with reliable services, APIs, concurrency, observability, and incident response.
  • Track record of leading complex technical initiatives while remaining hands-on across architecture, implementation, debugging, mentoring, and influencing direction.
  • Ability to troubleshoot complex systems and communicate clearly across technical backgrounds.

Responsibilities

  • Define and build the architecture for cloud-based AI inference services.
  • Develop Kubernetes controllers and platform capabilities supporting workload deployment, scheduling, recovery, scaling, upgrades, and lifecycle management.
  • Establish production practices for health validation, progressive rollout, rollback, observability, and service objectives, improving reliability and efficiency.
  • Lead production readiness reviews and resolve issues across services, Kubernetes, networking, and compute infrastructure.
  • Lead design and build reviews, mentor engineers, and drive technical alignment across teams.

Skills

Kubernetes production experience
Go / C++ / Rust / Python
Leadership of technical initiatives
Troubleshooting complex systems
Cross-team communication

Job description

Arm seeks an engineer for its AI Inference Cloud team to define the technical direction and develop highly available, scalable services for running AI inference workloads.

You will guide architecture and actively contribute across Kubernetes orchestration, workload management, service delivery, and observability, partnering with AI compute, Inference Runtime, and product teams to enhance performance and usability of Arm's AI platform.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal AI Inference Cloud Architect
Principal AI Inference Cloud Architect

Arm Limited • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Senior AI Compute Infrastructure Architect
Senior AI Compute Infrastructure Architect

Arm Limited • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Relocation package
Visa sponsorship
Staff Software Engineer, AI Inference Cloud
Staff Software Engineer, AI Inference Cloud

Arm Limited • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Relocation package
Visa sponsorship
Senior AI Inference Runtime Engineer - Distributed
Senior AI Inference Runtime Engineer - Distributed

Arm Limited • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Principal Software Engineer, AI Inference Cloud
Principal Software Engineer, AI Inference Cloud

Arm Limited • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Principal AI Inference Runtime Architect
Principal AI Inference Runtime Architect

Arm Limited • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Hybrid working
Accommodations during recruitment
Senior AI Inference Platform Engineer
Senior AI Inference Platform Engineer

Cloudflare • Austin (TX)

Hybrid
USD 180,000 - 240,000
Staff Cloud Inference Launch Engineer for Scalable AI
Staff Cloud Inference Launch Engineer for Scalable AI

Anthropic • San Francisco (CA)

On-site
USD 320,000 - 485,000
Principal Software Engineer, AI Inference Runtime
Principal Software Engineer, AI Inference Runtime

Arm Limited • Seattle (WA)

Hybrid
USD 263,000 - 355,000
Hybrid working
Accommodations during recruitment
Senior AI Solutions Engineer – Inference & Cloud
Senior AI Solutions Engineer – Inference & Cloud

Kindredventures • San Mateo (CA)

On-site
USD 180,000 - 240,000