Distinguished Software Engineer – AI/ML Engineer

Jobtailor

Sunnyvale (CA)

On-site

USD 230,000 - 350,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor in Sunnyvale, CA, seeks a senior engineer to lead technical development of agentic AI systems for reliability and automation, architect ML platforms, and build multi-agent orchestration for change management and scalability.

You will design observability, self-healing infrastructure, and fault-tolerant hybrid-cloud deployments, mentor teams, and advance MLOps/AIOps for autonomous optimization.

Qualifications

  • Bachelor’s degree or higher in a technical field with extensive software engineering experience.
  • 12+ years of hands-on experience in Reliability/AI-ML/Platform Engineering.
  • Proven track record influencing architecture and driving technical excellence in large organizations.
  • Deep experience operating mission-critical systems.
  • Expert-level cloud and observability expertise across multiple platforms.

Responsibilities

  • Lead technical development of next-generation agentic AI systems and intelligent automation solutions.
  • Architect and implement ML platforms and autonomous agents for change management, monitoring, prediction, and automated issue resolution.
  • Design and implement multi-agent orchestration platforms for change management, capacity planning, and performance optimization.
  • Build intelligent observability and monitoring platforms using ML-driven anomaly detection, predictive analytics, and autonomous resolution.
  • Develop self-healing infrastructure platforms that predict, prevent, and automatically remediate system issues.

Skills

Reliability Engineering
AI/ML Engineering
Platform Engineering
Multi-Agent Frameworks
Cloud Engineering

Education

Bachelor’s degree in CS/Engineering or related area
Master’s degree preferred

Tools

TensorFlow
PyTorch
Kubernetes
Docker
CI/CD Pipelines
Infrastructure as Code
MLflow
Kubeflow
Seldon
Azure
GCP
AWS
Kafka
Pulsar

Job description

  • Lead technical development of next-generation agentic AI systems and intelligent automation solutions for mission-critical reliability, scalability, and operational excellence
  • Architect and implement machine learning platforms and autonomous agents for change management, monitoring, prediction, and automated issue resolution
  • Design and implement multi-agent orchestration platforms for change management, capacity planning, and performance optimization
  • Build intelligent observability and monitoring platforms using ML-driven anomaly detection, predictive analytics, and autonomous resolution
  • Develop self-healing infrastructure platforms that predict, prevent, and automatically remediate system issues
  • Design and build tools improving latency, availability, scalability, and change management
  • Engineer reliability using metrics and measurements across domains
  • Enable system scaling through technical solutions, automation, and process optimization
  • Build failure-prevention tools and automation for mission-critical services
  • Enhance instrumentation for cohesive end-to-end system health visibility
  • Architect and implement fault-tolerant systems across hybrid cloud infrastructure
  • Reduce MTTD and MTTR through intelligent automation and predictive capabilities
  • Partner with service owners to define SLA breach detection and change-related anomaly handling
  • Troubleshoot and analyze large-scale distributed systems
  • Deliver autonomous reliability solutions using machine learning, NLP, and computer vision
  • Drive development of MLOps and AIOps platforms for continuous learning, deployment, monitoring, and autonomous optimization
  • Implement CI/CD pipelines with automated validation, deployment, rollback, and observability
  • Build reusable reliability infrastructure, intelligent monitoring platforms, and developer productivity tools
  • Provide technical mentorship and thought leadership through code reviews, design discussions, and knowledge sharing
Requirements
  • Bachelor’s degree in computer science, computer engineering, computer information systems, software engineering, or related area and 6 years’ experience in software engineering or related area; OR 8 years’ experience in software engineering or related area
  • 12+ years of hands-on experience in Reliability Engineering, AI/ML Engineering, or Platform Engineering
  • Proven record as a senior individual contributor influencing architecture and driving technical excellence across large organizations
  • Deep experience operating mission-critical systems
  • Expertise in MTTD, MTTR, availability, change management, model performance, and autonomous system reliability
  • Expert-level AI/ML engineering experience, including TensorFlow, PyTorch, and large-scale production ML deployments
  • Advanced experience with agentic AI systems, multi-agent frameworks, autonomous decision-making systems, LLM-based agents, and agent orchestration platforms
  • Comprehensive Reliability Engineering expertise, including Incident, Problem, and Change Management and performance and capacity engineering for AI/ML systems
  • Expert-level cloud engineering experience with Azure, GCP, or AWS
  • Experience with Kubernetes, Docker, serverless architectures, and cloud-native AI services
  • Deep observability experience across distributed tracing, metrics, logs, APM, and AI-driven anomaly detection
  • Strong platform engineering background including infrastructure as code, service mesh architectures, API gateways, and self-service developer platforms
  • Preferred: Master’s degree and 4 years' experience in software engineering or related area
  • Preferred: MLOps and model lifecycle management using MLflow, Kubeflow, or Seldon
  • Preferred: NLP and computer vision expertise
  • Preferred: Edge computing and distributed systems experience
  • Preferred: Kafka or Pulsar
  • Preferred: Chaos engineering, fault injection, and performance optimization
  • Preferred: Open-source contributions in reliability, observability, or infrastructure automation
  • Knowledge of WCAG 2.2 AA standards, assistive technologies, and digital accessibility best practices is preferred
Core Competencies

Demonstrates expertise in Reliability Engineering, AI/ML Engineering, and Platform Engineering, with a focus on architecting and implementing autonomous systems, intelligent automation solutions, and machine learning platforms. Proficient in cloud engineering and observability practices, ensuring mission-critical system reliability and performance optimization.

Highest-signal resume keywords
  • Reliability Engineering
  • AI/ML Engineering
  • Cloud Engineering
  • MLOps
  • Multi-Agent Frameworks
ATS Optimization Keywords
Hard Skills
  • Machine Learning
  • TensorFlow
  • PyTorch
  • Kubernetes
  • Docker
  • CI/CD Pipelines
  • Infrastructure as Code
  • Predictive Analytics
  • Anomaly Detection
  • Change Management
Soft Skills
  • Technical Mentorship
  • Thought Leadership
Industry Keywords
  • Incident Management
  • Problem Management
  • Change Management
  • Performance Engineering
  • Capacity Engineering
Tools & Technologies
  • Azure
  • GCP
  • AWS
  • MLflow
  • Kubeflow
  • Seldon
  • Kafka
  • Pulsar
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Staff Engineer – DevOps
Senior Staff Engineer – DevOps

Jobtailor • California (MO)

On-site
USD 140,000 - 180,000
Principal Engineer – Public Cloud Data
Principal Engineer – Public Cloud Data

Jobtailor • Arizona

On-site
USD 150,000 - 185,000
Senior Software Engineer – Application Reliability, Kubernetes, GCP, SQL
Senior Software Engineer – Application Reliability, Kubernetes, GCP, SQL

Jobtailor • California (MO)

On-site
USD 150,000 - 190,000
Principal Software Engineer – Backend
Principal Software Engineer – Backend

Jobtailor • Bentonville (AR)

On-site
USD 140,000 - 170,000
Lead Software Engineer
Lead Software Engineer

Jobtailor • Colorado

On-site
USD 150,000 - 190,000
Distinguished Engineer
Distinguished Engineer

Jobtailor • Atlanta (OH)

On-site
USD 180,000 - 280,000
Distinguished Engineer
Distinguished Engineer

Jobtailor • Kentucky

On-site
USD 180,000 - 240,000
Senior Manager, AI Reliability Engineering -Kroger Technology & Digital (P2498)
Senior Manager, AI Reliability Engineering -Kroger Technology & Digital (P2498)

Socket.dev • Cincinnati (OH)

On-site
USD 180,000 - 240,000
AI & Data Platform Engineering Manager
AI & Data Platform Engineering Manager

Jobtailor • Colorado

Hybrid
USD 180,000 - 240,000
Senior Engineering Team Lead – Infosec
Senior Engineering Team Lead – Infosec

Jobtailor • Northern (KY)

Hybrid
USD 140,000 - 200,000