Principal/Senior ML Platform Engineer (Cortex)

Jobs Paloaltonetworks

United Kingdom

On-site

GBP 120,000 - 180,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Palo Alto Networks is seeking a Principal MLOps Engineer to join the Data & AI group in Cortex Research. You will design and scale the MLOps and LLMOps platforms powering data scientists and security researchers, building high‑performance infrastructure to train, fine‑tune, and deploy advanced AI systems across SLMs and RAG workflows.

You will own the serving architecture for LLMs/SLMs, optimize GPU utilization, and implement robust monitoring and automated training loops.

Qualifications

  • 4+ years as Senior ML Engineer, MLOps Engineer, or Backend Platform Engineer (Hands‑On) in cloud environments.
  • Experience managing lifecycle of diverse model architectures (ML, LLMs/SLMs, agentic/RAG) and scalable data pipelines.
  • Experience with distributed training across multi-GPU setups using PyTorch/DeepSpeed/Megatron-LM or cloud-native infra.
  • Strong cloud infra knowledge in GCP/AWS/Azure and managed AI platforms.

Responsibilities

  • Scale distributed training with multi-GPU workloads and optimized compute.
  • Automate the ML lifecycle with end-to-end pipelines for CT/CD.
  • Own serving architecture for LLMs/SLMs, balancing latency and throughput.
  • Implement comprehensive observability for live model performance and data drift.
  • Collaborate with data scientists and security researchers; integrate with core cloud infra.

Skills

Distributed Training
ML Platform
Python
CI/CD
Cloud Platforms
GPU Compute
Observability
Collaboration

Tools

PyTorch
DeepSpeed
Megatron-LM
Kubernetes
GitHub Actions
GitLab CI

Job description

Our Mission

At Palo Alto Networks®, we’re united by a shared mission—to protect our digital way of life. We thrive at the intersection of innovation and impact, solving real-world problems with cutting‑edge technology and bold thinking. Here, everyone has a voice, and every idea counts. If you’re ready to do the most meaningful work of your career alongside people who are just as passionate as you are, you’re in the right place.

Who We Are

In order to be the cybersecurity partner of choice, we must trailblaze the path and shape the future of our industry. This is something our employees work at each day and is defined by our values: Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact every individual can have. If you are passionate about solving real-world problems and ideating beside the best and the brightest, we invite you to join us! We believe collaboration thrives in person. That’s why most of our teams work from the office full time, with flexibility when it’s needed. This model supports real-time problem-solving, stronger relationships, and the kind of precision that drives great outcomes.

Job Summary

We are looking for a Principal MLOps Engineer with a deep focus on ML Platforms and Infrastructure to join our Data & AI group at Cortex Research. Our team is responsible for designing, building, and scaling the foundational MLOps and LLMOps platforms that power both our Data Scientists and Security Researchers. You will architect the high‑performance core infrastructure that enables these roles to build, train, and deploy advanced AI systems—ranging from optimized Small Language Models (SLMs) to complex agentic workflows and RAG systems. If you are passionate about building scalable compute platforms and automating the full ML lifecycle to solve complex data and security challenges, we want to hear from you.

Key Responsibilities
  • Scale Distributed Training: Design and optimize infrastructure for training and fine‑tuning LLMs and SLMs, leveraging distributed GPU workloads, efficient clustering, and compute optimization.
  • Automate the ML Lifecycle: Architect robust, automated pipelines for continuous training (CT) and deployment (CD) of models, ensuring a seamless flow from raw data collection to production environments.
  • Build Model Infrastructure: Own the serving architecture for LLMs/SLMs, balancing latency, throughput, and GPU utilization under production traffic.
  • Implement Advanced Monitoring: Establish comprehensive observability systems to monitor live model performance, data drift, and computational metrics, feeding insights back into the automated training loops for continuous improvement.
  • Collaborative Architecture: Partner closely with data scientists and security researchers to productize complex model architectures and streamline their workflows, while collaborating with our DevOps team to integrate with core cloud infrastructure.
Qualifications
Required Qualifications
  • Core Engineering: 4+ years experience as a Senior ML Engineer, MLOps Engineer, or Backend Platform Engineer (Hands‑On) working with cloud environments.
  • Model Lifecycle Engineering: Hands‑on experience managing the technical lifecycle of diverse model architectures, spanning classic ML, LLMs/SLMs, and agentic/RAG systems. This includes engineering scalable data preparation and processing pipelines as well as implementing infrastructure for model training, fine‑tuning, optimization, and high‑throughput production serving.
  • Distributed Training & Compute: Strong foundational knowledge of Deep Learning concepts (neural network architectures, training dynamics, optimization techniques) paired with proven experience setting up and optimizing distributed training workloads across multiple GPUs (using PyTorch, DeepSpeed, Megatron‑LM, or cloud‑native training infrastructure).
  • Cloud & Infrastructure Architecture: Strong infrastructure knowledge within a major cloud provider ecosystem (GCP, AWS, or Azure), specifically leveraging managed AI platforms and services.
  • Python Expertise: Expert‑level Python skills focused on ML infrastructure, pipelines, and automation frameworks.
  • CI/CD Integration: Experience with modern CI/CD patterns (such as GitLab CI or GitHub Actions) for automating software and model delivery loops.
  • AI Tooling & Development: Proficient in leveraging day‑to‑day AI tools and ecosystems (e.g., Claude, Gemini, MCPs, custom skills, and markdown formatting) to generate, review, and test code dynamically within your development cycle.
Preferred Qualifications
  • Strong GCP ecosystem experience.
  • Background in data science or deep learning workflows.
  • Cybersecurity domain knowledge.
Our Commitment

We’re trailblazers that dream big, take risks, and challenge cybersecurity’s status quo. It’s simple: we can’t accomplish our mission without diverse teams innovating, together.

We’re committed to providing reasonable accommodations for all qualified individuals with a disability. If you require assistance or accommodation due to a disability or special need, please contact us at accommodations@paloaltonetworks.com.

Palo Alto Networks is an equal opportunity employer. We celebrate diversity in our workplace, and all qualified applicants will receive consideration for employment without regard to age, ancestry, color, family or medical care leave, gender identity or expression, genetic information, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran status, race, religion, sex (including pregnancy), sexual orientation or other legally protected characteristics.

All your information will be kept confidential according to EEO guidelines.

Is role eligible for Immigration Sponsorship? No. Please note that we will not sponsor applicants for work visas for this position.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal DevOps Engineer - AI Platforms
Principal DevOps Engineer - AI Platforms

Jobs Paloaltonetworks • United Kingdom

On-site
GBP 90,000 - 140,000
Senior Principal Software Engineer (Foundational Platform)
Senior Principal Software Engineer (Foundational Platform)

Jobs Paloaltonetworks • England

On-site
GBP 138,000 - 224,000
Lead MLOps Platform Architect
Lead MLOps Platform Architect

Jobs Paloaltonetworks • United Kingdom

On-site
GBP 120,000 - 180,000
ML Ops Engineer
ML Ops Engineer

Anaplan • Greater London

On-site
GBP 90,000 - 150,000
ML Ops Engineer
ML Ops Engineer

Anaplan Inc • Greater London

On-site
GBP 110,000 - 140,000
Platform Lead - ML Ops
Platform Lead - ML Ops

anaplan • Greater London

On-site
GBP 120,000 - 180,000
Platform Lead - ML Ops
Platform Lead - ML Ops

Anaplan Inc • Greater London

On-site
GBP 120,000 - 180,000
Ecosystems Strategic Analytics Manager
Ecosystems Strategic Analytics Manager

Jobs Paloaltonetworks • United Kingdom

On-site
GBP 93,000 - 150,000
Principal Software Engineer for SaaS (SIA for Databases)- Idira
Principal Software Engineer for SaaS (SIA for Databases)- Idira

Jobs Paloaltonetworks • United Kingdom

On-site
GBP 120,000 - 190,000
Backend Team Manager -AppSec (Cortex Cloud)
Backend Team Manager -AppSec (Cortex Cloud)

Jobs Paloaltonetworks • United Kingdom

On-site
GBP 110,000 - 170,000