Intermediate/ Senior Software Engineer - Cortex LLM Training Platform

Snowflake

Bellevue (WA)

On-site

USD 200,000 - 288,000

Full time

37 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Medical, dental, vision
401(k) retirement plan
Paid holidays and time off

Job summary

Snowflake is seeking an experienced Senior Software Engineer for Cortex Training to scale their ML platform. You will design the full stack from APIs to GPU data plane, ensuring fast, reliable training and inference at enterprise scale.

You will work with DeepSpeed, PyTorch, and GPU infrastructure, partnering with researchers to turn state-of-the-art methods into production-ready components. A strong distributed-systems background is essential.

Qualifications

  • BS in Computer Science or related field; MS/PhD a plus.
  • 3+ years building production ML systems; 6+ years for senior level.
  • Strong background in distributed systems and GPU infrastructure.

Responsibilities

  • Design and build across the full ML stack from public APIs to GPU data plane.
  • Scale distributed GPU compute with multi-tenant scheduling and fault tolerance.
  • Optimize training, inference and RL loops for high concurrency and throughput.
  • Productionize research building blocks with reliable, composable components.

Skills

Distributed systems
Kubernetes
PyTorch
DeepSpeed
Ray

Education

BS in Computer Science

Tools

CUDA/NCCL
vLLM
DeepSpeed/FSDP

Job description

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done.

Senior Software Engineer — Cortex Training

The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput.

The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it.

YOU WILL:
  • Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane.
  • Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional GPU pools, with fault tolerance built in.
  • Drive end-to-end performance at scale — keep the training, inference, and RL loops fast and the data plane responsive under heavy concurrent load, with GPUs kept saturated.
  • Productionize research building blocks — partner with Snowflake Research to turn state-of-the-art training and inference techniques into reliable, composable components customers can run at enterprise scale.
QUALIFICATIONS:
  • 3 + years (Intermediate) | 6+ years (Senior) building and shipping production ML systems
  • Strong distributed systems and infrastructure foundation — designing scalable, fault-tolerant services and operating them on Kubernetes in production.
  • Familiarity with GPU and LLM infrastructure — e.g., PyTorch, DeepSpeed/FSDP, Ray, CUDA/NCCL, vLLM; able to debug across the data, infrastructure, and GPU layers.
  • Demonstrated ability to harden complex systems for reliability, throughput, and cost efficiency.
  • BS in Computer Science or a related field (MS/PhD a plus).
  • (Bonus) Hands-on LLM post-training / modeling experience — the strongest candidates pair deep infra skills with real post-training intuition.

Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.

How do you want to make your impact?

For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com

The following represents the expected range of compensation for this role:

  • The estimated base salary range for this role is $200,000 - $287,500.
  • Additionally, this role is eligible to participate in Snowflake’s bonus and equity plan.

The successful candidate’s starting salary will be determined based on permissible, non-discriminatory factors such as skills, experience, and geographic location. This role is also eligible for a competitive benefits package that includes: medical, dental, vision, life, and disability insurance; 401(k) retirement plan; flexible spending & health savings account; at least 12 paid holidays; paid time off; parental leave; employee assistance program; and other company benefits.

To comply with pay transparency requirements and other statutes, you can notify us if you believe that a job posting is not compliant by completing this form.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff/Principal AI Software Engineer - Snowflake CoWork
Staff/Principal AI Software Engineer - Snowflake CoWork

Snowflake • Menlo Park (CA)

On-site
USD 264,000 - 380,000
Medical, Dental, Vision insurance
401(k) retirement plan
Equity plan
+1
Senior/Staff Software Engineer – LLM Inference & Reinforcement Learning Platform
Senior/Staff Software Engineer – LLM Inference & Reinforcement Learning Platform

Snowflake • Bellevue (WA)

On-site
USD 236,000 - 310,000
Medical, dental, vision, life, and‑dis
Disability insurance
401(k) retirement plan
+4
Senior Software Engineer - Cortex AI - FDE
Senior Software Engineer - Cortex AI - FDE

Snowflake • Menlo Park (CA)

On-site
USD 200,000 - 270,000
Medical Insurance
Dental Insurance
Vision Insurance
+9
Senior Software Engineer, Cortex Quality
Senior Software Engineer, Cortex Quality

Snowflake • Menlo Park (CA)

On-site
USD 200,000 - 270,000
Senior Forward Deployed Engineer, Applied AI
Senior Forward Deployed Engineer, Applied AI

Snowflake • Menlo Park (CA)

On-site
USD 200,000 - 263,000
Sr Manager, Applied Field Engineering - AI/ML
Sr Manager, Applied Field Engineering - AI/ML

Snowflake • Menlo Park (CA)

On-site
USD 207,000 - 272,000
Sr Manager, Applied Field Engineering - AI/ML
Sr Manager, Applied Field Engineering - AI/ML

Snowflake • New York (NY)

On-site
USD 207,000 - 272,000
Medical insurance
Dental insurance
Vision insurance
+4
Principal Technical Architect - AI/ML, Advanced Services
Principal Technical Architect - AI/ML, Advanced Services

Snowflake • New York (NY)

On-site
USD 196,000 - 257,000
Medical insurance
401(k) retirement plan
Paid holidays
Senior Software Developer - Full Stack
Senior Software Developer - Full Stack

Snowflake • Dublin (OH)

On-site
USD 140,000 - 210,000
Senior Forward Deployed Engineer, Applied AI
Senior Forward Deployed Engineer, Applied AI

Snowflake • Bellevue (WA)

On-site
USD 200,000 - 263,000