Senior Software Development Engineer

Workday

Boulder (CO)

On-site

USD 150,000 - 200,000

Full time

39 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Workday in Boulder, CO is seeking a Senior Software Development Engineer for the AI Model Serving team. You will design, build, and scale systems that host and serve production ML models, ranging from traditional models to large language models powering Workday’s agents.

You will lead architecture decisions, mentor engineers, and collaborate with ML engineers and data scientists to ensure reliability, performance, and cost efficiency in a high-throughput environment.

Qualifications

  • 6+ years of experience building and operating large-scale distributed systems.
  • Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent).
  • Excellent written and verbal communication skills.

Responsibilities

  • Lead the AI Model Serving team and shape platform vision.
  • Design, implement, and maintain large-scale distributed systems for production ML models.
  • Write design documents and drive consensus for new components.
  • Review PRs and enforce performance, security, and readability.
  • Collaborate with engineers, ML engineers, data scientists, and partner teams.
  • Respond to alerts and troubleshoot production issues.

Skills

Python
Distributed systems
Kubernetes
GPU infra
Observability
Mentorship
Communication

Education

BS in CS/Engineering or related field

Tools

vLLM
TGI
LoRA

Job description

  • The AI Model Serving team is the engine behind every production Workday agent and machine learning use case. We own the services that power all production AI workloads, acting as both the gateway to vendor-hosted LLMs (GCP, AWS Bedrock, Gemini) and the primary platform where Workday hosts and scales its internal models
  • We operate at scale, hosting thousands of traditional ML models across sharded Ray Serve clusters and maintaining Workday’s production model registry. Our platform consistently handles ~2,000 requests per second, peaking at over 10,000 RPS in our largest clusters
  • In the year ahead, our engineering roadmap is highly ambitious. We are focused on:
  • Scaling Architecture: Upgrading our systems to seamlessly support 20+ new AI agents going into production
  • Hosting Open-Weight LLMs: Designing the infrastructure to host and tune open-source LLMs directly within our stack
  • Performance & Reliability: Architecting optimizations to drive down core latency while maintaining the rock-solid stability our high-throughput production systems demand
  • Enterprise Governance: Hardening our unified vendor interface and implementing advanced cost-governance controls
  • Our culture is built on focus, camaraderie, and high performance. We are a friendly, dedicated group that takes pride in building and operating one of the most heavily used services at Workday. If you are energized by working on the infrastructure that sits at the very heart of Workday’s AI strategy, this is the team for you
  • As either a Senior Software Development Engineer on the AI Model Serving team, you will be a technical leader who helps shape the vision and direction of the platform alongside the engineering manager. You will play a central role in making critical design decisions, driving outcomes across the team, and setting a positive and inclusive team culture
  • Your work will directly impact Workday’s ability to serve AI at scale — from traditional ML models to the latest large language models powering Workday’s agents
  • Lead the team technically by making critical design decisions that drive performance, reliability, and scalability across the platform
  • Design, implement, and maintain large-scale systems that enable moving ML models to production
  • Write design documents to build consensus for new system components and enhancements to existing components
  • Evaluate and uptake new technologies made available within Workday and across the broader industry
  • Troubleshoot, improve, and scale continuous integration software pipelines
  • Develop relationships with software engineers, machine learning engineers, and data scientists on partner teams
  • Respond to alerts and debug production issues to maintain platform health and reliability
  • Review pull requests and enforce consistency, performance, readability, and security across code bases
  • Develop documentation to share knowledge with other engineers

6+ years of related work experience in software development, with a focus on building and operating large-scale distributed systemsBachelor’s degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience)Communication: Excellent written and verbal communication skills, including the ability to write clear design documents, articulate complex technical ideas, and build consensus across teamsPython: Deep proficiency in Python, with extensive experience writing production-level code and building systems in Python-based frameworksSoftware Development and Distributed Systems: Deep experience designing, building, and scaling production-grade distributed systems. You understand the full software development lifecycle - from coding standards and testing to code reviews, source control, and deployment, and can apply that knowledge to complex, high-throughput platformsKubernetes & GPU Infrastructure: Deep hands-on experience deploying and scaling workloads on Kubernetes, with a specific focus on GPU resource management. You understand how to optimize GPU utilization for hosting and tuning smaller open-weight LLMs using modern inference engines (e.g., vLLM, TGI). Familiarity with GPU memory constraints, serving tuned models (e.g., LoRA), and autoscaling hardware metricsObservability: You can design and maintain monitoring strategies that provide clear insight into system health, performance, and costLLMs and Traditional ML Models: Familiarity with both large language models and traditional ML models, including how they are served, scaled, and monitored in production. You understand the operational differences and can design abstractions that serve both effectivelyMentorship: A collaborative approach to engineering, with experience mentoring other engineers and fostering an inclusive team environment

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr. Software Development Engineer
Sr. Software Development Engineer

Engg • Colorado

On-site
USD 180,000 - 230,000
Sr. Software Development Engineer
Sr. Software Development Engineer

Workday • United States

On-site
USD 150,000 - 210,000
Senior Software Development Engineer
Senior Software Development Engineer

Workday • United States

On-site
USD 180,000 - 240,000
Senior Consultant, AI/ML Engineer
Senior Consultant, AI/ML Engineer

Hollstadt Consulting • Minnesota

On-site
USD 150,000 - 210,000
Senior AI Platform Engineer — Production ML at Scale
Senior AI Platform Engineer — Production ML at Scale

Workday • United States

On-site
USD 150,000 - 210,000
Sr. Software Development Engineer
Sr. Software Development Engineer

Workday • Boulder (CO)

Hybrid
USD 156,000 - 234,000
Senior AI Systems Engineer – Scalable LLM Infra
Senior AI Systems Engineer – Scalable LLM Infra

Engg • Colorado

On-site
USD 180,000 - 230,000
Senior AI Platform Engineer – Scale ML & Open-Weight LLMs
Senior AI Platform Engineer – Scale ML & Open-Weight LLMs

Workday • Boulder (CO)

On-site
USD 150,000 - 200,000
Staff Applied AI Inference Engineer
Staff Applied AI Inference Engineer

Crusoe Energy Systems • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health benefits
401(k) match
Paid time off
+1
Full Stack Developer
Full Stack Developer

MiddleGround Capital • Lexington (KY)

On-site
USD 120,000 - 190,000