Senior Software Engineer (Distributed Systems)

Workday

Santa Clara (CA)

Hybrid

USD 150,000 - 210,000

Full time

39 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Workday is hiring for a Senior or Mid-Level Software Engineer on the Data Platform and Observability team. You will design and scale critical observability services, including monitoring, logging, alerting, and tracing, across multi-cloud environments.

In a hybrid setup, you’ll collaborate with global teams and contribute to AI-powered workflows that reduce incidents and improve performance. We handle billions of messages daily and billions of time-series, building a trustworthy platform with

Qualifications

  • 7 years of coding in Python, Go, or Java
  • Experience with Linux
  • 8 years of software engineering/design experience
  • BS in Computer Science or a related field, or equivalent experience
  • MS Degree a plus
  • Ability to design, maintain, and optimize Time-Series DBs, document stores, and log collection/management systems
  • Hands-on experience with LLM orchestration frameworks and agentic AI build patterns

Responsibilities

  • Design, build, and improve critical observability services: Monitoring, Logging, Alerting, and Tracing
  • Own distributed tracing end to end instrumenting services
  • Build data capture and collection services using modern tech across environments
  • Design and develop core software modules for real-time and batch data processing
  • Build metrics ingestion pipelines, alert definitions, and automation for dashboard lifecycle management
  • Instrument AI-powered workflows for performance, cost, and quality
  • Define SLOs/SLIs and production-readiness criteria for new services
  • Build infrastructure components and deploy them in production
  • Work across all aspects of observability with focus on data quality and availability
  • Evaluate and implement new open-source and cloud-native tools

Skills

Python
Go
Java
Linux
Distributed systems
Communication

Education

BS in Computer Science or related field
MS Degree a plus

Tools

LangChain
LlamaIndex
Semantic Kernel
Docker
Kubernetes
Prometheus
Terraform
Ansible
Chef

Job description

  • The Data Platform and Observability team is based in Pleasanton, CA; Boston, MA; and Dublin, Ireland
  • We enable real-time insight across Workday's platforms, infrastructure, and applications including the AI agents and automated workflows that are becoming part of how those applications run
  • Our focus is building a large-scale distributed data and observability platform that keeps critical Workday services, and the AI-powered capabilities layered on top of them, fast, reliable, and trustworthy in production
  • We handle hundreds of terabytes of data and billions of messages produced daily across Workday's applications and underlying services, powering a platform that tracks over 2 billion time-series in production
  • We're also modernizing how we do observability itself bringing AI-driven approaches like anomaly detection, intelligent alerting, and automated root-cause analysis into the platform, so our systems get smarter about surfacing problems before they become incidents
  • If you enjoy writing efficient software, tuning and scaling large distributed systems, and applying AI to make observability itself sharper, you'll enjoy working with us
  • Do you want to solve interesting challenges at massive scale, across private and public cloud, for 4,000+ global customers while helping define what "observability" means for the next generation of AI-driven products?
  • Do you want to work alongside world-class engineers building the platforms that make that possible?
  • The Data Platform and Observability team is hiring a Senior or Mid-Level Software Engineer. We have a hybrid schedule you'll collaborate with Workmates in the office while having the flexibility to work up to 50% remote
  • Design, build, and improve critical observability services: Monitoring, Logging, Alerting, and Tracing
  • Own distributed tracing end to end instrumenting services, propagating context across service boundaries, and using trace data to understand system behavior and diagnose issues in complex, multi-hop request paths
  • Build data capture and collection services using the latest technologies across multiple infrastructure types (Kubernetes, Docker, OpenStack, bare metal, etc.)
  • Design and develop core software modules for real-time and batch data processing
  • Build metrics ingestion pipelines, alert definitions, and automation for dashboard lifecycle management
  • Instrument AI-powered and automated workflows for performance, cost, and quality, including multi-step processes and tool/service orchestration
  • Explore and apply AI-driven observability techniques anomaly detection, intelligent alerting, and automated root-cause analysis to reduce noise and speed up incident response
  • Partner with product and application teams to define SLOs/SLIs and production-readiness criteria for new services and automated workflows
  • Build infrastructure components and deploy them in production
  • Work across all aspects of observability with a keen eye for data quality, data integrity, and data availability
  • Evaluate and implement new open-source and cloud-native tools and technologies as needed
  • Participate in the on-call rotation supporting the observability platform
  • 7 years of coding in Python, Go, or Java
  • Experience with Linux
  • 8 years of software engineering/design experience
  • BS in Computer Science or a related technical field, or equivalent experience
  • MS Degree a plus
  • Ability to design, maintain, and optimize Time-Series DBs, document stores, and log collection/management systems
  • Hands-on experience with LLM orchestration frameworks (e.g., LangChain, LlamaIndex, Semantic Kernel) and a working knowledge of agentic AI build patterns how agents plan, call tools, hold state, and hand off work across multi-step chains
  • Excellent interpersonal, technical, and communication skills
  • Solid understanding of distributed tracing concepts (context propagation, spans, sampling) and hands-on experience with tracing tooling
  • Experience with service mesh, Prometheus, and cloud-native technologies
  • Ability to prioritize multiple tasks in a fast-paced environment
  • Experience with containerization and infrastructure automation (Docker, Kubernetes, Ansible, Chef, Terraform)
  • A knack for spotting where AI systems quietly go wrong in production: a model call that's suddenly slower than usual, a token bill creeping up for no clear reason, or answers that subtly drift in quality over time. You know how to build the dashboards, alerts, and instrumentation that catch these early turning "the AI feels off" into a measurable, debuggable signal
  • Public cloud experience (AWS/GCP), including native observability tooling such as AWS CloudWatch and GCP Cloud Operations (Stackdriver)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Distributed Systems Engineer - AI Observability
Senior Distributed Systems Engineer - AI Observability

Workday • Santa Clara (CA)

Hybrid
USD 190,000 - 285,000
Sr. Software Engineer - Distributed Systems
Sr. Software Engineer - Distributed Systems

Workday • Boulder (CO)

Hybrid
USD 169,000 - 253,000
Sr. Software Engineer - Distributed Systems
Sr. Software Engineer - Distributed Systems

Workday • Santa Clara (CA)

Hybrid
USD 190,000 - 285,000
Senior Software Development Engineer
Senior Software Development Engineer

Workday • Boulder (CO)

On-site
USD 150,000 - 200,000
Senior Distributed Systems Engineer AI-Driven Observability
Senior Distributed Systems Engineer AI-Driven Observability

Workday • Boulder (CO)

Hybrid
USD 169,000 - 253,000
Senior Software Engineer, AI-Driven Observability
Senior Software Engineer, AI-Driven Observability

HR Tech Job • Santa Clara (CA)

Hybrid
USD 190,000 - 285,000
Senior Software Engineer, Distributed Systems & AI
Senior Software Engineer, Distributed Systems & AI

Workday • Santa Clara (CA)

Hybrid
USD 150,000 - 210,000
Principal Distributed Systems Engineer - Observability
Principal Distributed Systems Engineer - Observability

Workday, Inc. • Pleasanton (CA)

On-site
USD 223,000 - 334,000
Flexible work options
Sr. Software Engineer - Distributed Systems
Sr. Software Engineer - Distributed Systems

HR Tech Job • Santa Clara (CA)

On-site
USD 190,000 - 285,000
Senior Software Engineer, Agentic AI and Observability
Senior Software Engineer, Agentic AI and Observability

NVIDIA AI • Santa Clara (CA)

On-site
USD 190,000 - 240,000
Equity
Benefits