Stand out for this role — generate a tailored resume and cover letter in about a minute.
NVIDIA AI in Santa Clara, CA, seeks a senior software engineer to build observability for agentic AI applications in production. You will collaborate with AI teams to define SLOs/SLIs and lead incident response and reliability improvements for BizApps AI services.
Proficiency in Python, Kubernetes, Docker, Datadog, OpenTelemetry, Grafana, Prometheus, and CI/CD is required. Experience with LangChain, LlamaIndex, Semantic Kernel, Terraform, Pulumi, and GitOps is a plus.
Build and implement observability solutions, including metrics and tracing, for agentic AI applications in production. Partner with AI teams to define SLOs/SLIs and drive incident response and reliability improvements for BizApps AI services.
Requirements: Requires a BS or MS in Computer Science or a related field with over 8 years of software engineering experience in Python and modern CI/CD practices. Must have hands-on experience with container orchestration and observability platforms like Datadog or Prometheus.
Key Skills: Python, Kubernetes, Docker, Datadog, OpenTelemetry, Grafana, Prometheus, CI/CD, LangChain, LlamaIndex, Semantic Kernel, Terraform, Pulumi, GitOps, SRE, Distributed Systems
Benefits: Equity, Benefits