Get more replies from employers
Send a job-specific resume in minutes.
Design, implement, and scale production‑grade observability for ML/LLM applications covering model performance, data quality/drift, safety compliance, cost, latency, reliability (SLOs), and user experience.
Evaluate, integrate, or build observability tooling for metrics, logs, and traces, including model telemetry.
Partner closely with ML engineers and platform teams to enable trustworthy AI systems in production.
Build telemetry for models including latency, token throughput, error rates, and SLOs for AI endpoints.
Create self‑service dashboards for stakeholders.
Salary Range: $100,000 - $120,000 a year.