Engagement summary: Join a small engineering team building custom internal MCP (Model Context Protocol) servers in Python on Kubernetes, connecting enterprise systems such as Salesforce, JIRA, and Snowflake to internal AI/LLM applications. The shared framework, CI/CD, and platform infrastructure already exist; this role focuses on designing and implementing production-grade integration servers on top of them.
Responsibilities
- Design and implement MCP servers in Python using the MCP SDK (e.g., FastMCP) on the existing internal framework and scaffolding template.
- Author per-server specifications: define a focused v1 tool set (5–10 tools), identity and auth model, rate-limit/quota strategy, and pagination/truncation rules, and drive stakeholder sign-off before build
- Build robust async API clients (httpx or similar) for enterprise SaaS and data platforms, with retries, backoff, circuit breakers, and clean mapping of downstream errors (401/403/429, entitlement failures) into meaningful MCP responsesß
- Define tool interfaces with Pydantic schemas and model-friendly descriptions; iterate based on LLM evaluation runs to eliminate ambiguous or model-hostile tool designs
- Integrate with the organization's OAuth abstraction layer for per-user and service-account identity flows, and with secrets management for credential handling — never handling raw secrets in code or logs
- Implement caching, internal rate limiting, and quota-protection logic for metered vendor APIs
- Write unit and integration tests (including sandbox-tenant testing) and participate in LLM eval passes as part of the standard definition of done
- Support security reviews: least-privilege scope design, audit logging (user → tool → resource), guardrails on write/destructive actions, and prompt-injection-aware handling of untrusted data returned from downstream systems
- Deploy through the existing Kubernetes pipeline (Helm/Kustomize), configure dashboards and alerts (error rates, latency, quota burn), and support two-week pilot rollouts per server
- Document each server in the internal catalog and contribute improvements back to the shared framework and scaffold template
Required qualifications
- 5+ years of professional backend development experience, with 3+ years in Python (async programming, typing, Pydantic)
- Strong track record integrating third-party REST APIs at production scale: authentication flows, pagination, rate limiting, webhooks, and error-handling patterns
- Working knowledge of OAuth 2.0 concepts (authorization code flow, token refresh, scopes) sufficient to consume an internal auth abstraction correctly
- Hands-on experience deploying and operating services on Kubernetes (containers, Helm or Kustomize, health probes, HPA basics) and working within established CI/CD pipelines
- Solid testing discipline: pytest, mocking external APIs, integration testing against sandbox environments
- Experience with observability tooling (structured logging, Prometheus metrics, OpenTelemetry tracing or equivalent)
- Strong written communication — this role produces specs and catalog documentation, not just code
Preferred qualifications
- Direct experience with MCP, LLM tool/function calling, or building AI agent integrations
- Prior integration work with one or more target systems (Salesforce APIs, Atlassian/JIRA, Snowflake, ServiceNow, Workday, or similar enterprise platforms)
- Familiarity with secrets management (Vault, External Secrets Operator) and enterprise security review processes
- Experience with Redis or similar for caching strategies against metered APIs
- Exposure to prompt-injection risks and secure design patterns for LLM-facing services
Success measures (first 90 days)
- 3–5 servers shipped to production through all quality gates with pilot sign-off
- Zero security-review findings related to secret handling or scope over-provisioning
- LLM eval pass rates meeting team standards without repeated malformed tool calls