Get more replies from employers
Send a job-specific resume in minutes.
Baseten in New York builds the internal AI development platform that enables teams to ship AI features quickly and safely. You will own the end-to-end platform, from architecture and rollout to operation and measurement, ensuring the defaults are easy to adopt and work well across our monorepo.
You will evaluate third-party AI coding tools, build agent frameworks, and create an evaluation practice that guides investment decisions.
Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products.
Baseten's engineers want to work in an AI-first way. What's missing isn't enthusiasm — it's the platform underneath it. Today everyone assembles their own agent config, context files, and MCP servers, so the good patterns stay trapped in individual setups instead of becoming defaults everyone inherits.
You'll build that platform: the agent configurations tuned to our monorepo, the context and tooling layer that makes agents competent in our codebase, the evals that tell us which approaches actually work, and the rollout mechanics that get a new engineer productive with agents in week one.
You are not here to mandate how engineers use AI — you're here to make the good path the easy path. Success looks like teams adopting what you build because it beats what they'd cobble together themselves, not because a policy requires it. Platform engineer, not AI evangelist. Ship infrastructure, measure it, kill what doesn't work, let adoption be the referee.
The playbook for AI-first SDLC doesn't exist at any company yet. You'll write ours.
Agent substrate — Repo-level context infrastructure that makes agents competent in our codebase (CLAUDE.md/AGENTS.md conventions, architecture and domain context, and the tooling to keep it accurate as code moves). Internal MCP servers giving agents scoped access to CI, observability, incident tooling, deployment state, and docs. Shared skills, subagents, and hooks that encode Baseten workflows. Sandboxed environments where agents can build and test safely.
The golden path — Project templates and onboarding that ship with AI tooling configured and working. Self-serve infrastructure so teams build their own agents without you as the bottleneck. Gateway, auth, cost controls, and audit logging for internal model access.
The feedback loop — Eval harnesses that answer 'is this config better than that one' against real Baseten tasks, not vibes. Instrumentation of AI tool usage and its downstream effects on cycle time, review latency, and change failure rate. Honest reporting, including on what you built that didn't pan out.
Agents in the SDLC — Automation where agents earn their keep: PR review triage, test gap-filling, incident context assembly, migrations and refactors, codebase Q&A. Integrating agents into CI/CD with guardrails that make it trustworthy.
At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.
We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).