Principal AI/Ops Engineer

RxSense

United States

On-site

USD 160,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

RxSense is seeking a Principal AIOps Engineer to build the platform that makes AI cheaper, faster, safer, and observable. You will own the infrastructure powering all AI-powered RxSense products and directly report to the Director of AI Engineering.

This is a hands-on role from day one, partnering with AI engineers, software developers, data scientists, security, and finance to deliver deployment pipelines, agent runtimes, eval frameworks, and self-hosted model serving.

Qualifications

  • BS in Computer Science or a related technical field.
  • 6+ years in platform, infrastructure, or DevOps engineering, with at least 2 years building AI/ML or LLM-powered systems.
  • Hands-on Python or TypeScript coding experience.
  • Strong AWS background including IAM, networking, and container orchestration.
  • Proven experience designing and operating CI/CD pipelines for high-velocity teams.
  • Excellent written and verbal communication skills.

Responsibilities

  • Build and maintain end-to-end deployment pipelines for AI-powered applications.
  • Stand up and operate runtime infrastructure for production agents and manage deployment contracts.
  • Own provisioning, rotation, and metering of access to model provider APIs; enforce quotas and cost visibility.
  • Integrate eval frameworks into CI/CD and enable teams to author their own evals.
  • Stand up self-hosted inference for latency-sensitive workloads and manage serving stack and GPU economics.
  • Develop shared developer harness for prompt management, model routing, retries, tracing, and policy enforcement.
  • Collaborate with finance on token accounting and real-time spend observability.
  • Write documentation and runbooks; promote best practices and reusable interfaces.

Skills

Python
TypeScript
AWS
CI/CD
DevOps
Distributed systems
Communication
Problem solving

Education

BS in Computer Science or related technical field

Tools

Kubernetes
Docker
Terraform

Job description

Position Summary:

The Principal AIOps Engineer will build the platform that makes AI cheap, fast, safe, and observable at RxSense. As a direct report to the Director of AI Engineering, this role will own the infrastructure that every AI-powered product at RxSense depends on. This is a hands-on-keyboard position from day one, partnering with AI engineers, software engineers, data scientists, security, and finance to deliver deployment pipelines, agent runtime, eval frameworks, self-hosted model serving, and the developer harness that determines how fast every other engineer in the company can ship.

Essential Duties and Responsibilities:
  • Build and maintain end-to-end deployment pipelines for AI-powered applications, including artifact builds, environment promotion, rollback, and observability hooks. Drive new greenfield deployment platforms from initial build to the default that AI teams ship on.
  • Stand up and operate the runtime and lifecycle infrastructure for production agents, including deployment, versioning, monitoring, rate-limiting, and retirement. Define the deployment contract (config, secrets, tools, memory, evals) and the operational SLOs.
  • Own how the organization provisions, rotates, scopes, and meters access to model provider APIs (Anthropic, OpenAI, and others). Build a key management layer that enforces per-team and per-app quotas, prevents leakage, and gives finance and engineering a clear view of spend.
  • Build evals into the CI/CD pipeline so no agent or LLM-powered service ships without passing a defined eval bar. Design the framework so product teams can author their own evals against a shared harness, and so eval results gate promotion across environments.
  • Stand up self-hosted inference for workloads where managed APIs aren't the right fit, including latency-sensitive paths, regulated data, cost optimization, and vendor redundancy. Own the serving stack, the autoscaling and GPU economics behind it, and the playbook for when a workload belongs to a managed provider versus internal infrastructure.
  • Design and build the shared developer harness that every AI-powered service uses: prompt management, model routing, retries, tracing, eval hooks, and policy enforcement. Set the abstractions that determine how fast every other AI engineer can ship for the next three years.
  • Partner with finance on cost visibility, including token accounting, per-feature cost attribution, and real-time spend observability.
  • Write documentation, runbooks, and clear interfaces so the platform is adoptable by other engineering teams without hand-holding.
  • Participate in code review and promote collaboration and best practices including simplicity, automation, sound design patterns, test coverage, and reusability.
  • BS (or higher, e.g., MS or Ph.D.) in Computer Science or related technical field involving coding, or equivalent technical experience.
  • 6+ years of platform, infrastructure, or DevOps engineering, with at least 2 years building production infrastructure for AI/ML or LLM-powered systems. We care more about depth and drive than years on a resume.
  • Deep hands-on experience designing and operating CI/CD pipelines for high-velocity engineering organizations, including artifact management, environment promotion, and progressive rollout.
  • Strong AWS background, comfortable down to the IAM, networking, and container orchestration layers.
  • Proven track record building developer platforms or internal tools that other engineering teams adopted by choice, not by mandate.
  • Production experience with LLM-powered applications, including prompt management, model routing, retries, tracing, and the operational realities of running agents or chains in production.
  • Hands-on coding fluency in Python or TypeScript, ideally both. This is a keyboard role, not an architecture-only role.
  • Comfortable operating in a polyglot environment. The RxSense AI engineering stack spans Python, .NET, and TypeScript, and you will deploy and support services across all three.
  • Comfortable owning the cost and reliability conversation with both engineering leadership and finance partners.
  • Strong written communication and a bias toward documentation, runbooks, and clear interfaces.
  • Proven analytical thinking and problem-solving skills.
  • Excellent communication skills, both verbal and written.
Bonus Qualifications:
  • Direct experience integrating with Anthropic, OpenAI, or other frontier model provider APIs at scale, including key management, quota enforcement, and capacity planning.
  • Hands-on experience self-hosting models with vLLM, TGI, SGLang, or similar inference servers, including GPU autoscaling and cost optimization.
  • Built or contributed to an eval framework that gated production deployments.
  • Familiarity with agent runtimes and frameworks such as the Claude Agent SDK, LangGraph, or in-house equivalents.
  • Working familiarity with .NET, enough to read code, debug a deploy, and pair with service owners.
  • Background in healthcare, PBM, pharmacy, or another regulated data environment.
  • FinOps experience, particularly attributing AI spend to features or business units.
  • Kubernetes operator experience or comfort with custom controllers.
  • Experience with Agile development methodologies, preferably both Scrum and Kanban

RxSense believes that a diverse workforce is a more talented and productive workforce. As such, we are an Equal Opportunity and Affirmative Action employer. Our recruitment process is free from discriminatory hiring practices and all qualified applicants are considered for employment without regard to race, color, religion, sex, gender, sexual orientation, gender identity, ancestry, age, or national origin. Neither will qualified applicants be discriminated against on the basis of disability or protected veteran status. We believe in the strength of the collaboration, creativity and sense of community a diverse workforce brings.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Platform Engineer
Principal Platform Engineer

Rxsense • Boston (MA)

On-site
USD 190,000 - 225,000
Principal Platform Engineer, AI Engineering
Principal Platform Engineer, AI Engineering

RxSense • United States

Remote
USD 190,000 - 235,000
Lead Software Platform Engineer, MLOps
Lead Software Platform Engineer, MLOps

TetraScience • Cambridge (MA)

On-site
USD 200,000 - 270,000
Employer-paid benefits
Unlimited PTO
401K
+3
Lead Software Platform Engineer, MLOps
Lead Software Platform Engineer, MLOps

TetraScience, Inc. • Cambridge (MA)

Hybrid
USD 200,000 - 270,000
Employer-paid benefits
Unlimited PTO
401K
+3
Senior AI Engineer/AI Platform Developer
Senior AI Engineer/AI Platform Developer

American Screening Corporation • Shreveport (LA)

On-site
USD 140,000 - 220,000
Staff AI Engineer
Staff AI Engineer

Robots & Pencils • Houston (TX)

Hybrid
USD 177,000 - 210,000
Paid time off
Medical/dental/vision insurance
401(k)
Applied AI Engineer
Applied AI Engineer

Uncover • New York (NY), Northern (KY)

Hybrid
USD 140,000 - 210,000
Equity participation
High autonomy
Competitive salary
Applied AI Engineer
Applied AI Engineer

Norbert Health • New York (NY)

On-site
USD 100,000 - 150,000
Equity participation
Competitive salary
High autonomy and technical ownership
Applied AI Engineer
Applied AI Engineer

Uncover • New York (NY)

Hybrid
USD 140,000 - 200,000
Equity
Competitive salary
AI Integration Architect – AI Accelerator
AI Integration Architect – AI Accelerator

RTX • East Hartford (CT)

Hybrid
USD 157,000 - 299,000
medical insurance
dental insurance
vision insurance
+11