AI Engineer - US

ScaleOps

United States

On-site

USD 140,000 - 190,000

Full time

17 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

ScaleOps is redefining autonomous cloud and AI infrastructure. We are forming a new AI Group and seeking a Senior AI Engineer to shape it from an early stage—a greenfield, long-term effort to evolve decisions across the ScaleOps platform using AI-driven systems.

You'll build the AI brain that informs and drives our core automation engine with real production impact from day one, with room to grow into technical leadership as the group scales.

Qualifications

  • 4+ years of software engineering experience with Python and backend fundamentals.
  • Experience building and operating production systems in cloud environments.
  • Practical GenAI deployment in production, with latency, cost control, and failure modes.
  • Strong ownership and ability to work independently while collaborating across teams.
  • Experience enabling LLM-driven systems with retrieval / vector databases is a plus.

Responsibilities

  • Design and build autonomous AI agents that analyze infrastructure in real time and make decisions safely.
  • Develop integrations to expose AI capabilities to agents across tools like Slack and Jira.
  • Own ML pipeline from training to production, ensuring reliability and monitoring.
  • Define AI architecture and best practices as a founding member of the AI team.
  • Lead technical decisions on frameworks, multi-agent systems, and data governance.
  • Build internal AI tools to accelerate engineering, development, and support.

Skills

Python
Backend engineering
Cloud environments
GenAI
AI systems design
Ownership
Technical leadership

Tools

LangGraph
PydanticAI
LangChain
MetaGPT
Slack
Jira
Cursor
Windsurf

Job description

ScaleOps is redefining autonomous cloud and AI infrastructure. We're on a mission to free DevOps and platform engineers from manual resource management so they can focus on innovation, not tuning resources. The results: maximized performance and a reduction of cloud costs by up to 80%.

As the category leader in Autonomous Cloud and AI Infrastructure Resource Management, we're trusted by leading enterprises including Adobe, Wiz, Epic Games, Northwestern Mutual, Coinbase, DocuSign, and Fortune 100 companies to autonomously manage their most critical production environments.

Backed by Insight Partners, Lightspeed Venture Partners, and other leading VCs with over $210M in funding, ScaleOps is the leading player in a massive and growing market. We are building the autonomous infrastructure management platform that will power the next decade of enterprise compute.

What You'll Be Doing

We're forming a new AI Group and looking for a Senior AI Engineer to help shape it from an early stage - a greenfield, long-term effort to evolve how decisions are made across the ScaleOps platform using AI-driven systems. You won't just integrate APIs or build demos; you'll build the AI brain that works alongside (and increasingly drives) our core automation engine, with real production impact from day one and room to grow into technical leadership as the group scales.

  • Agentic AI Architecture: Design and build autonomous AI agents that analyze infrastructure in real time and make intelligent decisions. Work with modern agentic frameworks (LangGraph, PydanticAI) and conversational AI to create multi-agent systems - including troubleshooting, optimization, FinOps, and how-to agents. Leverage core LLM capabilities (tool-use, memory, retrieval) to operate safely in production.
  • Platform Integration & Intelligent Decision Systems: Develop MCPs to expose ScaleOps capabilities to AI agents that reason over infrastructure environments, metrics, configurations, and cost signals. Build integrations with tools like Slack, Jira, and AI-powered IDEs (Cursor, Windsurf) to deliver context-aware insights, from "why is this pod not scheduling?" to "how can we reduce costs by 30% safely?"
  • AI Model Development & MLOps: Build and deploy machine learning models that learn from infrastructure patterns - detecting the right resource policies for workloads, predicting optimal scaling triggers, and recommending GPU configurations. Own the complete ML pipeline from training to production, ensuring models are reliable, monitored, and continuously improving.
  • R&D AI Tools Development & Adoption: Build and embed internal AI tools to accelerate engineering, development, research, and support.
  • AI Tools for Business Impact: Develop AI-powered tools that help Sales and Support teams demonstrate value instantly - agents that analyze customer infrastructure, generate cost optimization reports automatically, and turn technical data into clear business recommendations.
  • End-to-End Ownership: Own AI systems from concept to production, ensuring they're fast (sub-2-second responses), reliable, safe, and cost-effective. Build evaluation frameworks to measure quality, implement security controls, and balance performance tradeoffs in production.
  • Technical Leadership: Define AI architecture and best practices as a founding member of the AI team. Make key technical decisions - choosing frameworks, designing multi-agent systems, establishing data governance - and shape how ScaleOps evolves from AI-enhanced internal tools to customer-facing AI products.
What You'll Bring
  • Core Engineering: Significant software engineering experience (typically 4+ years) with strong Python skills and solid backend engineering fundamentals.
  • Production Experience: Experience building and operating production systems in cloud environments.
  • Real-World GenAI Experience: Practical experience bringing LLM-based systems into production, including handling latency, cost control, and failure modes. Familiarity with additional agentic frameworks (e.g., LangChain, MetaGPT) and evaluation frameworks.
  • Builder Mentality: Strong ownership and the ability to operate independently while collaborating closely across teams, with the motivation to grow into technical leadership as the group expands.
  • (Advantage) Data & RAG: Experience enabling LLMs to consume structured or operational data (configurations, logs, metrics) and experience with retrieval systems (RAG) or vector databases.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead AI Engineer
Lead AI Engineer

Salt • New York (NY)

On-site
USD 180,000 - 230,000
AI Engineer
AI Engineer

Pinpoint Global Communications • United States

On-site
USD 120,000 - 180,000
Senior AI Engineer (Fully Remote)
Senior AI Engineer (Fully Remote)

ScaleOps • United States

Remote
USD 140,000 - 190,000
Staff Engineer / AI Builder
Staff Engineer / AI Builder

Robots & Pencils • California

On-site
USD 124,000 - 171,000
Lead Engineer� Data Platforms, Performance & Agentic AI
Lead Engineer� Data Platforms, Performance & Agentic AI

Accylerate, LLC. • United States

On-site
USD 180,000 - 240,000
Senior AI Engineer
Senior AI Engineer

Harnham • San Francisco (CA)

On-site
USD 180,000 - 240,000
Automation AI Engineer
Automation AI Engineer

GigaBrands • United States

On-site
MXN 600,000 - 1,200,000
Competitive salary
High-impact role
Scale AI systems
Senior Agentic AI Engineer
Senior Agentic AI Engineer

ImagineX LLC • Northern (KY)

Hybrid
USD 140,000 - 190,000
Senior Agentic AI Engineer
Senior Agentic AI Engineer

Build • Northern (KY)

Hybrid
USD 150,000 - 190,000
AI/ML Engineer
AI/ML Engineer

Harnham • San Francisco (CA)

On-site
USD 180,000 - 280,000