AI Infrastructure / MLOps Engineer — NYC

LaStellar Group

New York (NY)

On-site

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

LaStellar Group is seeking a Platform Engineer to own and operate our production AI and data platform infrastructure across Azure and GCP. You will manage cloud infra, CI/CD pipelines, container orchestration, and agent operations to keep systems reliable, scalable, and production-grade.

You will build observability tooling, support data platform integration, and implement auto-scaling for container workloads. The role demands 3–5 years in DevOps/MLOps with strong curiosity about AI systems.

Qualifications

  • 3–5 years in software engineering, DevOps, MLOps, or platform engineering with production ownership.
  • Hands‑on Docker and container orchestration experience in Azure and/or GCP.
  • Terraform across cloud providers — you've designed it, not just configured it.
  • CI/CD pipeline experience with Git-based release management.
  • Systems thinker — you troubleshoot end‑to‑end, not just at the surface.
  • Genuine curiosity about AI and agentic systems — excited to grow into deeper platform concepts.
  • Strong signals in policy-as-code, observability depth, and multi-cloud tooling are a plus.

Responsibilities

  • Operate and scale live agentic AI systems across Azure and GCP, ensuring availability, performance, and resilience.
  • Build and maintain observability tooling for agent execution—logging, tracing, alerting, and performance monitoring.
  • Support integration of agents with data platforms and MCP servers.
  • Implement auto-scaling strategies for containerized workloads across Azure Container Apps, Cloud Run, and GKE.
  • Contribute to evaluation frameworks and quality standards for AI agents in production.
  • Own Python-based services’ lifecycle—from containerization to deployment and runtime behavior.
  • Build shared tooling and internal packages to accelerate data science workflows.
  • Write and maintain Terraform across Azure and GCP for registries, identities, secrets, and networks.
  • Develop CI/CD pipelines and release workflows across data science and engineering repos, enforcing security and runbooks.

Skills

Docker
Container orchestration
CI/CD pipelines
Infrastructure as code
End-to-end troubleshooting
AI/agented systems curiosity

Tools

Azure
GCP
Azure Container Apps
GKE
GCP Cloud Run
Vertex AI
Managed Identities
VNets
Key Vault
Secret Manager
OpenTelemetry
LangChain
MLflow
Prefect
dbt
Snowflake

Job description

We are a fast-growing fintech and investment platform operating at the intersection of AI and financial markets. Our production AI systems are live, our data platform is scaling fast, and we need someone to help us build the infrastructure that keeps it all running reliably.

This is not a research or modeling role. You will own the systems — cloud infrastructure, CI/CD pipelines, container orchestration, agent operations, observability, and enforcement — that make our AI and data platforms reliable, scalable, and production-grade.

What You’ll Own

AI Platform & Agent Operations

  • Operate and scale live agentic AI systems across Azure and GCP — ensuring agents are highly available, performant, and resilient under load.
  • Build and maintain observability tooling for agent execution — logging, tracing, alerting, and performance monitoring.
  • Support integration of agents with data platforms and Model Context Protocol (MCP) servers.
  • Implement auto-scaling strategies for containerized agent workloads across Azure Container Apps, GCP Cloud Run, and GKE.
  • Contribute to evaluation frameworks and quality standards for AI agents in production.
MLOps & Python Engineering
  • Write, maintain, and improve production Python code powering data pipelines, agent workflows, and platform tooling.
  • Own the full lifecycle of Python-based services — containerization, deployment, versioning, and runtime behavior.
  • Deploy and operate workflow orchestration using Prefect — scheduling, error handling, retry logic, and human-in-the-loop patterns.
  • Build shared Python tooling and internal packages that enable data science teams to develop and deploy faster.
  • Write and maintain Terraform across Azure and GCP — container registries, managed identities, Key Vault, GCP Secret Manager, storage backends, and VNet configurations.
  • Build and maintain CI/CD pipelines and release management workflows across data science and engineering repositories.
  • Enforce coding standards, security policies, and compliance controls directly in the pipeline.
  • Ensure all production systems are well-documented with clear runbooks and data lineage.
Observability & Reliability
  • Build and own the observability stack — metrics, logging, distributed tracing, alerting.
  • Drive SLO/SLI frameworks and incident response as the platform matures.
  • Troubleshoot production issues end-to-end — from application logic through to infrastructure.
What We're Looking For
Required
  • 3–5 years in software engineering, DevOps, MLOps, or platform engineering with clear production system ownership.
  • Hands‑on Docker and container orchestration experience in Azure and/or GCP.
  • Terraform across cloud providers — you've designed it, not just configured it.
  • CI/CD pipeline experience with Git-based release management.
  • Systems thinker — you troubleshoot end‑to‑end, not just at the surface.
  • Genuine curiosity about AI and agentic systems — excited to grow into deeper platform concepts.
Strong Signal
  • Policy-as-code enforcement — OPA or equivalent in a production CI/CD context. Tell us about it.
  • Observability depth — if you describe OpenTelemetry as a specification and protocol rather than a tool, we want to talk.
  • Experience with Azure Container Apps, ACI, ACR, Managed Identities, VNets.
  • Experience with GCP Cloud Run, GKE, Vertex AI, IAM, Secret Manager.
  • Familiarity with agentic frameworks — MCP, LangChain, or similar.
  • AI observability platforms — Langfuse, MLflow, or similar.
  • dbt, Snowflake, or similar data transformation and warehousing tools.
Nice to Have
  • Prefect or similar workflow orchestration in production.
  • Multi-cloud networking and identity management experience.
  • Financial services or fintech domain exposure.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

MLOps Engineer
MLOps Engineer

Compunnel, Inc. • San Antonio (TX)

On-site
USD 100,000 - 130,000
MLOps Engineer: Scalable ML Pipelines & Infra
MLOps Engineer: Scalable ML Pipelines & Infra

Compunnel, Inc. • San Antonio (TX)

On-site
AI Engineer, AIOps & Infrastructure
AI Engineer, AIOps & Infrastructure

eloquentai • San Francisco (CA)

On-site
USD 130,000 - 160,000
AI Infrastructure Engineer MLOps
AI Infrastructure Engineer MLOps

EITACIES Inc. • San Francisco (CA)

On-site
USD 120,000 - 150,000
401(k)
AIOps Engineer
AIOps Engineer

Compunnel, Inc. • Reston (VA)

On-site
USD 120,000 - 150,000
MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
AI Engineer
AI Engineer

Teserac, Inc. • Sunnyvale (CA)

On-site
USD 100,000 - 130,000
Health Care Plan (Medical, Dental & Vision)
Paid Time Off (Vacation, Sick & Public Holidays)
Free Food & Snacks
+2
ML-Ops / Platform Engineer
ML-Ops / Platform Engineer

Veriipro • Charlotte (NC)

On-site
USD 120,000 - 180,000
ML Ops Engineer — Agentic AI Lab (Founding Team)
ML Ops Engineer — Agentic AI Lab (Founding Team)

Fabrion • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Meaningful equity
MLOps Technical Lead
MLOps Technical Lead

TechDigital Group • Austin (TX)

On-site
USD 100,000 - 130,000