AI Infrastructure / MLOps Engineer — NYC

LaStellar Group

New York (NY)

On-site

USD 140,000 - 180,000

Full time

21 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

LaStellar Group, a fast-growing fintech and AI-enabled investment platform, seeks a Platform Engineer to own production infrastructure. You will manage cloud infrastructure, CI/CD pipelines, container orchestration, agent operations, observability, and policy enforcement to keep our AI and data platforms reliable at scale.

You will collaborate across data science and engineering teams, implement auto-scaling, build observability tooling, and maintain Terraform configurations across Azure and

Qualifications

  • 3–5 years in software engineering, DevOps, MLOps, or platform engineering with clear production system ownership.
  • Hands‑on Docker and container orchestration experience in Azure and/or GCP.
  • Terraform across cloud providers — you’ve designed it, not just configured.
  • CI/CD pipeline experience with Git‑based release management.

Responsibilities

  • Operate and scale live agentic AI systems across Azure and GCP — ensuring availability, performance, and resilience under load.
  • Build and maintain observability tooling for agent execution — logging, tracing, alerting, and performance monitoring.
  • Support integration of agents with data platforms and MCP servers.
  • Implement auto-scaling strategies for containerized agent workloads across Azure Container Apps, GCP Cloud Run, and GKE.
  • Write, maintain production Python code powering data pipelines, agent workflows, and platform tooling.
  • Build shared Python tooling and internal packages for data science teams to deploy faster.
  • Write and maintain Terraform across Azure and GCP — container registries, managed identities, Key Vault, Secret Manager, storage backends, and VNet configurations.
  • Build and maintain CI/CD pipelines and release management workflows across data science and engineering repositories.
  • Enforce coding standards, security policies, and compliance controls directly in the pipeline.
  • Ensure all production systems are well-documented with clear runbooks and data lineage.

Skills

Docker
Container orchestration
Terraform
CI/CD pipelines
End-to-end troubleshooting

Tools

OpenTelemetry
Prefect

Job description

We are a fast-growing fintech and investment platform operating at the intersection of AI and financial markets. Our production AI systems are live, our data platform is scaling fast, and we need someone to help us build the infrastructure that keeps it all running reliably.

This is not a research or modeling role. You will own the systems — cloud infrastructure, CI/CD pipelines, container orchestration, agent operations, observability, and enforcement — that make our AI and data platforms reliable, scalable, and production-grade.

What You’ll Own
AI Platform & Agent Operations
  • Operate and scale live agentic AI systems across Azure and GCP — ensuring agents are highly available, performant, and resilient under load.
  • Build and maintain observability tooling for agent execution — logging, tracing, alerting, and performance monitoring.
  • Support integration of agents with data platforms and Model Context Protocol (MCP) servers.
  • Implement auto-scaling strategies for containerized agent workloads across Azure Container Apps, GCP Cloud Run, and GKE.
  • Contribute to evaluation frameworks and quality standards for AI agents in production.
MLOps & Python Engineering
  • Write, maintain, and improve production Python code powering data pipelines, agent workflows, and platform tooling.
  • Own the full lifecycle of Python-based services — containerization, deployment, versioning, and runtime behavior.
  • Deploy and operate workflow orchestration using Prefect — scheduling, error handling, retry logic, and human-in-the-loop patterns.
  • Build shared Python tooling and internal packages that enable data science teams to develop and deploy faster.
  • Write and maintain Terraform across Azure and GCP — container registries, managed identities, Key Vault, GCP Secret Manager, storage backends, and VNet configurations.
  • Build and maintain CI/CD pipelines and release management workflows across data science and engineering repositories.
  • Enforce coding standards, security policies, and compliance controls directly in the pipeline.
  • Ensure all production systems are well-documented with clear runbooks and data lineage.
Observability & Reliability
  • Build and own the observability stack — metrics, logging, distributed tracing, alerting.
  • Drive SLO/SLI frameworks and incident response as the platform matures.
  • Troubleshoot production issues end-to-end — from application logic through to infrastructure.
What We’re Looking For
Required
  • 3–5 years in software engineering, DevOps, MLOps, or platform engineering with clear production system ownership.
  • Hands‑on Docker and container orchestration experience in Azure and/or GCP.
  • Terraform across cloud providers — you’ve designed it, not just configured it.
  • CI/CD pipeline experience with Git‑based release management.
  • Systems thinker — you troubleshoot end‑to‑end, not just at the surface.
  • Genuine curiosity about AI and agentic systems — excited to grow into deeper platform concepts.
Strong Signal
  • Policy-as-code enforcement — OPA or equivalent in a production CI/CD context. Tell us about it.
  • Observability depth — if you describe OpenTelemetry as a specification and protocol rather than a tool, we want to talk.
  • Experience with Azure Container Apps, ACI, ACR, Managed Identities, VNets.
  • Experience with GCP Cloud Run, GKE, Vertex AI, IAM, Secret Manager.
  • Familiarity with agentic frameworks — MCP, LangChain, or similar.
  • AI observability platforms — Langfuse, MLflow, or similar.
  • dbt, Snowflake, or similar data transformation and warehousing tools.
Nice to Have
  • Prefect or similar workflow orchestration in production.
  • Multi‑cloud networking and identity management experience.
  • Financial services or fintech domain exposure.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior MLOps Engineer
Senior MLOps Engineer

Jobot • Atlanta (GA)

On-site
USD 150,000 - 175,000
Remote work 100%
Competitive salary + bonus + equity
Medical, dental, vision insurance
+3
MLOps / AIOps / LLMOps / AgentOps Engineer
MLOps / AIOps / LLMOps / AgentOps Engineer

FinOps Weekly • Northern (KY)

Hybrid
USD 120,000 - 160,000
MLOps Engineer
MLOps Engineer

Compunnel, Inc. • San Antonio (TX)

On-site
USD 100,000 - 130,000
MLOps Engineer: Scalable ML Pipelines & Infra
MLOps Engineer: Scalable ML Pipelines & Infra

Compunnel, Inc. • San Antonio (TX)

On-site
AI Infrastructure Engineer MLOps
AI Infrastructure Engineer MLOps

EITACIES Inc. • San Francisco (CA)

On-site
USD 120,000 - 150,000
401(k)
Platform Engineer
Platform Engineer

Atlas Search • New York (NY)

On-site
USD 140,000 - 200,000
ML Ops Engineer — Agentic AI Lab (Founding Team)
ML Ops Engineer — Agentic AI Lab (Founding Team)

Fabrion • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Meaningful equity
Senior AI DevOps Engineer (AI Ops / Platform Engineering)
Senior AI DevOps Engineer (AI Ops / Platform Engineering)

DeepCamp • Tucker (GA)

On-site
USD 96,000 - 165,000
AI Engineer
AI Engineer

Pinpoint Global Communications • United States

On-site
USD 120,000 - 180,000
AI Infrastructure & Platform Engineer
AI Infrastructure & Platform Engineer

International Materials, LLC • Delray Beach (FL), Northern (KY)

Hybrid
USD 120,000 - 180,000