Staff MLOps Engineer – ML Platform

BrightAI Corporation

Palo Alto (CA)

On-site

USD 150,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

BrightAI Corporation in Palo Alto is seeking a Staff MLOps Engineer to lead the build-out of their cloud-native ML developer platform. This role is crucial for developing scalable workflows and deploying reliable ML services on AWS, enhancing decision-making through intelligent automation.

Applicants should have at least 8 years in software/ML engineering, and strong skills in Python, AWS, and CI/CD processes. The company offers competitive compensation and a collaborative work environment.

Qualifications

  • 8+ years in software/ML engineering, including 4+ years in MLOps or a similar role.
  • Strong programming skills in Python; fluent with Docker and Terraform.
  • Experience with AWS services: SageMaker, S3, IAM, CloudWatch.

Responsibilities

  • Design, build, and operate the ML/AI development platform on AWS.
  • Implement CI/CD for ML; ship real-time endpoints and batch jobs.
  • Partner with backend engineers to productionize notebooks and prototypes.

Skills

Software/ML Engineering
MLOps
Python
AWS
Docker
Terraform

Education

B.S. or M.S. in Computer Science or related field

Tools

SageMaker
Terraform
Airflow
Glue
EMR/Spark

Job description

Overview

Bright.AI is a high-growth Physical AI company transforming how infrastructure businesses interact with the physical world through intelligent automation. Our AI platform processes visual, spatial, and temporal data from billions of real-world events—captured across edge devices, mobile sensors, and cloud infrastructure—to enable intelligent decision-making at scale.

You will join as a Staff MLOps Engineer to lead the build-out of our cloud-native ML developer platform and production pipelines. This role is pivotal to building an integrated ML/AI development platform with programmatic data analysis and algorithm development capability on AWS—so teams can move from notebook to secure, reliable, and cost-efficient production services quickly.

You’ll work at the intersection of ML engineering, cloud infrastructure, and developer experience, designing scalable data/model workflows, CI/CD for ML, observability, and governance that turn ideas into durable, monitored ML services.

Key Responsibilities
  • Design, build, and operate our ML/AI development platform on AWS—including Amazon SageMaker AI (Studio/Notebooks, Training/Processing/Batch Transform, Real-Time & Async Inference, Pipelines, Feature Store) and supporting services.
  • Establish golden-path project templates, base Docker images, and internal Python libraries to standardize experiments, data processing, training, and deployment workflows.
  • Implement Infrastructure-as-Code (e.g., Terraform) and workflow orchestration (Step Functions, Airflow); optionally support EKS for training/inference.
  • Build automated data pipelines with S3, Glue, EMR/Spark (PySpark), Athena/Redshift; add data quality (Great Expectations/Deequ) and lineage.
  • Stand up experiment tracking and a model registry (SageMaker Experiments & Model Registry or MLflow); enforce versioning for data, code, and models.
  • Implement CI/CD for ML (CodeBuild/CodePipeline or GitHub Actions): unit/integration tests, data contracts, model tests, canary/shadow deployments, and safe rollback.
  • Ship real-time endpoints (SageMaker endpoints/FastAPI on Lambda/ECS/EKS) and batch jobs; set SLOs and autoscaling, and optimize for cost/performance.
  • Build monitoring & observability for production models and services (drift, performance, bias with SageMaker Model Monitor; service telemetry with CloudWatch/Prometheus/Grafana).
  • Enforce security & governance: least-privilege IAM, VPC isolation/PrivateLink, encryption, secret management.
  • Partner with backend engineers to productionize notebooks and prototypes.
  • Help integrate GenAI/Bedrock services where appropriate; support RAG pipelines with vector stores (OpenSearch) and evaluation harnesses.
Educational Background
  • B.S. or M.S. in Computer Science, Electrical/Computer Engineering, or related field; advanced degree a plus.
  • Strong foundation in machine learning systems, distributed computing, and data engineering; applied experience building production grade ML platforms.
Required Skills & Expertise
  • 8+ years in software/ML engineering, including 4+ years in MLOps or in a similar role.
  • Strong programming skills (proficient in Python), fluent with Docker and Terraform or AWS CDK.
  • Hands-on with AWS: SageMaker, S3, IAM, CloudWatch, ECR, and ECS/EKS/Lambda.
  • Built and operated CI/CD for ML (tests for code/data/models; automated deploys) and shipped real-time & batch ML workloads to production.
  • Experience with experiment tracking & model registry (e.g., SageMaker Experiments/Model Registry or MLflow) and data versioning.
  • Implemented monitoring & quality (SageMaker Model Monitor, EvidentlyAI, Great Expectations/Deequ) and created on-call/runbooks for model & service incidents.
  • Solid grasp of security & compliance in cloud ML (IAM policy design, VPC/private networking, KMS encryption, secrets management, audit logging).
Bonus Qualifications
  • Distributed training at scale (SageMaker Training, PyTorch DDP, Hugging Face on SageMaker).
  • Data engineering at scale (e.g., Spark/EMR, Glue, Redshift).
  • Observability stacks (e.g., Grafana), performance tuning, and capacity planning for ML services.
  • LLMOps/RAG (Bedrock, vector databases, evals) as optional capabilities.
  • Prior startup experience building ML platforms and products from the ground up.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

MLOps Platform Engineer (SageMaker)
MLOps Platform Engineer (SageMaker)

IVIDTEK INC • Plano (TX)

On-site
USD 170,000 - 210,000
Platform Engineer
Platform Engineer

B12 Consulting • Plano (TX)

On-site
USD 150,000 - 190,000
MLOps Engineer
MLOps Engineer

XM • Town of Poland (NY)

On-site
USD 120,000 - 180,000
Private health insurance
International training opportunities
MLOps Engineer
MLOps Engineer

Evlo AI • Miami (FL)

On-site
USD 110,000 - 170,000
MLOps Platform Engineer (SageMaker)
MLOps Platform Engineer (SageMaker)

Ampcus, Inc • Plano (TX)

On-site
USD 180,000 - 240,000
MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
Senior Consultant, AI/ML Engineer
Senior Consultant, AI/ML Engineer

Horizontal Talent • Minneapolis (MN)

On-site
USD 95,000 - 147,000
Medical benefits
Dental benefits
Vision benefits
+1
ML Ops Engineer
ML Ops Engineer

Veriipro • Town of Brookfield (WI)

On-site
USD 120,000 - 160,000
Senior MLOps Engineer
Senior MLOps Engineer

Jobtailor • Ipswich (MA)

On-site
USD 120,000 - 180,000
Machine Learning Engineer
Machine Learning Engineer

Errgo • Town of Boston (NY)

Hybrid
USD 120,000 - 160,000
Medical, dental, and vision insurance
401(k)
Equity
+2