Machine Learning Operations Engineer II

S&P Global, Inc.

Cambridge (MA)

On-site

USD 130,000 - 175,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, Dental, and Vision insurance
Unlimited Paid Time Off
26 weeks paid parental leave
401(k) with 6% matching
Up to $20,000 tuition assistance
Dog-friendly office
Bike-sharing program

Job summary

S&P Global, Inc. is looking for a candidate to join Kensho as a ML Engineer. This role focuses on enhancing the Machine Learning processes, developing tools, and provides training on effective ML practices. Ideal candidates should have 2+ years of experience related to Machine Learning infrastructure and demonstrate proficiency in Python, Kubernetes, and AWS.

We offer a competitive salary of $130,000–$175,000 along with a range of benefits including unlimited PTO, paid parental leave, and robust insurance coverage.

Qualifications

  • 2+ years of experience in ML infra, ML Ops, or a similar skill set.
  • Experience managing distributed systems with Kubernetes.
  • Cloud Platform (AWS) knowledge including EKS and managed ML services.

Responsibilities

  • Iterate on ML processes to develop robust ML workflow tools.
  • Work closely with ML engineers to identify pain points.
  • Provide resources and training on best practices for ML teams.

Skills

Experience in ML infra
Kubernetes management
Cloud Platform (AWS) Knowledge
Python proficiency
Distributed computing frameworks
Software engineering best practices
Debugging distributed systems
Communication skills
Curiosity and eagerness to learn

Tools

Python
Bash
LangGraph
PyTorch
Ray
Amazon EKS
Airflow
Jsonnet
Terraform
Git
GitHub
Sentry
Prometheus
W&B

Job description

About Kensho

Kensho is S&P Global’s hub for AI innovation and transformation. We develop and deploy novel solutions to innovate and drive progress at S&P Global and its customers worldwide, focusing on business and financial generative AI applications, agents, data retrieval APIs, data extraction, and more. Our MLOps team is the de‑facto ML platform team, working at the intersection of infrastructure and ML to empower engineers with state‑of‑the‑art processes, tooling, and infrastructure.

What You’ll Do
  • Iterate on Kensho’s ML processes to develop tools, services, and frameworks that make every stage of the ML workflow robust, auditable, and usable.
  • Work closely with ML engineers to understand their unique processes, identify pain points, and form effective solutions.
  • Empower engineers with stable tooling to rapidly experiment and actualize research into demonstrable prototypes and mature products.
  • Provide resources and training for ML teams on best practices, enabling efficient productionization of their work.
  • Evaluate, select, and champion open‑source and third‑party solutions, driving their adoption across teams and integrating them into Kensho’s platform ecosystem.
  • Ship scalable, efficient, and automated processes for model fine‑tuning, reinforcement learning, and evaluation of LLMs/Agents.
  • Improve LLM and agentic observability to monitor agentic applications in production, detecting performance, decay, and drift issues.
  • Stay at the frontier by tracking emerging tools and frameworks, promoting best practices, and strengthening the technical expertise of the team.
What You’ll Need
  • 2+ years of experience in ML infra, ML Ops, ML Engineering, or a similar skill set.
  • Experience managing distributed systems with Kubernetes (understanding concepts and trade‑offs).
  • Cloud Platform (AWS) knowledge, including EKS and managed ML services such as Bedrock and SageMaker.
  • Python proficiency (we are a Python shop).
  • Familiarity with distributed computing frameworks and workflow orchestration (e.g., Ray, Airflow).
  • Understanding of software engineering best practices in an ML context.
  • Basic understanding of ML concepts, LLMs, and agents.
  • Ability to debug distributed systems across infrastructure, networking, and application layers.
  • Excellent communication skills to drive adoption of new tools and best practices across multiple teams.
  • A curious, driven, low‑ego mindset eager to learn across a range of engineering disciplines.
Technologies & Tools We Use
  • Development: Python, Bash, LangGraph, PyTorch
  • Infrastructure: Ray, Amazon EKS, Airflow, Jsonnet, Terraform
  • Ops: Git, GitHub, AWS, LangFuse, Sentry, Prometheus, W&B
Benefits
  • Base salary range: $130,000–$175,000, plus annual incentive bonus and equity plans.
  • Medical, Dental, and Vision insurance (100% company‑paid premiums).
  • Unlimited Paid Time Off.
  • 26 weeks of fully paid parental leave (paternity and maternity).
  • 401(k) plan with 6% employer matching.
  • Generous company matching on donations to non‑profit charities.
  • Up to $20,000 tuition assistance toward degree programs, plus up to $4,000 per year for professional education such as industry conferences.
  • Plentiful snacks, drinks, and regularly catered lunches.
  • Dog‑friendly office (CAM office).
  • Bike sharing program memberships.
  • Compassion leave and elder care leave.
  • Mentoring and additional learning opportunities.
  • Opportunity to expand professional network and participate in conferences and events.
Equal Opportunity

Kensho is an equal opportunity employer that welcomes future Kenshins with all experiences and perspectives. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, or national origin.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Operations Engineer II
Machine Learning Operations Engineer II

Kensho • Cambridge (MA)

On-site
USD 130,000 - 175,000
Medical, Dental, Vision insurance
401(k) with matching
Parental Leave
Machine Learning Operations Engineer II
Machine Learning Operations Engineer II

S&P Global • Cambridge (MA)

On-site
USD 130,000 - 175,000
Medical, Dental, and Vision insurance
Unlimited Paid Time Off
401(k) plan with company matching
+3
Machine Learning Operations Engineer II
Machine Learning Operations Engineer II

Kensho Technologies • New York (NY)

On-site
USD 120,000 - 150,000
Medical, Dental, and Vision insurance
Unlimited Paid Time Off
26 weeks paid Parental Leave
+7
Machine Learning Engineer II
Machine Learning Engineer II

S&P Global • New York (NY)

On-site
USD 140,000 - 180,000
Medical benefits
Unlimited PTO
Parental leave
+3
Machine Learning Engineer II
Machine Learning Engineer II

S&P Global, Inc. • New York (NY)

On-site
USD 140,000 - 180,000
Medical, Dental, and Vision insurance
Unlimited Paid Time Off
Parental Leave (26 weeks)
Machine Learning Engineer II
Machine Learning Engineer II

Kensho Technologies • New York (NY)

On-site
USD 140,000 - 180,000
Health benefits
Unlimited PTO
Parental leave
+3
Machine Learning Engineer II
Machine Learning Engineer II

S&P Global • Cambridge (MA)

On-site
USD 140,000 - 180,000
Medical benefits
Unlimited PTO
Parental Leave
+2
Machine Learning Engineer II
Machine Learning Engineer II

Kensho Technologies • Cambridge (MA)

On-site
USD 140,000 - 180,000
Medical insurance
Dental insurance
Vision insurance
+6
Machine Learning Engineer II
Machine Learning Engineer II

spgi • New York (NY)

On-site
USD 140,000 - 180,000
Equity / stock options
Performance bonus
Senior Backend Software Engineer
Senior Backend Software Engineer

S&P Global, Inc. • Cambridge (MA)

On-site
USD 150,000 - 225,000
Medical, Dental, and Vision insurance
Unlimited Paid Time Off
26 weeks of paid Parental Leave
+8