Senior AI/ML Systems Engineer - Scalable Infra & Automation

Socket.dev

Santa Clara (CA)

On-site

USD 150,000 - 170,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Unlimited PTO
Generous parental leave
Entrepreneurial culture
Open communication
Small teams = high impact
Medical, Dental, Vision
Disability & Life insurance
Mental health support
Annual bonus program
ESPP
Team building experiences
Mentorship programs
Manager resources

Job summary

Socket.dev is seeking a Software Engineer focused on AI/ML applications to own the end-to-end lifecycle of AI-powered systems. You will deploy, test, and scale models on distributed GPU infrastructure, while building automated testing and secure CI/CD pipelines.

You will drive experiments, benchmark models, and analyze errors with dashboards to communicate performance to stakeholders. Independent ownership and cross-team collaboration are essential.

Qualifications

  • 6+ years of professional experience building production-grade AI/ML systems.

Responsibilities

  • Architect and scale AI infrastructure across multi-node clusters using Kubernetes, Ray, or Slurm.
  • Design and evaluate production-grade AI models and agents with benchmarking capabilities.
  • Run comprehensive model benchmarks and build analytics dashboards for performance.
  • Develop extensive automated testing that covers deterministic code and stochastic AI outputs.
  • Automate DevSecOps with vulnerability scanning on MR workflows and fast remediation pipelines.
  • Own features from ideation to production, including architecture decisions and repo synchronization.
  • Collaborate across teams and with open-source communities.

Skills

AI infrastructure
Python (async)
Kubernetes
Ray
Slurm
LangChain
Hugging Face
MLOps
GitLab pipelines
PyTest
Git advanced workflows

Education

Bachelor's or Master's degree in Computer Science, Engineering, or related field

Tools

Kubernetes
Ray
Slurm
LangChain
Hugging Face
MLOps
vLLM
SGLang
PyTest
GitLab

Job description

Socket.dev is seeking a Software Engineer focused on AI/ML applications to own the end-to-end lifecycle of AI-powered systems. You will deploy, test, and scale models on distributed GPU infrastructure, while building automated testing and secure CI/CD pipelines.

You will drive experiments, benchmark models, and analyze errors with dashboards to communicate performance to stakeholders. Independent ownership and cross-team collaboration are essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Solutions Engineer — Platform & Delivery
AI Solutions Engineer — Platform & Delivery

Socket.dev • Indiana (PA)

On-site
USD 120,000 - 160,000
Junior AI/ML Engineer - Mentored, Production-Ready ML
Junior AI/ML Engineer - Mentored, Production-Ready ML

Socket.dev • Auburn Hills (MI)

On-site
USD 60,000 - 90,000
AI/ML Engineer (Mid) Flexible Work & Growth
AI/ML Engineer (Mid) Flexible Work & Growth

Socket.dev • Washington

On-site
USD 110,000 - 170,000
PTO 15 days
11 paid holidays
Medical Insurance options
+8
Senior ML Systems Engineer — Scalable AI Infra
Senior ML Systems Engineer — Scalable AI Infra

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000
GenAI Software Engineer | ML Infra, Python/C++ | Equity
GenAI Software Engineer | ML Infra, Python/C++ | Equity

Socket.dev • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Autonomy ML Engineer: Real-Time End-to-End Deployment
Autonomy ML Engineer: Real-Time End-to-End Deployment

Socket.dev • Santa Clara (CA)

On-site
USD 170,000 - 240,000
Senior AI Agentic Engineer
Senior AI Agentic Engineer

Socket.dev • Atlanta (GA)

On-site
USD 180,000 - 240,000
AI Engineer — Production-Grade Model Reliability
AI Engineer — Production-Grade Model Reliability

Socket.dev • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Systems ML Engineer — Scalable AI Infrastructure
Senior Systems ML Engineer — Scalable AI Infrastructure

Meta • Jackson (MS)

On-site
USD 154,000 - 217,000
ML Engineer - Scalable AI Systems & GPU Orchestration
ML Engineer - Scalable AI Systems & GPU Orchestration

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits