Software Engineer AI/ML Systems - USA

Socket.dev

Santa Clara (CA)

On-site

USD 150,000 - 170,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Unlimited PTO
Generous parental leave
Entrepreneurial culture
Open communication
Small teams = high impact
Medical, Dental, Vision
Disability & Life insurance
Mental health support
Annual bonus program
ESPP
Team building experiences
Mentorship programs
Manager resources

Job summary

Socket.dev is seeking a Software Engineer focused on AI/ML applications to own the end-to-end lifecycle of AI-powered systems. You will deploy, test, and scale models on distributed GPU infrastructure, while building automated testing and secure CI/CD pipelines.

You will drive experiments, benchmark models, and analyze errors with dashboards to communicate performance to stakeholders. Independent ownership and cross-team collaboration are essential.

Qualifications

  • 6+ years of professional experience building production-grade AI/ML systems.

Responsibilities

  • Architect and scale AI infrastructure across multi-node clusters using Kubernetes, Ray, or Slurm.
  • Design and evaluate production-grade AI models and agents with benchmarking capabilities.
  • Run comprehensive model benchmarks and build analytics dashboards for performance.
  • Develop extensive automated testing that covers deterministic code and stochastic AI outputs.
  • Automate DevSecOps with vulnerability scanning on MR workflows and fast remediation pipelines.
  • Own features from ideation to production, including architecture decisions and repo synchronization.
  • Collaborate across teams and with open-source communities.

Skills

AI infrastructure
Python (async)
Kubernetes
Ray
Slurm
LangChain
Hugging Face
MLOps
GitLab pipelines
PyTest
Git advanced workflows

Education

Bachelor's or Master's degree in Computer Science, Engineering, or related field

Tools

Kubernetes
Ray
Slurm
LangChain
Hugging Face
MLOps
vLLM
SGLang
PyTest
GitLab

Job description

The Role

We're seeking a Software Engineer specialized in AI/ML applications to independently drive the development, evaluation, deployment, and end-to-end lifecycle management of AI-powered systems. This role sits at the intersection of advanced AI application development, robust software engineering, and continuous automation: you'll build the systems, and you'll build the machinery that keeps them tested, secure, and running at scale.

This position goes deep on infrastructure and reliability: deploying and scaling open-source models across distributed GPU infrastructure, designing automated testing frameworks that cover both deterministic code and stochastic AI outputs, and implementing secure CI/CD pipelines that enforce quality on every merge.

This role is built for an engineer who takes high ownership: comfortable carrying a feature from ideation through architecture, build, evaluation, and production, and rigorous enough to prove with benchmarks and dashboards that the system actually works.

What You'll Do
  • Architect and scale AI infrastructure: deploy and scale open-source models using distributed orchestration frameworks (Kubernetes, Ray, or Slurm) to run highly available, fault-tolerant AI workloads across multi-node clusters.

  • Design and evaluate AI systems: run experiments, prompt-tune, evaluate, and deploy production-grade models and AI agents, with flexible mechanisms to benchmark performance and swap models quickly as use cases evolve.

  • Drive error and gap analysis: run comprehensive model benchmarks, perform deep error and gap analysis on model outputs, and build analytics dashboards that communicate system performance clearly to stakeholders.

  • Build automated testing at depth: develop extensive automated test suites that validate end-to-end application code, aggressively improving coverage across both standard software and stochastic AI outputs.

  • Automate DevSecOps: keep systems clean and secure by automating vulnerability scanning on GitLab Merge Requests and implementing fast remediation pipelines for identified issues.

  • Execute independently: own features from ideation to production, including architectural decisions, public/private repository synchronization, and open-source community interactions.

What We’re Looking For
  • 6+ years of professional experience writing production-grade, asynchronous Python, with a strong focus on decoupled, clean system architecture and design patterns.

  • Deployment & orchestration: hands-on experience deploying, monitoring, and scaling models in production using Kubernetes, Ray, or Slurm, including multi-node cluster configurations.

  • Hardware & scaling optimization: strong understanding of GPU memory management and infrastructure-level tuning for high-throughput, low-latency AI inference workloads.

  • AI evaluation & frameworks: deep experience building with LangChain, Hugging Face libraries, MLOps, vLLM, and SGLang, with proven expertise in prompt engineering, automated model benchmarking, and systematic LLM evaluations.

  • Data analysis: proficient in Python-based analysis (pandas, NumPy, or similar), able to extract insights from evaluation results and communicate findings clearly to technical and non-technical audiences.

  • CI/CD & security automation: advanced knowledge of GitLab pipelines, including automated test jobs and vulnerability scanners integrated directly into the MR workflow.

  • Testing toolchains: expert familiarity with Python testing frameworks (PyTest), mocking libraries, and automated test generation approaches for AI workloads.

  • Advanced version control: high proficiency in advanced Git workflows, including rebase strategies, cryptographic commit signing, and complex public/private repository mirroring.

  • Bachelor's or Master's degree in Computer Science, Engineering, or a related field (or equivalent experience).

Salary Range: US East/West Coast: $150000 - $170000

Disclaimer: The base salary range is a guideline and may vary based on factors such as candidate experience, specialized skills, and geographical location. Actual compensation may include additional benefits and bonuses.

Perks and Benefits of Working With Us
  • Unlimited PTO.

  • Please ask us about our very generous parental leave, much above industry standards!

  • Entrepreneurial culture where pushing limits and taking risks is everyday business.

  • Open communication with management and company leadership.

  • Small, dynamic teams = massive impact.

  • Medical, Dental and Vision coverage for employees.

  • Access to Disability & Life insurance.

  • Mental health and wellbeing support.

  • Annual bonus program.

  • Employer Stock Purchase Program (ESPP).

  • Yearly team building experiences.

  • Mentorship and sponsorship opportunities.

  • Manager resources and support.

Cogniify is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or any other protected characteristic

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Data Scientist Applied AI - USA
Lead Data Scientist Applied AI - USA

Socket.dev • Santa Clara (CA)

On-site
USD 150,000 - 170,000
Unlimited PTO
Generous parental leave
Annual bonus program
+3
Senior Generative AI Engineer - USA
Senior Generative AI Engineer - USA

Socket.dev • San Francisco (CA)

On-site
USD 150,000 - 170,000
Unlimited PTO
Generous parental leave
Medical, Dental and Vision coverage
+1
Senior AI PM - USA Onsite (Mountain View, CA)
Senior AI PM - USA Onsite (Mountain View, CA)

S27a • Mountain View (CA), Northern (KY)

Hybrid
USD 120,000 - 150,000
Unlimited PTO
Parental leave
Employee stock purchase program (ESPP)
+3
Mid - Forward Deployment Engineer - USA Onsite (Mountain View, CA; New York City, NY)
Mid - Forward Deployment Engineer - USA Onsite (Mountain View, CA; New York City, NY)

S27a • New York (NY)

On-site
USD 165,000 - 175,000
Unlimited PTO
Parental leave
Stock Purchase Program (ESPP)
+2
Mid - Forward Deployment Engineer - USA
Mid - Forward Deployment Engineer - USA

Worky • Mountain View (CA)

On-site
USD 165,000 - 175,000
Unlimited PTO
Generous parental leave
Competitive bonus program
+5
AI Software Engineer
AI Software Engineer

webAI • Austin (TX)

On-site
USD 100,000 - 140,000
Competitive salary
Equity options
Comprehensive health benefits
+8
AI Product Manager Mid Level - USA Hybrid (Atlanta, GA)
AI Product Manager Mid Level - USA Hybrid (Atlanta, GA)

S27a • Atlanta (GA), Northern (KY)

Hybrid
USD 130,000 - 170,000
Unlimited PTO
Parental leave, generous
Entrepreneurial culture
+10
AI Developer
AI Developer

Talentfoot Executive Search • Chevy Chase (MD)

On-site
USD 145,000 - 200,000
AI/ML Engineer (US)
AI/ML Engineer (US)

Latitude • Boston (MA), Northern (KY)

Hybrid
USD 120,000 - 200,000
Medical, dental, and vision insurance
401(k) retirement benefits
Paid time off and holidays
+3
AI/ML Software Engineer / Senior software engineer/ Tech lead
AI/ML Software Engineer / Senior software engineer/ Tech lead

Expert Intelligence™ • San Francisco (CA)

On-site
USD 150,000 - 230,000
Competitive pay bonuses
Personalized PTO
Equal Opportunity employer