Principal Artificial Intelligence Engineer

Socket.dev

San Francisco (CA)

On-site

USD 260,000 - 340,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Generous PTO
Parental Leave
Rewards & Recognition
Insurance Benefits

Job summary

Innovaccer seeks a Principal AI Engineer to build production-grade AI systems at scale. You will take ideas from paper to prototype to production, design multi-model pipelines, and ensure product-level accuracy through rigorous evaluation.

Role requires deep expertise in Python, PyTorch, and distributed training, with prior production ML experience. You will mentor engineers and work with clinicians and operators to ensure reliable outputs.

Qualifications

  • MS or PhD in Computer Science, Machine Learning, or related quantitative field; exceptional BS candidates considered.
  • First-author publications at NeurIPS/ICML/ICLR/ACL/EMNLP or similar; open-source ML contributions; model training and shipping experience.
  • Fine-tuned open-weight models; explain choice between parameter-efficient vs full fine-tuning.
  • Strong Python and PyTorch; familiarity with training and serving stacks (HuggingFace, DeepSpeed, vLLM or equivalents).
  • Exposure to multi-GPU training and understanding sharding strategy and its importance.
  • Experience finishing projects; production-ready models; data pipelines for training data (dedup, filtering, decontamination).
  • Ability to scope ambiguous problems into plans; mentoring or direction for other engineers is a plus.

Responsibilities

  • Take ideas from paper to prototype to production and build real product model layers.
  • Choose model sizes, compose models into a working system, and maintain a product-level accuracy bar.
  • Write real code in Python and PyTorch; optimize training runs and system efficiency.
  • Design experiments with hypotheses, ablations, and report results.
  • Explain work to clinicians and operators who will rely on outputs.
  • Lead post-training on large models and own outcomes across experiments.

Skills

Python
PyTorch
Distributed training
Model deployment
Research publications
Multi-GPU training
System design
Mentoring

Education

MS or PhD in Computer Science / ML

Tools

HuggingFace
DeepSpeed
FSDP
vLLM
SGLang

Job description

About the Role

We are looking for a Principal AI Engineer to join our growing AI team and help build intelligent, production-grade AI systems that solve complex problems at scale.

A Day in the Life
  • You take an idea from paper to prototype to production. If you have only ever done one of those three, this role will stretch you, and we are fine with that if the rest is strong.
  • You can build the model layer of a real product, not just a model. That means choosing model sizes, composing several models into a working system, and holding a product-level accuracy bar.
  • You write real code. Python fluently, PyTorch fluently, and enough systems sense to know why your training run is slow.
  • You design experiments. You state the hypothesis, run the ablation, and report the result that disagrees with you.
  • You measure things. You are suspicious of results that look good, and you build the eval before you build the model.
  • You read current research and can tell the difference between a technique that will hold up and one that will not.
  • You explain your work to people who are not AI engineers, including clinicians and operators who will tell you when your output is wrong.
What You Need
  • MS or PhD in Computer Science, Machine Learning, or a related quantitative field. Exceptional BS candidates with substantial research or open-source work will be considered.
  • Depth beyond coursework: first-author publications at NeurIPS, ICML, ICLR, ACL, EMNLP, or similar; meaningful open-source ML contributions; a research internship at an AI lab; or models you trained and shipped that people actually used.
  • You have fine-tuned an open-weight model yourself, understand the difference between parameter-efficient and full fine-tuning, and can explain why you chose one.
  • Strong Python and PyTorch. Familiarity with the current training and serving stack (HuggingFace, FSDP or DeepSpeed, vLLM or SGLang, or equivalents).
  • Some exposure to multi-GPU training, even at lab scale. You should know what a sharding strategy is and why it matters.
  • Evidence you can finish things.
  • You have post-trained models that ran in production and you know what broke.
  • You have built the model layer behind a real product feature, including fine-tuning at more than one model size and composing multiple models into a system that met a production accuracy and latency bar.
  • Hands-on multi-node training experience, tens of GPUs at minimum, on models in the tens of billions of parameters or larger.
  • You have built or substantially owned a data pipeline that fed a real training run, including the unglamorous parts: dedup, filtering, decontamination, format normalization.
  • You can scope an ambiguous problem into a plan and tell the difference between a research question and an engineering task.
  • Mentoring or setting direction for other engineers is a plus and will matter more as the team grows.
  • You have led post-training on models in the hundreds of billions of parameters, on runs spanning hundreds of GPUs, and you owned the outcome rather than a slice of it.
  • You have designed the complete model architecture behind a shipped product, spanning small, medium, and large models, and made the tier and composition decisions that determined whether it worked economically.
  • You have made the calls that decide whether a model program works: data mixture, objective design, when to stop a run, when to abandon an approach that a team is emotionally invested in.
  • Deep familiarity with large-scale distributed training failure modes and the operational discipline that keeps a long run alive.
  • A public or verifiable record: papers, model releases, systems, or a training program whose results speak for themselves.
  • You can set a two-to-four-quarter technical direction and then get a team to execute it. You will effectively define what in-house modeling means at Innovaccer.

We offer competitive benefits to set you up for success in and outside of work.

Here's What We Offer
  • Generous PTO Benefits: Enjoy PTO benefit accrual of 20 days per year.
  • Parental Leave: Experience one of the industry's best parental leave policies to spend time with your new addition.
  • Rewards & Recognition: Unlock your potential and be rewarded generously with both monetary incentives and widespread recognition for your dedication and outstanding performance. Unlock your potential and be rewarded generously with both monetary incentives and widespread recognition for your dedication and outstanding performance.
  • Insurance Benefits: We offer medical, dental, and vision benefits along with 100% company-sponsored short and long-term disability and basic life insurance. Legal aid and pet insurance options are available at a discounted rate.

Innovaccer is an equal opportunity employer. We celebrate diversity, and we are committed to fostering an inclusive and diverse workplace where all employees, regardless of race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, marital status, or veteran status, feel valued and empowered.

About Innovaccer

Innovaccer activates the flow of healthcare data, empowering providers, payers, and government organizations to deliver intelligent and connected experiences that advance health outcomes. The Healthcare Intelligence Cloud equips every stakeholder in the patient journey to turn fragmented data into proactive, coordinated actions that elevate the quality of care and drive operational performance. Leading healthcare organizations like CommonSpirit Health, Atlantic Health, and Banner Health trust Innovaccer to integrate a system of intelligence into their existing infrastructure- extending the human touch in healthcare. For more information, visit www.innovaccer.com.

Check us out on YouTube, Glassdoor, LinkedIn, Instagram, and the Web.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Artificial Intelligence Engineer
Principal Artificial Intelligence Engineer

Innovaccer Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Generous PTO
Parental Leave
Rewards & Recognition
+1
Principal Artificial Intelligence Engineer
Principal Artificial Intelligence Engineer

Innovaccer Analytics • San Francisco (CA)

On-site
USD 250,000 - 420,000
PTO benefits
Parental leave
Rewards & recognition
+1
Principal Artificial Intelligence Engineer
Principal Artificial Intelligence Engineer

Innovaccer • San Francisco (CA)

On-site
USD 210,000 - 300,000
PTO 20 days/year
Parental leave
Rewards & Recognition
+1
Senior Artificial Intelligence Engineer
Senior Artificial Intelligence Engineer

Innovaccer Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Generous PTO: 20 days per year
Parental Leave
Rewards & Recognition program
+1
Senior Artificial Intelligence Engineer
Senior Artificial Intelligence Engineer

Innovaccer Analytics • San Francisco (CA)

On-site
USD 230,000 - 340,000
PTO benefits
Parental leave
Rewards & recognition
+1
Senior Artificial Intelligence Engineer
Senior Artificial Intelligence Engineer

Socket.dev • San Francisco (CA)

On-site
USD 170,000 - 260,000
Generous PTO
Parental Leave
Rewards & Recognition
+1
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Innovaccer • San Francisco (CA)

On-site
USD 150,000 - 230,000
Generous PTO benefits
Parental Leave
Rewards & Recognition
+1
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Innovaccer Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Generous PTO
Parental Leave
Rewards & Recognition
+1
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Innovaccer Analytics • San Francisco (CA)

On-site
USD 140,000 - 210,000
Generous PTO
Parental Leave
Rewards & Recognition
+1
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Socket.dev • San Francisco (CA)

On-site
USD 140,000 - 230,000
PTO 20 days per year
Parental Leave
Rewards & Recognition
+1