Applied ML Engineer

Knowtex

United States

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Knowtex is building the future of voice AI operating systems for clinicians, transforming how healthcare documentation happens at the point of care. We seek an Applied ML Engineer to productionize and scale machine learning systems powering our voice AI platform.

You will collaborate with ML Scientists and Platform teams to optimize inference, deploy models in regulated healthcare environments, and build robust evaluation and monitoring pipelines to ensure low latency and high reliability in

Qualifications

  • 3–7+ years of experience in ML engineering or Applied ML roles.
  • Strong Python and PyTorch (or TensorFlow) skills.
  • Experience deploying ML models in production environments.
  • Familiarity with transformer architectures and large language models.
  • Experience with model optimization techniques like quantization, distillation, pruning.
  • Experience with cloud infrastructure (AWS preferred).
  • Strong software engineering fundamentals and debugging skills.

Responsibilities

  • Productionize ML models for real-time clinical applications.
  • Optimize inference pipelines for low latency and high throughput.
  • Deploy and scale models using AWS-based infrastructure.
  • Build automated evaluation and regression testing frameworks for LLM outputs.
  • Implement monitoring systems for model performance and drift detection.
  • Collaborate with Backend teams to integrate ML services into APIs and workflows.
  • Improve model efficiency through quantization, batching, caching, and optimization techniques.
  • Support specialty-level model evaluation and performance analysis.
  • Contribute to CI/CD workflows for ML deployment.

Skills

Python
PyTorch
TensorFlow
Production ML deployment
Transformer models
Cloud (AWS)
Software debugging

Tools

SageMaker
Triton Inference Server
CI/CD for ML
AWS

Job description

About Knowtex

Knowtex is building the future of voice AI operating systems for clinicians, transforming how healthcare documentation happens at the point of care. Founded by Stanford AI scientists with deep clinical experience, we're experiencing explosive growth across both commercial health systems and federal healthcare, with our ambient documentation platform scaling rapidly to thousands of clinicians across hundreds of specialties. We're at an inflection point where cutting‑edge AI meets real clinical impact, giving clinicians hours back each day to focus on what matters most - their patients.

Position Overview

We are seeking an Applied ML Engineer to productionize and scale machine learning systems powering our voice AI platform. This role bridges research and engineering — transforming models into reliable, low-latency, production-grade systems deployed across enterprise healthcare environments.

You will work closely with ML Scientists, Backend Engineers, and Platform teams to optimize inference performance, build evaluation pipelines, and ensure robust model deployment in regulated environments.

Key Responsibilities
  • Productionize ML models for real-time clinical applications
  • Optimize inference pipelines for low latency and high throughput
  • Deploy and scale models using AWS-based infrastructure
  • Build automated evaluation and regression testing frameworks for LLM outputs
  • Implement monitoring systems for model performance and drift detection
  • Collaborate with Backend teams to integrate ML services into APIs and workflows
  • Improve model efficiency through quantization, batching, caching, and optimization techniques
    Support specialty-level model evaluation and performance analysis
  • Contribute to CI/CD workflows for ML deployment
Required Qualifications
  • 3-7+ years of experience in machine learning engineering or applied ML roles
  • Strong proficiency in Python and PyTorch (or TensorFlow)
  • Experience deploying ML models in production environments
  • Familiarity with transformer architectures and large language models
  • Experience with model optimization techniques (quantization, distillation, pruning)
  • Experience working with cloud infrastructure (AWS preferred)
  • Strong software engineering fundamentals and debugging skills
Preferred Qualifications
  • Experience with speech recognition systems or NLP pipelines
  • Experience with Triton Inference Server or similar deployment frameworks
  • Familiarity with healthcare data or clinical documentation workflows
  • Experience working in regulated environments (HIPAA, GovCloud, etc.)
  • Knowledge of medical coding systems (ICD-10, CPT)
Technical Environment
  • Python, PyTorch / TensorFlow
  • Transformer-based LLM architectures
  • AWS (SageMaker, ECS, Lambda, S3)
  • Triton Inference Server
  • CI/CD pipelines for ML deployment
  • Observability tools for performance and drift monitoring
Compensation & Benefits
  • Meaningful equity compensation
  • Unlimited PTO
  • Premium health, dental, and vision coverage
  • 401(k) plan
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Engineer: Speech & LLMs
ML Engineer: Speech & LLMs

Socket.dev • San Francisco (CA)

On-site
USD 140,000 - 210,000
Unlimited PTO
Premium health coverage
401(k) plan
+2
ML Engineer: Speech & LLMs
ML Engineer: Speech & LLMs

Knowtex • San Francisco (CA)

Hybrid
USD 120,000 - 180,000
Competitive salary
Equity
Unlimited PTO
+3
Applied ML Engineer: Scale Real-Time Voice AI in Healthcare
Applied ML Engineer: Scale Real-Time Voice AI in Healthcare

Knowtex • United States

On-site
USD 120,000 - 160,000
Software Engineer (Applications Engineering)
Software Engineer (Applications Engineering)

Knowtex • United States

On-site
USD 90,000 - 120,000
Meaningful equity compensation
Unlimited PTO
Premium health, dental, and vision coverage
+1
Director of Machine Learning (Healthcare AI)
Director of Machine Learning (Healthcare AI)

Nxt Level • United States

Hybrid
USD 150,000 - 200,000
Competitive salary
Meaningful equity
Direct line to CEO
+1
Applied AI Engineer
Applied AI Engineer

Norbert Health • New York (NY)

On-site
USD 100,000 - 150,000
Equity participation
Competitive salary
High autonomy and technical ownership
AI / ML Engineer
AI / ML Engineer

Third Way Health, Inc. • Cambridge (MA)

On-site
USD 120,000 - 160,000
Senior Software Engineer Applied AI
Senior Software Engineer Applied AI

Advanced Monitored Caregiving Inc. • Jefferson City (MO)

On-site
USD 140,000 - 190,000
ML Engineer
ML Engineer

FSS Gov Solutions • United States

On-site
USD 120,000 - 150,000
Senior Software Engineer Applied AI
Senior Software Engineer Applied AI

Advanced Monitored Caregiving Inc. • Atlanta (GA)

On-site
USD 140,000 - 190,000