Software Engineer (SE / Sr SE), Applied ML & Data Mining

PlusAI

Santa Clara (CA)

On-site

USD 130,000 - 220,000

Full time

12 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

PlusAI is seeking Software Engineers to advance data mining, retrieval, and ranking methods for autonomous-driving datasets. You will design, implement, and evaluate scalable pipelines handling fleet-scale imagery, video, and time-series data, delivering end-to-end mining products with Python APIs, TypeScript/React interfaces, and robust production systems.

The role welcomes candidates with ML or data-mining foundations who are eager to grow across scalable systems and the product stack.

Qualifications

  • BS, MS, or PhD in Computer Science, Machine Learning, or a related technical field, or equivalent practical experience.
  • Hands-on experience developing and evaluating machine-learning, data-mining, computer-vision, or information-retrieval methods with PyTorch or TensorFlow, including experimentation, error analysis, and principled metrics.
  • Foundations for data-centric machine learning: an understanding of model uncertainty and evaluation, embeddings and similarity search, and how training-data composition shapes model behavior.
  • Self-driven with a strong sense of ownership: a quick learner who is eager to take responsibility and drive projects forward end to end.

Responsibilities

  • Develop and evaluate mining, retrieval, and ranking methods using signals such as model confidence, disagreement, embeddings, anomalies, temporal behavior, and learned representations.
  • Build and evolve semantic image/video/scenario search, including text-to-image/video and image-to-image or video-to-video retrieval, vector search, metadata and temporal or spatial filters, task-specific ranking, and search quality, freshness, latency, and reliability.
  • Build and operate distributed mining, inference, and indexing pipelines over fleet-scale imagery, video, time-series, and autonomy-system data, including GPU batch inference, embedding generation, reproducible candidate datasets, and reliable index refreshes.
  • Design and ship mining products end to end: Python APIs and services, relational data models, asynchronous jobs, modern TypeScript/React search and review experiences, deployment, access control, testing, observability, and production reliability.
  • Ensure that your work is performed in accordance with the company’s Quality Management System (QMS) requirements and contribute to continuous improvement efforts.

Skills

ML fundamentals
Data-centric ML
Ownership
Learning ability

Education

BS/MS/PhD in CS or related

Tools

PyTorch
TensorFlow
Milvus
FAISS
pgvector
React

Job description

Finding the right data is central to improving autonomous-driving models. Among petabytes of fleet data, you will develop methods that identify and rank the most valuable moments for training and evaluation, then turn those methods into reliable tools that autonomy and ML engineers use to search, review, and curate datasets. You will work at the intersection of applied machine learning, information retrieval, large-scale data processing, and product engineering. We welcome candidates with ML or data-mining foundations who are excited to grow across scalable systems and the product stack.
We are open to candidates at either the Software Engineer or Senior Software Engineer level. Level will be determined by experience, technical depth, scope of ownership, and demonstrated impact. You do not need experience with every technology in our stack; we value strong fundamentals, ownership, and the ability to learn.

Responsibilities:
  • Develop and evaluate mining, retrieval, and ranking methods using signals such as model confidence, disagreement, embeddings, anomalies, temporal behavior, and learned representations
  • Build and evolve semantic image/video/scenario search, including text-to-image/video and image-to-image or video-to-video retrieval, vector search, metadata and temporal or spatial filters, task-specific ranking, and search quality, freshness, latency, and reliability
  • Build and operate distributed mining, inference, and indexing pipelines over fleet-scale imagery, video, time-series, and autonomy-system data, including GPU batch inference, embedding generation, reproducible candidate datasets, and reliable index refreshes
  • Design and ship mining products end to end: Python APIs and services, relational data models, asynchronous jobs, modern TypeScript/React search and review experiences, deployment, access control, testing, observability, and production reliability
  • Ensure that your work is performed in accordance with the company’s Quality Management System (QMS) requirements and contribute to continuous improvement efforts
Required Skills:
  • BS, MS, or PhD in Computer Science, Machine Learning, or a related technical field, or equivalent practical experience
  • Hands-on experience developing and evaluating machine-learning, data-mining, computer-vision, or information-retrieval methods with a framework such as PyTorch or TensorFlow, including experimentation, error analysis, and principled metrics
  • Foundations for data-centric machine learning: an understanding of model uncertainty and evaluation, embeddings and similarity search, and how training-data composition shapes model behavior
  • Self-driven with a strong sense of ownership: a quick learner who is eager to take responsibility and drive projects forward end to end
Preferred Skills:
  • Production full-stack experience spanning backend services and REST APIs, relational data modeling and SQL, and modern JavaScript or TypeScript frontend development using React or a comparable framework
  • Experience improving models through data — training or fine-tuning, active learning and data flywheels, hard-example mining, uncertainty or disagreement signals, dataset curation — ideally in autonomous driving, robotics, or perception
  • Experience with embedding and multimodal models (CLIP-style models, VLMs), vector databases (Milvus, FAISS, pgvector), or GPU batch inference at scale
$130,000 - $220,000 a year
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer/Senior Software Engineer, Applied ML & Data Mining
Software Engineer/Senior Software Engineer, Applied ML & Data Mining

PlusAI, Inc. • Santa Clara (CA)

On-site
USD 140,000 - 210,000
Applied ML & Data Mining Engineer for Autonomy
Applied ML & Data Mining Engineer for Autonomy

PlusAI, Inc. • Santa Clara (CA)

Hybrid
USD 140,000 - 210,000
Senior AI Data Pipeline Engineer (Autonomous Driving)
Senior AI Data Pipeline Engineer (Autonomous Driving)

42dot • Sunnyvale (CA)

On-site
USD 133,000 - 254,000
Senior AI Data Pipeline Engineer (Autonomous Driving)
Senior AI Data Pipeline Engineer (Autonomous Driving)

42dot Inc. • Sunnyvale (CA)

On-site
USD 133,000 - 254,000
Senior ML Platform Engineer (Autonomous Driving)
Senior ML Platform Engineer (Autonomous Driving)

42dot Inc. • Sunnyvale (CA)

On-site
USD 133,000 - 254,000
Senior/Principal Machine Learning Engineer
Senior/Principal Machine Learning Engineer

University of Michigan Health-West • Northern (KY)

On-site
USD 200,000 - 300,000
Principal Technical Lead Manager - AI Data Applications
Principal Technical Lead Manager - AI Data Applications

Motional • Boston (MA)

On-site
USD 240,000 - 330,000
Hybrid work model
Remote-friendly option
Senior/Staff Software Engineer, ML Data
Senior/Staff Software Engineer, ML Data

Icehouseventures • Mountain View (CA)

On-site
USD 193,930 - 352,290
Staff Software Engineer, Behavior ML Data
Staff Software Engineer, Behavior ML Data

Kindredventures • Mountain View (CA)

On-site
USD 193,930 - 352,290
Annual performance bonus
Equity options
Competitive benefits package
Senior Software Engineer, ML Loop
Senior Software Engineer, ML Loop

NVIDIA • Santa Clara (CA)

On-site
USD 168,000 - 270,000
Equity
Benefits