Turn this role into an interview — a resume and cover letter built around what this employer wants.
Cbase Inc. in Austin, TX seeks a Senior Data Architect / Data Scientist to work with manufacturing operations. You will analyze factory data, build robust models, and work with operations to turn insights into action.
The role focuses on data understanding, feature creation, and model validation, partnering with manufacturing, quality, maintenance, data engineering, and software teams. It is not a data platform owner role.
We are hiring a Senior Data Scientist to work with factory and industrial data: find what is going wrong (or about to), build models that hold up against plant reality, and help the team act on them.
Most of the job is in the data — understanding how a process behaves, cleaning noisy and incomplete signals, defining “normal” vs “abnormal” with people who run the line, creating features, validating against real outcomes, and explaining limits when the data cannot support a model. You will partner with manufacturing, quality, maintenance, data engineering, and software. You will not own the data platform. This is not a research lab role and not a platform-engineering role.
Someone who can walk a real example: this was the grain of the data, this is what I found, this is the model, this is how I knew it was wrong or right, this is what operations did with it.
Typical problems: process drift, abnormal machine behavior, quality prediction, equipment health, bottlenecks, downtime, scrap/rework, root-cause support. Methods follow the problem (statistical limits, clustering, isolation forest, time series, autoencoders, supervised models when labels exist) — we do not hire to a method list.
Manufacturing experience is a plus. We will also consider people from industrial IoT, equipment, quality, automotive, semiconductor, energy, telecom/ops, or similar operational environments who have done this loop on messy sensor or process data.
Those skills exist on the team or in partner teams. We need the person who works the data.
Frame manufacturing problems with plant and engineering partners; push back when labels, ground truth, or “accuracy” expectations are not real.
Explore, clean, and join fragmented operational data (machines, sensors, quality, maintenance, production, MES/historian extracts — you do not need to have used every acronym).
Build and validate statistical and machine-learning models for anomaly, quality, health, and process monitoring; report false positives/negatives and business cost, not only a leaderboard metric.
Hand usable outputs to engineers and operators (thresholds, explanations, “what to do when this fires”), and support models after they are in use.
Work with data engineering and software on pipelines, Databricks, and production — you are the customer of the platform, not the person hired to build it.
Bachelor’s or master’s in a quantitative or engineering field (data science, CS, statistics, industrial/mechanical/manufacturing engineering, OR, applied math, or related).
10+ years of applied data science (analysis, feature work, statistical or ML modeling on real operational or business datasets). Count data-science years, not total years in IT, DBA, or software engineering.
Strong Python and SQL; evidence of working large, messy tables — not only notebooks on clean extracts.
Production of models you can defend: classification, regression, clustering, anomaly detection, or time series, with a clear target and validation approach.
Experience creating features from machine, sensor, process, quality, maintenance, or other operational data (industrial preferred; high-volume ops data from adjacent domains is acceptable).
Comfort telling stakeholders when a model should not ship.
Ability to learn an unfamiliar plant process quickly.
Time in manufacturing, industrial IoT, semiconductor, automotive, aerospace, energy, or equipment-heavy operations.
Databricks, Spark/PySpark, or similar cloud analytics (we use Databricks; we do not require you to have been the lakehouse owner).
Familiarity with MLOps (tracking, monitoring, drift) as a partner to platform teams.
SPC, explainability, or prior work with historians/MES data.
Business Group – manufacturing, sup group is handling manufacturing propulsions, software interface, production lines, manage product quality as well as propulsions system data
Team structure – they will not really interface with people outside of their group, mainly just working with their current team, their Product Owner interacts with the other departments as needed, the coding team size is 4-5 people, the larger team is around 20 people
Motivator for this position – they are building a brand-new application/system and need help with this process
Chance for an extension – yes at least a year
From scratch build the whole system up
Have access to all manufacturing data, quality data, etc
Want to build AI application
Build a data model
This person can be creative after reviewing the data, help them come up with more ideas and help their business with AI
Education - Bachelor's degree in a technical field such as computer science, computer engineering or related field required
Years of experience – at least 10 years of experience
Must have experience with data ***very important***
Application AI platform skill set is a nice to have, not required
This is a true Data Scientist role
1-Data modeling at least 10 years of experience
2-Data pipeline at least 10 years of experience
3-Data analytics – able to build something out of messy data at least 10 years of experience