Tech Lead, Data & Inference Engineer

Catalyst Labs

New York (NY)

Hybrid

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading tech company in New York City is seeking a Tech Lead, Data & Inference Engineer to lead the design and development of a robust data platform. This role involves scaling data pipelines, ensuring reliability and data quality, and mentoring engineers across teams. Candidates should have 6 to 12 years of experience in building production-grade data systems, proficiency in SQL and Python, and familiarity with distributed data technologies. The position supports hybrid or distributed work environments.

Qualifications

  • 6 to 12 years of experience building and scaling production-grade data systems.
  • Deep expertise in data architecture, modeling, and pipeline design.
  • Excellent written and verbal communication; proactive and collaborative mindset.

Responsibilities

  • Lead the design, development and scaling of an end-to-end data platform.
  • Build and maintain scalable batch and streaming pipelines.
  • Take full ownership of reliability, cost, and service-level objectives.

Skills

SQL (query optimization on large datasets)
Python
Data architecture
Data pipeline design
Distributed data technologies (Spark, Flink, Kafka)
Modern orchestration tools (Airflow, Dagster, Prefect)
Node.js

Education

Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or Mathematics

Tools

Kubernetes
Cloud infrastructure (AWS, GCP, or Azure)
dbt
DuckDB
IaC
CI/CD

Job description

Tech Lead, Data & Inference Engineer

Lead the design, development and scaling of an end‑to‑end data platform from ingestion to insights, ensuring that data is fast, reliable and ready for business use.

Build and maintain scalable batch and streaming pipelines, transforming diverse data sources and third‑party APIs into trusted and low‑latency systems.

Take full ownership of reliability, cost, and service‑level objectives, including achieving 99.9% uptime, maintaining minute‑level latency and optimizing cost per terabyte. Conduct root‑cause analysis and provide long‑lasting solutions.

Operate inference pipelines that enhance and enrich data—enrichment, scoring and quality assurance using large language models and retrieval‑augmented generation. Manage version control, caching, and evaluation loops.

Work across teams to deliver data as a product through the creation of clear data contracts, ownership models, lifecycle processes and usage‑based decision making.

Guide architectural decisions across the data lake and the entire pipeline stack. Document lineage, trade‑offs, and reversibility while making practical decisions on whether to build internally or buy externally.

Scale integration with APIs and internal services while ensuring data consistency, high data quality, and support for both real‑time and batch‑oriented use cases.

Mentor engineers, review code and raise the overall technical standard across teams. Promote data‑driven best practices throughout the organization.

Work type: Full Time.

Qualifications

Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or Mathematics.

Excellent written and verbal communication; proactive and collaborative mindset.

Comfortable in hybrid or distributed environments with strong ownership and accountability.

A founder‑level bias for action to identify bottlenecks, automate workflows, and iterate rapidly based on measurable outcomes.

Demonstrated ability to teach, mentor, and document technical decisions and schemas clearly.

Core Experience

6 to 12 years of experience building and scaling production‑grade data systems, with deep expertise in data architecture, modeling, and pipeline design.

Expert SQL (query optimization on large datasets) and Python skills.

Hands‑on experience with distributed data technologies (Spark, Flink, Kafka) and modern orchestration tools (Airflow, Dagster, Prefect).

Familiarity with dbt, DuckDB, and the modern data stack; experience with IaC, CI/CD, and observability.

Exposure to Kubernetes and cloud infrastructure (AWS, GCP, or Azure).

Bonus: Strong Node.js skills for faster onboarding and system integration.

Previous experience at a high‑growth startup (10 to 200 people) or early‑stage environment with a strong product mindset.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Tech Lead, Data & Inference Engineer
Tech Lead, Data & Inference Engineer

Catalyst Labs • Greenwich (CT)

Hybrid
USD 120,000 - 160,000
Tech Lead, Data & Inference Engineer
Tech Lead, Data & Inference Engineer

Catalyst Labs • Jacksonville (FL)

Hybrid
USD 130,000 - 160,000
Tech Lead, Data & Inference Engineer
Tech Lead, Data & Inference Engineer

Catalyst Labs • Connecticut

Hybrid
USD 120,000 - 180,000
Tech Lead, Data & Inference Engineer
Tech Lead, Data & Inference Engineer

Catalyst Labs • Seattle (WA)

Hybrid
USD 120,000 - 160,000
Tech Lead, Data & Inference Engineer
Tech Lead, Data & Inference Engineer

Catalyst Labs • Miami (FL)

Hybrid
USD 120,000 - 160,000
Tech Lead, Data & Inference Engineer
Tech Lead, Data & Inference Engineer

Catalyst Labs • Massachusetts

Hybrid
USD 120,000 - 160,000
Tech Lead, Data & Inference Engineer
Tech Lead, Data & Inference Engineer

Catalyst Labs • San Francisco (CA)

On-site
USD 180,000 - 280,000
Tech Lead, Data & Inference Engineer
Tech Lead, Data & Inference Engineer

Catalyst Labs • Illinois

Hybrid
USD 120,000 - 160,000
Tech Lead, Data & Inference Engineer
Tech Lead, Data & Inference Engineer

Catalyst Labs • New Jersey

Hybrid
USD 130,000 - 180,000
Tech Lead, Data & Inference Engineer
Tech Lead, Data & Inference Engineer

Catalyst Labs • Palo Alto (CA)

On-site
USD 150,000 - 180,000