Data Engineer, AI & Distributed Systems

Zignal Labs

San Francisco (CA)

Remote

USD 120,000 - 180,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Zignal Labs is seeking a Data Engineer to build and operate pipelines that ingest, enrich, and structure massive volumes of unstructured data into real‑time intelligence. You’ll own components end‑to‑end and work with engineers across teams to scale streaming, AI, and search services.

You’ll collaborate on production‑grade systems, write robust code, and contribute to CI/CD and infrastructure‑as‑code practices while operating in a fully remote environment across U.S. time zones.

Qualifications

  • 3+ years building and operating data pipelines in production.
  • Strong programming skills in a JVM language — Scala, Java, or Kotlin.
  • Python proficiency for data and scripting work.
  • Hands‑on with Spark and Kafka in a distributed environment.
  • Experience with AWS, Docker, and Kubernetes in production.
  • Proficient SQL plus at least one NoSQL or caching store.

Responsibilities

  • Build and maintain batch and streaming pipelines ingesting high‑volume unstructured data.
  • Develop data pathways feeding NLP, LLM, and retrieval services.
  • Collaborate with search and storage layers for semantic search and real‑time retrieval.
  • Design and implement scalable data services and APIs.
  • Operate and monitor pipelines; write reliable, tested code and review others' work.
  • Collaborate with Data Science, ML, Product, and Security teams to move ideas to production.

Skills

Scala
Java
Kotlin
Python
Apache Spark
Kafka
SQL
NoSQL
Redis
AWS
Docker
Kubernetes
Airflow
Dagster

Education

Bachelor’s degree in Computer Science, Engineering, or a related field

Tools

Databricks
Delta Lake
Flink
Airflow
Prefect
Dagster

Job description

About Zignal Labs

Zignal Labs’ real-time intelligence technology helps the world’s largest organizations protect their people, places, and position. Analyzing billions of data points in real time, Zignal’s AI-powered platform accelerates mission-critical decision making by empowering leaders with contextual situational awareness of the information environment.

Fully remote, with Silicon Valley roots and team members in over 20 states, Zignal serves customers around the world. Learn more at zignallabs.com.

About The Role

We ingest, enrich, and structure massive volumes of unstructured data — from social platforms and news outlets to broadcast media — and turn it into real-time intelligence for our customers.

As a Data Engineer on this team, you’ll build and operate the pipelines that make that possible. You’ll work on systems that process billions of events a day, and on the data pathways that feed our search, NLP, and AI services. You’ll own meaningful pieces of the pipeline end to end, and you’ll do it alongside engineers who have been running these systems at scale for years.

This is a hands-on build-and-operate role. You don’t need to have designed a distributed system from scratch before — you need to be someone who writes solid code, reasons carefully about data correctness and failure modes, and wants to go deep on streaming and AI infrastructure.

What You’ll Do
  • Build and maintain pipelines. Develop and operate batch and streaming pipelines that ingest and enrich high-volume unstructured data. Own components end to end, from implementation through production monitoring.
  • Support our AI systems. Build and extend the data pathways that feed downstream NLP, LLM, and retrieval services — including data preparation, embedding generation, and indexing workflows.
  • Work with search and storage layers. Integrate with and tune our search and vector stores to support semantic search, clustering, and real-time retrieval.
  • Build services and APIs. Implement and improve the microservices and APIs that deliver analytics to enterprise customers.
  • Operate what you build. Write clean, tested, maintainable code. Participate in code review, CI/CD, and infrastructure-as-code practices. Debug production issues and improve reliability over time.
  • Collaborate across teams. Work with Data Science, ML, Product, and Security to take ideas from prototype into production.
What You’ll Need

These are the things we genuinely need on day one.

  • 3+ years building and operating data pipelines in production.
  • Strong programming skills in a JVM language — Scala, Java, or Kotlin. Our core pipeline code is Scala. If you’re strong in Java or Kotlin and want to learn Scala, we’ll support that; we care more about your fundamentals than your current syntax.
  • Working proficiency in Python for data and scripting work.
  • Hands‑on experience with a distributed processing framework, most likely Apache Spark.
  • Hands‑on experience with a streaming platform, most likely Kafka — including a real understanding of consumer groups, offsets, partitioning, and what happens when things fall behind.
  • Practical AWS experience and comfort with Docker. You should be able to work in a Kubernetes environment; you don’t need to administer one.
  • Experience with a workflow orchestrator such as Airflow, Prefect, or Dagster.
  • Solid SQL and experience with at least one NoSQL or caching layer (Redis, MongoDB, DynamoDB, or similar).
  • Sound CS fundamentals — data structures, algorithms, and the judgment to reason about performance and correctness in a distributed setting.
  • Strong written communication and the ability to work asynchronously across U.S. time zones. We’re fully remote; writing clearly is part of the job.
Nice to Have

Genuinely optional. We don’t expect any one candidate to have all of these, and we’re prepared to teach them.

  • Databricks or Delta Lake specifically
  • Flink, or other stream‑processing frameworks beyond Kafka
  • Vector databases (Pinecone, Qdrant, Milvus, pgvector) and hands‑on RAG or embedding pipeline work
  • Elasticsearch or OpenSearch
  • Experience parsing messy, unstructured, or multilingual text at scale
  • Deep database performance tuning, or experience with distributed consensus systems
  • Bachelor’s degree in Computer Science, Engineering, or a related field

Zignal Labs’ real-time intelligence technology helps the world’s largest organizations protect their people, places, and position. Analyzing billions of data points in real time, Zignal’s AI-powered platform accelerates mission-critical decision making by empowering leaders with contextual situational awareness of the information environment.

Fully remote, with Silicon Valley roots and team members in over 20 states, Zignal serves customers around the world.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Data Engineer, AI & Distributed Systems
Remote Data Engineer, AI & Distributed Systems

Zignal Labs, Inc. • United States

On-site
USD 120,000 - 160,000
Remote Data Engineer - Real‑Time AI & Distributed Pipelines
Remote Data Engineer - Real‑Time AI & Distributed Pipelines

Zignal Labs • San Francisco (CA)

Remote
USD 120,000 - 180,000
Data Engineer
Data Engineer

SZNS • Reston (VA)

Hybrid
USD 100,000 - 130,000
Competitive salary and benefits package
Hybrid work environment
Collaborative work environment
+1
Data Engineer
Data Engineer

SZNS Solutions • Reston (VA)

Hybrid
USD 120,000 - 150,000
Competitive salary and benefits package
Hybrid work environment
Continuous learning and development opportunities
+1
Data Engineer
Data Engineer

Unchain Data • Reston (VA)

Hybrid
USD 100,000 - 130,000
Competitive salary and benefits package
Hybrid work environment (MWF in-person)
Collaborative and innovative work environment
+1
AI Engineering Specialist
AI Engineering Specialist

ZS • Evanston (IL)

Hybrid
USD 100,000 - 130,000
Health and well-being benefits
Financial planning support
Professional development programs
Data Engineer
Data Engineer

Sigma • New York (NY)

On-site
USD 140,000 - 180,000
Equity
Generous health benefits
Flexible time off policy
+5
Senior AI Engineer
Senior AI Engineer

Zs Associates • South San Francisco (CA)

On-site
USD 120,000 - 150,000
Health and well-being benefits
Financial planning
Professional development opportunities
Engineering Manager, Data Cloud
Engineering Manager, Data Cloud

chainalysis-careers • Ontario (CA)

On-site
USD 120,000 - 160,000
Diversity and Inclusion initiatives
Dev / ML Ops Engineer
Dev / ML Ops Engineer

Zensors Inc • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive base salary + equity options
Comprehensive health, dental, and vision benefits