Software Engineer, ML Data Systems

Cursor

New York, San Francisco (NY, CA)

On-site

USD 100,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cursor is seeking a qualified Data Infrastructure Engineer based in New York. This role involves designing, maintaining, and optimizing data systems with a focus on correctness and performance. The ideal candidate will have deep expertise in Spark, production experience with Ray Data, and hands-on ownership of large data pipelines.

Responsibilities include ensuring data quality and performance, implementing data retention strategies, and maintaining system reliability. Candidates with experience in ClickHouse and dbt will be preferred. Flexible work arrangements are available.

Qualifications

  • Deep experience with building real systems at scale.
  • Hands-on ownership of large data pipelines and storage systems.
  • Clear thinking about data modeling and long-term maintainability.

Responsibilities

  • Own the design and maintenance of data systems and infrastructure.
  • Ensure correctness and performance in data workflow.
  • Implement retention and compression strategies for data storage.

Skills

Experience with Spark
Production experience with Ray Data
Debugging performance issues
Data modeling

Tools

ClickHouse
dbt
Dagster

Job description

About the Role

Cursor ships daily. Every release leaves signals behind: telemetry, prompts, completions, agent runs, sessions. Those signals power model improvement, evals, and experimentation. Data infrastructure is what turns them into something teams can trust.

A lot of systems here started simple so we could move fast. Over time, the constraints change and the “good enough” version becomes the bottleneck. This role owns the full ladder: patch what should be patched, redesign what should be redesigned, ship the replacement, and operate it.

Privacy guarantees are part of correctness. What we can retain and use depends on Privacy Mode and org configuration, and getting that wrong breaks a product promise. We choose work by business impact: what blocks product and model teams today, and what will block them next month.

Sample projects include...
  • A core pipeline started as a pragmatic reuse of infrastructure built for something else. It works, but it cannot guarantee properties downstream consumers now need (for example, point‑in‑time consistency). You design and ship the replacement while keeping the existing system running.

  • A new product surface ships without instrumentation. You talk to the team, define what needs to be captured, and wire it through before the absence becomes anyone else’s problem.

  • Eval coverage drops. You trace it to an instrumentation gap introduced weeks ago by a product change nobody flagged. You fix the gap, add a contract so it cannot recur, and ship the dashboard that would have caught it earlier.

  • Multiple consumers depend on overlapping data. You design schema evolution and validation so changes in one place do not silently degrade the others.

  • Storage costs rise faster than usage. You decide what is worth keeping, implement retention and compression, and delete what is not.

What we're looking for

We’re looking for someone who has built real systems at scale and cares about correctness, cost, and ergonomics.

Strong signals include:

  • Deep experience with Spark (Databricks or open‑source Spark both count)

  • Production experience with Ray Data

  • Hands‑on ownership of large data pipelines and storage systems

  • Comfort debugging performance issues across client instrumentation, streaming, storage, and model‑facing workflows, as well as, compute, storage, and networking layers

  • Clear thinking about data modeling and long‑term maintainability

  • You have good judgment about when to patch and when to rebuild

Nice to have
  • Experience running or scaling ClickHouse

  • Familiarity with dbt, Dagster, or similar orchestration and modeling tools

We're in‑person with cozy offices in North Beach, San Francisco and Manhattan, New York, replete with well‑stocked libraries.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Data/ML
Software Engineer, Data/ML

XOXO AI Inc. • San Francisco (CA)

On-site
USD 250,000 - 500,000
Top-tier health benefits
Dental & vision coverage
Equity 1–5%
Backend/Infra Engineer
Backend/Infra Engineer

Judgment Labs • San Francisco (CA)

On-site
USD 140,000 - 180,000
Software Engineer III
Software Engineer III

United States Digital Space LLC • Town of Oakland (WI)

Hybrid
USD 120,000 - 160,000
Full Lifecycle Data Engineer
Full Lifecycle Data Engineer

Lockton • Kansas City (MO)

On-site
USD 110,000 - 160,000
Software Engineer (Data Platform)
Software Engineer (Data Platform)

The Recruiting Guy • Boston (MA)

Remote
USD 125,000 - 200,000
Backend/Infra Engineer
Backend/Infra Engineer

Judgment Labs Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff — Data Infrastructure
Member of Technical Staff — Data Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 130,000 - 170,000
Data & Analytics Engineering Manager
Data & Analytics Engineering Manager

understood • New York (NY)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff — Data Infrastructure
Member of Technical Staff — Data Infrastructure

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Member of Technical Staff — Data Infrastructure
Member of Technical Staff — Data Infrastructure

Causal • San Francisco (CA)

On-site
USD 150,000 - 190,000