Data Engineer

Solve Intelligence

Greater London

On-site

GBP 70,000 - 90,000

Full time

4 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity stake
Founding impact
Visa sponsorship
Private medical insurance
Free meals

Job summary

Solve Intelligence is building the ingestion and search backbone for AI products across patent literature, case law, standards and scientific data. You will own end-to-end pipelines from source acquisition to serving queries, handling large, varied datasets with high throughput and reliability.

You'll collaborate with AI researchers and product engineers, owning schema design, indexing and performance optimisations to enable fast, scalable search across hundreds of millions of records.

Qualifications

  • Strong Python and SQL; experience designing and operating production databases.
  • Experience building and operating production data pipelines over large, messy datasets.
  • Expertise with running search systems over large document collections.
  • End-to-end ownership from raw data to user-facing functionality.
  • Understanding of schema design, indexing and query optimisation.
  • Experience diagnosing and fixing performance bottlenecks in live systems via profiling and measurement.

Responsibilities

  • Large-scale ingestion: build and operate high-throughput pipelines for large datasets with incremental updates and recovery.
  • Document processing and data quality: extract content, handle malformed records, validate outputs while preserving structure.
  • Search and serving: build keyword, vector and structured search; design schemas and partitioning for fast queries.
  • Connecting information across sources: link patents and records with traceability to sources.
  • Performance engineering: profile parsing, ingestion, DB builds and queries; optimize for throughput, latency and cost.

Skills

Python
SQL
Data pipelines
Search systems
Performance optimization

Tools

PostgreSQL
pgvector
OpenSearch
Spark/Delta Lake
AWS

Job description

About Us

We’re the fastest-growing startup transforming the IP industry.

  • Traction: 20-30% MoM revenue growth; selling to 700+ global IP teams (DLA Piper, tech boutiques, and global enterprises).
  • Proven Value: Users report 50-90% efficiency gains using our AI platform.
  • Backing: Recently featured in Sifted following our $40M Series B announcement, bringing our total funding to $55M from elite investors including Y Combinator, 20VC, Visionaries, Microsoft and Thomson Reuters.
About The Role

We’re hiring a data engineer to build the ingestion and search systems behind Solve Intelligence’s AI products.

Our sources include global patent literature, case law, technical standards and contributions, scientific databases, academic papers, and content from across the web. The data spans structured records, documents, images, audio and video. You’ll work across bulk ingestion and on-demand retrieval, making this information searchable and useful in our products.

You’ll own systems from source acquisition through to serving queries. The work includes:

  • Large-scale ingestion. Build and operate high-throughput, resumable pipelines for large datasets, with efficient incremental updates, monitoring and recovery from failures.
  • Document processing and data quality. Extract useful content from complex documents and other formats. Handle malformed records and changing schemas, and validate outputs while preserving structure and metadata.
  • Search and serving. Build keyword, vector and structured search, and design schemas, indexes and partitioning for fast queries over tens to hundreds of millions of records.
  • Connecting information across sources. Link patents, scientific records and supporting documents, preserve dates and versions, and make results traceable to their original sources.
  • Performance engineering. Profile parsing, ingestion, database builds and queries throughout development, testing against representative datasets at realistic scale. Diagnose CPU, memory and storage I/O bottlenecks, and tune jobs and infrastructure for throughput, latency and cost.

You’ll work closely with our AI researchers and product engineers, with substantial freedom to choose the approach and build the systems yourself.

What you bring
Must Haves
  • Strong Python and SQL, with experience designing and operating production databases.
  • Solid experience building and operating production data pipelines over large, messy datasets.
  • Expertise with running search systems over large document collections.
  • End-to-end ownership from raw data to user-facing functionality.
  • A good understanding of schema design, indexing and query optimisation.
  • A track record of diagnosing and fixing performance bottlenecks in live systems through profiling and measurement.
Nice To Have
  • Experience with PostgreSQL/pgvector, OpenSearch (or Elasticsearch), Spark/Delta Lake, AWS, NoSQL databases, or Rust/C++ is useful.
The Founders

You’ll partner with a founding team of AI PhDs and elite systems engineers:

  • Sanj (CRO): PhD in AI (Gatsby Unit, UCL), ex-Huawei R&D, former lead at Magic Carpet AI (acquired).
  • Chris (CEO): PhD in AI (UCL), published researcher, ex-Dyson and Alan Turing Institute.
  • Angus (CTO): MEng Computer Science, ex-Qualcomm and Coremont (Brevan Howard).
What we offer
  • Competitive Salary + Significant Equity: We want you to have true ownership in the success of the company.
  • Founding Impact: You’ll have a direct hand in how we build out the data infrastructure the rest of the product depends on.
  • Support: Full visa sponsorship and private medical insurance.
  • The Environment: Free meals and a seat at the table with an incredibly smart, ambitious team.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

United States Digital Space LLC • Greater London

On-site
GBP 90,000 - 130,000
Visa sponsorship
Private medical insurance
Free meals
Data Engineer
Data Engineer

Solve Intelligence, Inc. • Greater London

Hybrid
GBP 100,000 - 160,000
Significant equity
Founding impact
Visa sponsorship
+2
Infrastructure Engineer
Infrastructure Engineer

Solve Intelligence • Greater London

On-site
GBP 90,000 - 130,000
Competitive salary
Significant equity
Full visa sponsorship
+2
Full-Stack Engineer (Front-End Leaning)
Full-Stack Engineer (Front-End Leaning)

Solve Intelligence • Greater London

On-site
GBP 60,000 - 80,000
Competitive Salary + Significant Equity
Full visa sponsorship
Private medical insurance
+1
AI Engineer
AI Engineer

Solve Intelligence • Greater London

On-site
GBP 65,000 - 85,000
Competitive Salary + Significant Equity
Full visa sponsorship
Private medical insurance
+2
Legal and Product Engineer (Patent Prosecution)
Legal and Product Engineer (Patent Prosecution)

Solve Intelligence • Greater London

On-site
GBP 90,000 - 130,000
Full visa sponsorship
Private medical insurance
Significant equity
+1
Data Engineer
Data Engineer

Senzo • England

Hybrid
GBP 60,000 - 70,000
Competitive salary
Equity
Hybrid working
+1
Founding Infrastructure Engineer
Founding Infrastructure Engineer

Solve Intelligence • England

On-site
GBP 88,000 - 186,000
Competitive Salary + Significant Equity
Full visa sponsorship
Private medical insurance
+1
Data Engineer - End-to-End AI Search Pipelines (Equity)
Data Engineer - End-to-End AI Search Pipelines (Equity)

United States Digital Space LLC • Greater London

On-site
GBP 90,000 - 130,000
Visa sponsorship
Private medical insurance
Free meals
Data Engineer, AI Infrastructure & Search (Equity)
Data Engineer, AI Infrastructure & Search (Equity)

Solve Intelligence, Inc. • Greater London

Hybrid
GBP 100,000 - 160,000
Significant equity
Founding impact
Visa sponsorship
+2