Python Data Engineer: Modernize SAS, Scalable Pipelines

Katmai

United States

Remote

USD 145,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote-friendly

Job summary

Katmai is seeking a seasoned Data Engineer to lead the modernization of core data processing systems. You will translate SAS logic into robust Python solutions and build scalable pipelines for large-scale datasets in a cloud-native environment.

You will work with DSD IT Specialists in an integrated team to deliver maintainable and high-performance data products, with emphasis on quality, documentation, and collaboration.

Qualifications

  • Ability to read and interpret SAS code.
  • Experience with Parquet, SQL, JSON, and larger-than-memory datasets.
  • Experience with PySpark, multiprocessing, Dask, and code profiling tools.
  • Experience in Linux environments and cloud-native resources (e.g., AWS).
  • Strong SQL skills, specifically with PostgreSQL integration.
  • Version control (Git), unit testing, and technical documentation.
  • Experience using Python to read and reformat large data files.
  • Experience improving code performance, including profiling and refactoring code to identify bottlenecks and overhead, preferably in production environments.
  • Experience writing high-quality code that is well-documented and easy to read, use, maintain, and extend.

Responsibilities

  • Analyze existing SAS-based systems and re-implement logic in Python, ensuring functional parity while improving maintainability, scalability, and performance.
  • Design and develop end-to-end data pipelines to ingest, clean, validate, and transform diverse data formats—including SAS datasets, Parquet, JSON, and SQL—using larger-than-memory processing techniques.
  • Profile, optimize, and refactor code to eliminate bottlenecks and improve performance. Implement parallel computing strategies using libraries such as Dask and multiprocessing to efficiently process production-scale workloads.
  • Develop Python solutions that interface with PostgreSQL for high-volume data storage, retrieval, and processing.
  • Implement automated testing, data validation, and quality assurance processes to ensure the integrity, accuracy, and reliability of critical data assets throughout the data lifecycle.
  • Maintain high standards for code quality by producing well-documented, readable, maintainable, and extensible code. Participate in peer reviews and knowledge sharing to promote engineering best practices.
  • Collaborate with DSD IT Specialists and other stakeholders to implement efficient, scalable Python solutions that meet functional and technical requirements.
  • Partner with DSD stakeholders to validate business and technical requirements, troubleshoot implementation issues, and support the successful modernization of legacy systems.
  • Attend branch meetings and communicate technical concepts, project updates, and recommendations clearly and effectively.
  • Maintain regular communication with the team as required.
  • Maintain regular and punctual attendance.
  • Perform other duties as assigned.

Skills

SAS reading
Python data processing
Data pipeline design
Performance optimization
Cloud computing AWS
SQL & PostgreSQL
Linux proficiency
Testing & QA
Documentation

Education

Bachelor's degree in CS or related field

Tools

PySpark
Dask
PostgreSQL
Git
AWS
Parquet
JSON
SQL

Job description

Katmai is seeking a seasoned Data Engineer to lead the modernization of core data processing systems. You will translate SAS logic into robust Python solutions and build scalable pipelines for large-scale datasets in a cloud-native environment.

You will work with DSD IT Specialists in an integrated team to deliver maintainable and high-performance data products, with emphasis on quality, documentation, and collaboration.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Python Data Engineer: SAS-to-Python Modernization (Remote)
Python Data Engineer: SAS-to-Python Modernization (Remote)

Katmai Government Services • United States

On-site
USD 145,000 - 150,000
Medical insurance
Dental insurance
Vision insurance
+3
Python Data Platform Engineer - Scalable Pipelines
Python Data Platform Engineer - Scalable Pipelines

Recru, LLC. • Houston (TX)

On-site
USD 90,000 - 130,000
Senior Data Engineer
Senior Data Engineer

GENNTE Technologies • Pittsburgh

Hybrid
USD 100,000 - 140,000
Lead Data Engineer — AI-Driven Pipelines on Databricks (MN)
Lead Data Engineer — AI-Driven Pipelines on Databricks (MN)

US staffing Inc • Eagan (MN)

On-site
USD 120,000 - 170,000
Senior Data Engineer - Data Pipelines & Data Lake
Senior Data Engineer - Data Pipelines & Data Lake

Modern Technology Solutions, Inc. • Alexandria (VA)

On-site
USD 100,000 - 130,000
Senior Data Engineer — Scalable Pipelines in PySpark/AWS
Senior Data Engineer — Scalable Pipelines in PySpark/AWS

Aviva • Town of Poland (NY)

Hybrid
USD 48,000 - 75,000
Performance Bonus
Private medical care (ENEL-MED)
Cafeteria benefits (MultiSport card)
+1
Senior Python Data Engineer — Cloud-Native Pipelines & APIs
Senior Python Data Engineer — Cloud-Native Pipelines & APIs

Aktra • McLean (VA)

On-site
USD 100,000 - 130,000
Data Engineer
Data Engineer

Silversmith Capital Partners • United States

On-site
USD 85,000 - 110,000
Senior Data Engineer
Senior Data Engineer

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
Senior Data Platform Engineer - AI-Driven Pipelines
Senior Data Platform Engineer - AI-Driven Pipelines

Craft • United States

Hybrid
USD 150,000 - 190,000
Equity
Unlimited vacation
Health + dental + vision insurance (99
+1