We are seeking an experienced Geospatial Data Engineer with Palantir Foundry experience to support a large-scale wildfire modeling and grid safety initiatives.
We have developed and currently utilizes wildfire models that leverage significant volumes of ArcGIS, geospatial, utility asset, and location-based data to support wildfire risk analysis and operational decision-making.
The Geospatial Data Engineer will be responsible for building and supporting high-performance data pipelines and geospatial processing workflows that feed these models. The role requires strong Data Engineering expertise combined with hands-on experience processing and manipulating large-scale spatial datasets, particularly using Apache Sedona, PySpark, Python, and SQL.
The ideal candidate will also have experience working within Palantir Foundry or a comparable enterprise cloud/big-data environment and understand how to efficiently partition, transform, validate, and process complex geospatial datasets at scale.
Core Responsibilities
- Design, build, optimize, and support high-performance automated data pipelines in AWS and cloud environments.
- Utilize PySpark and Python to ingest, transform, and process large-scale utility asset, outage, and geospatial datasets.
- Build scalable ETL/data engineering solutions capable of handling large historical and location-based datasets.
- Optimize pipelines for performance, reliability, scalability, and maintainability.
- Develop and support automated end-to-end data workflows within Palantir Foundry.
- Build and maintain Foundry data transformations, pipelines, and code repositories.
- Work within the Foundry ecosystem to prepare and deliver datasets supporting wildfire models and downstream analytical applications.
- Troubleshoot and optimize existing Foundry pipelines and data processing workflows.
Distributed Geospatial Processing
- Implement advanced distributed geospatial data processing using Apache Sedona.
- Utilize spatial libraries and frameworks such as GeoPandas, Shapely, and Apache Sedona/GeoSpark.
- Process, manipulate, and analyze large volumes of ArcGIS and other location-based data.
- Develop efficient approaches for partitioning and processing extremely large spatial datasets.
- Perform and optimize spatial joins, indexing, geometry operations, and other distributed geospatial workloads.
Wildfire Modeling & Grid Safety Support
- Support the underlying data engineering capabilities and datasets used by existing wildfire models.
- Ensure geospatial and utility data is appropriately processed, transformed, validated, and delivered to support wildfire-related analysis and operational decision-making.
- Interface with systems and datasets supporting:
- Remote Inspections
- Public Safety Power Shutoffs (PSPS)
- Work with large-scale utility asset and outage data used in grid safety and wildfire risk initiatives.
- Optimize distributed data processing to improve performance and scalability.
- Work with complex database structures, network topology, and spatial analytics frameworks.
- Support large historical datasets and enterprise-scale data models.
- Troubleshoot data quality, pipeline, geometry, performance, and scalability issues.
- Participate actively in Agile/Scrum ceremonies.
- Apply strong software engineering principles including:
- Unit testing
- CI/CD
- Source/version control
- Code reviews
- Reusable and maintainable development practices
- Collaborate with Data Scientists, Engineers, GIS specialists, modeling teams, and other stakeholders supporting wildfire and grid safety initiatives.
Required Qualifications
Education
- Bachelor's degree in Computer Science, Engineering, GIS, or another related quantitative/technical discipline.
- 5+ years of experience working within Data Engineering, ETL, or large-scale data processing ecosystems.
- Strong hands-on proficiency with:
- PySpark
- Python
- SQL
- Demonstrated experience building scalable data pipelines for large and complex datasets.
- Strong understanding of distributed data processing and performance optimization.
- Strong hands-on experience working with large-scale geospatial and spatial datasets.
- Experience with geospatial frameworks/libraries including:
- Experience working with ArcGIS-related or comparable enterprise geospatial datasets.
Coordinate Reference Systems
Strong understanding of Coordinate Reference Systems (CRS), including:
- EPSG codes
- NAD83
- WGS84
- Coordinate transformations and projections
Candidates should understand how differences between coordinate systems impact spatial processing, transformations, joins, distance calculations, and analytical results.
Spatial Data Formats
Hands-on knowledge of common vector and spatial data formats, including:
- Shapefile
- GeoJSON
- GeoParquet
- GeoPackage
- KML
Experience working with raster formats
Strong understanding of distributed spatial indexing and partitioning concepts
- R-trees
- Grid indexing
- Quadtree indexing
- Apache Sedona partitioning and optimization techniques
Candidates should understand how to determine the appropriate processing and partitioning strategy for very large geospatial datasets.
Geometry Operations at Scale
Experience performing and optimizing large-scale geometry operations including:
- Intersections
- Nearest-neighbor calculations
- Topology validation
- Geometry simplification
- Handling invalid geometries
The consultant should understand both the functional and performance implications of executing these operations across large distributed datasets.
Workflow Orchestration
Experience with workflow orchestration technologies such as:
- Airflow
- Palantir Foundry-native orchestration/workflow capabilities
Data Modeling & Warehousing
Strong understanding of enterprise Data Engineering concepts including:
- Dimensional data modeling
- Slowly Changing Dimensions (SCD)
Historical data management - Large-scale data warehousing
- ETL/ELT patterns
- Data quality and validation
- Hands-on experience working within Palantir Foundry is strongly preferred.
- Candidates with comparable experience on large-scale cloud or enterprise big-data platforms may also be considered.
- Experience developing automated data transformations, pipelines, workflows, and code repositories within an enterprise data platform.
- Experience dealing with:
- Complex database structures
- Spatial analytics frameworks
- Ability to use technologies such as Apache Sedona/GeoSpark to efficiently resolve and process complex spatial datasets.
Preferred Experience
- Experience supporting utility, energy, infrastructure, or other asset-intensive organizations.
- Experience working with wildfire modeling or wildfire risk data.
- Experience with electric utility asset and outage datasets.
- Experience with Public Safety Power Shutoffs (PSPS).
- Experience with LiDAR and vegetation management datasets.
- Experience with ArcGIS and enterprise GIS environments.
- Experience supporting Data Science, predictive modeling, or risk modeling teams.