An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Creditsafe Group is building a new data and AI platform with data quality, lineage and auditability as core goals. The team will use AWS services like S3, Iceberg, Aurora, DynamoDB, Lambda, Fargate, ECS, Airflow, Spark and will integrate with Snowflake and Databricks to serve customers with high-throughput data insights.
You will work on end-to-end data pipelines, design scalable architectures, and contribute to AI capabilities on top of the platform, using Python and Terraform for
With the largest owned database supported by a partner network, Creditsafe provides data on 320 million companies in over 100 countries to support its customers in this regard. Creditsafe is scaling business and currently going through a significant growth stage. Being a data company, Creditsafe collects and processes large amounts of complex data, received from multiple channels with the only purpose of serving our clients with opportunities to gain actionable insights. Our holistic data aggregation helps our clients to fill the knowledge gaps, influence decision making, and find hidden patterns.
This opportunity is to join a team of highly technically skilled engineers who are designing Creditsafe’s new data and AI platform with data quality, data lineage, and auditability as the primary goals whilst maintaining high throughput and scalability. The data platform will be built using AWS infrastructure such as S3, Iceberg, Aurora, DynamoDB, Lambda, Fargate, ECS, Airflow, Spark. A key aspect of this role with be helping to implement the integration of the AWS data platform across Snowflake and Databricks to meet the data needs of our customers in their native data cloud platform. The platform is expected to manage over one hundred million objects with a high volume of daily updates handling addition, deletion, and correction of our data and indexes in an auditable manner. We will also be building AI capabilities on top of these platforms based on MCP and AI agentic workflows. Our data processing application is entirely based on Python and we use Terraform to build re-creatable infrastructure.
Our Technology Stack: Python, Linux, Airflow, AWS DynamoDB, S3, Glue, Athena, Redshift, lambda, API Gateway, Terraform, CI/CD