Senior Software Engineer – Big Data & Agentic AI Workflows
Location: Mountainview, CA
12 months
Onsite Only
About the Role
We are looking for a Software Engineer in Data with deep experience building large-scale data platforms, data products and pipelines to join our team. You will design, build, and operate highly scalable, fault-tolerant data processing systems, and help modernize how we deliver data by building agentic workflow solutions that use LLM tools to automate steps across the Data Development Lifecycle (DDLC). This role blends strong traditional big-data engineering with hands-on application of AI agents to accelerate data design, development, testing, and operations.
What You'll Do
- Design, build, and maintain highly scalable, robust, and fault-tolerant data processing pipelines from the ground up.
- Develop database schemas and build ETL/ELT pipelines that process large data volumes reliably and efficiently.
- Build and operate solutions on MPP/Hadoop-based systems, with hands-on work in Hive and Hive on Spark.
- Leverage AWS services (EC2, S3, EMR, Redshift, or equivalent cloud platforms) to build and scale data infrastructure.
- Write advanced scripts and automation in Python, Shell, or similar languages to support pipeline development and operations.
- Build agentic workflow solutions using LLM tools to automate and accelerate tasks across the Data Development Lifecycle (DDLC) — including schema design, code generation, data quality checks, testing, and documentation.
- Design and write analytical SQL for data marts, data warehousing, and broader analytic architecture.
- Create and consume JSON/REST web services to integrate with upstream and downstream systems.
- Own operational health of production pipelines, including problem management, SLA management, and incident response.
- Participate actively in Agile/Scrum teams — sprint planning, estimation, code reviews, and retrospectives.
- Data Maturity Data Products using Data Capabilities with Data Quality ( Data Completeness and Accuracy) and Observability for SLA
- Cost and Performance Optimization of the data pipelines.
What You'll Bring
- Deep software development experience, including building database schemas, developing ETLs, and working with MPP/Hadoop systems.
- Advanced proficiency in a scripting language such as Python or Shell.
- Experience with AWS (EC2, S3, EMR) and Redshift, or equivalent cloud computing platforms.
- Hands-on experience with the Hadoop stack, particularly Hive and Hive on Spark.
- Proven experience designing, building, and maintaining large-scale, fault-tolerant data pipelines.
- Experience working with large data volumes in production environments.
- Experience creating and consuming JSON/REST web services and integrating across systems.
- Practical experience building agentic workflow solutions or applying LLM tools to automate parts of the Data Development Lifecycle.
- Strong grounding in software development methodologies and best practices.
- Experience working in Agile development teams, with working knowledge of Scrum.
- An operational mindset, comfortable with Problem, SLA, and Incident Management.
- Strong expertise writing analytical SQL, and experience with data marts, data warehousing, and analytic architecture.
Education
- BS/MS in Computer Science, or equivalent work experience.