Job Summary
The Principal DevOps Platform Engineer will be responsible for driving development and strategy for multiple data and analytics platforms with an emphasis on designing highly available and self‑healing platforms. The role is expected to leverage relationships with business and technology users to learn the business processes the applications support, play a pivotal role in delivering data needs, identifying and driving documentation and process improvements, performing support activities for Technology Services, and engaging technology and business users for ideas to improve the user experience.
Key Accountabilities
- Implement and maintain cloud native machine learning analytics platforms; build services to deploy data ingestion pipelines and machine learning models on a Kubernetes‑based platform.
- Create data policies, standards, guidelines, and procedures to ensure data supporting analytics is available, responsive, and achieves business outcomes.
- Create and implement data and analytics solutions or enhancements to improve overall Vistra data architecture; conduct research and make recommendations on new data management processes and innovations.
- Partner closely with vendors, service providers, the business and internal team to define and understand analytics needs and achieve key performance indicators and service level agreements.
- Determine analytics business needs and translate them into functional requirements, update or create documentation.
- Participate in analytics projects and Agile teams; make recommendations and implement changes to mitigate risks and optimize analytics performance.
Education, Experience, and Skill Requirements
- 10+ years of professional experience in a technology or business field.
- Associate level AWS certification preferred.
- Well versed in DevOps or Site Reliability Engineering (SRE) tenets; 2+ years in an SRE/Operations/DevOps role as part of a team.
- 3+ years hands‑on AWS experience, Kubernetes subject matter expert (EKS preferred).
- Experience in writing and maintaining Helm charts.
- Experience in maintaining ML stack workloads on Kubernetes including Airflow, data pipelines, REST APIs.
- Hands‑on experience with Prometheus or other Kubernetes monitoring solutions.
- 2+ years of hands‑on experience with Python.
- Deliver cloud native DevSecOps on AWS solutions.
- Proven experience analyzing code for non‑standard practices, errors, and security vulnerabilities and presenting detailed remediation plans.
- Experience managing and delivering enterprise‑wide initiatives related to the platforms supported by the team.
- Ability to learn quickly, be self‑motivated to improve knowledge and tackle new challenges.
- Superior organizational, communication, and attention to detail skills; skill at creating executive‑level communications and presenting to leaders.
- Bachelor’s degree in MIS, Computer Science, Mathematics, Engineering, or a related field from an accredited college or university (or equivalent).
Key Metrics
- System availability
- Data availability
- Data quality
Location
Irving, Texas
Equal Employment Opportunity Statement
The Company complies with all employment laws and affords equal employment opportunity to individuals in all aspects of employment, including selection for job opportunities, without regard to race, color, religion, sex, sexual orientation, gender identity, pregnancy, national origin, age, disability, genetic information, military service, protected veteran status, or any other consideration protected by federal, state or local laws. If you are an individual with a disability and need assistance submitting an application or would like to request an accommodation, please email assistance@vistraenergy.com.