Database Administrator

Paycom

Spring (TX)

Hybrid

USD 110,000 - 140,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Brillient Corporation seeks a Data Scientist / ETL Engineer to support an IRS data warehousing initiative in a hybrid Spring, TX role. Design, build, and maintain ETL processes moving data into the CDW while applying statistical and ML techniques.

You will collaborate with data architects, analysts, data quality specialists, and stakeholders to ensure data accuracy, quality, and actionable insights, documenting workflows and models in Jupyter and SQL environments.

Qualifications

  • Bachelor's degree in Data Science, CS, Statistics or related field.
  • US citizenship required.
  • Ability to obtain Public Trust clearance.
  • 8+ years of ETL, data science, analytics or ML experience.
  • Strong SQL ETL experience focusing on load, select, and update commands.
  • Proficiency in Linux/Unix, Sybase IQ, and scripting.
  • Experience with Postgres and Postgres SQL scripting.
  • Strong Python experience for ETL and ML (Pandas, NumPy, Jupyter).
  • Experience building and evaluating ML models with scikit-learn or similar.

Responsibilities

  • ETL design and development for data loading into the CDW.
  • Data science and analytics using Pandas and NumPy.
  • Develop, evaluate, and maintain ML models with scikit-learn.
  • Optimize ETL processes and monitor jobs; ensure data quality.
  • Collaborate with data architects, analysts, and stakeholders.
  • Document ETL processes, workflows, and models.
  • Support data availability and analytical value for stakeholders.

Skills

ETL development
SQL
Python
Linux/Unix
Data science
Machine learning
PostgreSQL
Sybase IQ
OpenShift/Kubernetes
SAS

Education

Bachelor's degree in Data Science, CS, Statistics

Tools

PostgreSQL
Sybase IQ
Python
Jupyter Notebook
OpenShift/Kubernetes
Git
VS Code

Job description

Job Details: Level: Experienced, Job Location: IRS: Spring, TX REMOTE - Spring, TX 77380

Brillient Corporation is seeking a Data Scientist / ETL Engineer to support a mission-critical IRS data warehousing initiative. In this hybrid role, you will design, build, and maintain ETL processes that move data into the Compliance Data Warehouse (CDW). You will also apply statistical and machine learning techniques to that data to produce predictive models and actionable insights. You'll work closely with data architects, business analysts, data quality specialists, and mission stakeholders to make sure data is accurate, consistent, and analytically useful.

This role supports CDW operations within RAAS. That work includes data analysis, process improvement recommendations, and research and evaluation of emerging technologies to improve RAAS data availability, analytics, and value to stakeholders.

The CDW is a non-IT data warehouse containing all of the IRS's return, entity, information return, regulatory, enforcement, web, and security data. The team provides technical guidance and executes work across ETL, database administration, SQL and Bash development, SAS administration, system security, COTS ETL development, metadata, data quality, customer service, web design, and intergovernmental data exchanges.

Key Responsibilities
ETL Design and Development
  • Extract data and tables from Unix/Linux systems, and transform and load them into Sybase IQ and Postgres.
  • Refactor existing ETL jobs and build new solutions as needed. This includes building an object-oriented, Unix-based framework of scripts and stored procedures to run and monitor multiple ETL and statistics processes.
  • Develop and run shell, SQL, and Python scripts to support data processing, automation, and analytical workflows.
  • Build repeatable, scalable data pipelines and analytical processes.
  • Troubleshoot errors in shell and SQL, and configure and operate SSH clients across servers.
Data Science and Analytics
  • Prepare, clean, transform, and explore data, and engineer features, using Pandas and NumPy.
  • Analyze large, complex datasets and tables to find trends, patterns, relationships, and actionable insights.
  • Develop, implement, evaluate, and maintain statistical and machine learning models with scikit-learn or comparable frameworks.
  • Validate model results, assess performance, and find ways to improve accuracy and reliability.
  • Conduct exploratory and statistical analysis to support data-driven decision-making.
  • Develop and document analytical workflows and models in Jupyter Notebook.
  • Work with structured and unstructured data from multiple sources and formats.
Optimization, Performance, and Data Quality
  • Optimize ETL processes for efficient, timely data processing.
  • Monitor ETL jobs and resolve performance bottlenecks and failures promptly.
  • Keep data accurate and intact through validation, cleansing, and auditing.
  • Work with the Data Quality team to resolve data inconsistencies and issues.
Collaboration and Documentation
  • Work with Data Architects to design data models and schemas that meet business needs.
  • Translate business and mission requirements from analysts, SMEs, and stakeholders into ETL and analytical solutions.
  • Document ETL processes, workflows, data dictionaries, methodologies, models, assumptions, and results.
  • Help users move data into Sybase IQ and provide metadata for the metadata repository.
  • Present complex technical findings clearly to both technical and non-technical audiences.
  • Take part in Scrum and client meetings, and keep Kanban board cards up to date.
Environment and Continuous Improvement
  • Work in Linux/Unix and Windows environments, including containerized and distributed platforms (OpenShift, Kubernetes).
  • Keep up with emerging ETL, data science, ML, and AI technologies, and recommend improvements to processes and infrastructure.
Qualifications: Required Education and Qualifications
  • Bachelor's degree in Data Science, Computer Science, Statistics, Mathematics, Engineering, Information Systems, or a related technical field.
  • US citizenship.
  • Ability to obtain and maintain a Public Trust security clearance.
  • At least eight (8) years of combined professional experience in ETL development and data science, analytics, or machine learning.
  • Strong SQL ETL experience, focused on load, select, and update commands.
  • High proficiency in Linux/Unix, Sybase IQ, and Sybase SQL scripting.
  • Working experience with Postgres and Postgres SQL scripting.
  • Strong hands-on Python experience for ETL, data analysis, and machine learning, including Pandas, NumPy, and Jupyter Notebook.
  • Demonstrated experience building and evaluating ML models with scikit-learn or comparable frameworks.
  • Experience with data cleaning, transformation, EDA, feature engineering, and model evaluation.
  • Shell scripting for automation and data processing (KSH, Bash, SSH, sed, awk).
  • Experience working in Linux and Windows environments.
  • Experience with OpenShift and/or Kubernetes.
  • Strong communication skills, and the ability to work both independently and collaboratively in a fast-paced environment.
Preferred Qualifications
  • Experience with AI, LLMs, Retrieval-Augmented Generation (RAG), or model fine-tuning.
  • Experience with Apache Airflow and automated ETL pipelines.
  • Experience with SAP Data Services or other COTS ETL tools.
  • Experience with Git/GitLab, DBeaver, VS Code, SecureCRT/SecureFX, Wiki, JSON, and Perl.
  • Experience deploying ML models in containerized or cloud environments.
  • Experience with MLOps: model deployment, monitoring, and lifecycle management.
DISCLAIMER

The above statements are intended to describe the general nature and level of work performed. They are not intended to be an exhaustive list of all responsibilities, duties, skills, efforts, requirements, or working conditions. Management reserves the right to revise the job or to require that other or different tasks be performed as assigned in accordance with business demands and/or contractual requirements.

Brillient is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to any status protected under applicable federal, state, or local law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote Data Scientist & ETL Engineer for Data Warehouse
Remote Data Scientist & ETL Engineer for Data Warehouse

Paycom - ATS • Spring (TX)

Hybrid
USD 110,000 - 140,000
Database Administrator 1778
Database Administrator 1778

Sistema Technologies Inc. • Austin (TX)

On-site
USD 90,000 - 120,000
Data Scientist
Data Scientist

Brillient-Corporation • Hyattsville (MD)

On-site
USD 120,000 - 140,000
Paid time off
Medical, Dental, and Vision Insurance
Company-Paid Life Insurance and Short‑
+5
Senior Data Engineer (ETL / Python Developer)
Senior Data Engineer (ETL / Python Developer)

CSpring • Indianapolis (IN)

On-site
USD 110,000 - 150,000
Database Administrator 1775
Database Administrator 1775

Sistema Technologies Inc. • Austin (TX)

Hybrid
USD 90,000 - 120,000
Data Engineer
Data Engineer

Jobtailor • Vienna (VA)

On-site
USD 130,000 - 180,000
ETL Database Administrator (local to Austin, TX)
ETL Database Administrator (local to Austin, TX)

Apptad Inc • Austin (TX)

Hybrid
USD 95,000 - 120,000
Sr. Data Warehouse Analyst
Sr. Data Warehouse Analyst

Masterapp Labs • Austin (TX)

Hybrid
USD 96,000 - 138,000
Database Administrator 2 (Hybrid)
Database Administrator 2 (Hybrid)

Serigor Inc. • Austin (TX)

Hybrid
USD 90,000 - 120,000
Senior Data Engineer
Senior Data Engineer

VS Tech Solutions • Boise (ID)

On-site
USD 110,000 - 150,000
Comprehensive Healthcare: Medical, Dental and Vision
401(k) with company match
Fully remote opportunities available
+1