Conducts data analytics, data engineering, data mining, exploratory analysis, predictive analysis, and statistical analysis, and uses scientific techniques to correlate data into graphical, written, visual and verbal narrative products, enabling more informed analytic decisions.
Proactively retrieves information from various sources, analyzes it for better understanding about the data set, and builds AI tools that automate certain processes. Duties typically include:
Responsibilities
- Create data packages in the form of databases, reports, and visualizations.
- Communicate ongoing data science activities, technical findings, and data products for both technical and non‑technical customers.
- Extract relevant features from large data stores containing open source, PIA, and CAI data that may have bad records, partial records, errors, or other forms of noise.
- Extract features from open source information stored in a wide range of possible formats, including JSON, XML, raw text logs, industry‑specific encodings, and graph link data.
- Apply natural language processing, computer vision, signal processing, and speaker and speech recognition algorithms to identify objects in text, image, video, and audio files.
- Apply descriptive and inferential statistics to describe data and make predictions about the data, including statistical tests to determine confidence for a hypothesis, common summary statistics (e.g., mean, variance, and counts), fit distributions to datasets, and use those distributions to predict event likelihoods.
- Execute data science methods using parallel computing frameworks such as deeplearning4j, Torch, TensorFlow, Caffe, Neon, NVIOFFICE CUDA Deep Neural Network library (cuDNN), and OpenCV, as well as distributed data processing frameworks (e.g., Hadoop, including HDFS, HBase, Hive, Impala, Giraph, Sqoop, Spark, including MLlib, GraphX, SQL, and DataFrames).
- Execute data science methods using common programming and scripting languages: Python, Java, Scala, and R (statistics).
Qualifications
- Skilled in data visualization and use of graphical applications such as Microsoft Power BI and Tableau.
- Proficient in major data science languages such as R and Python, and in managing and merging disparate data sources, preferably through R, Python, or SQL.
- Experience with large data multi‑INT analytics, machine learning, and automated predictive analytics.
- Ability to build AI tools and automated processes such as recommendation engines or automated lead‑scoring systems.
- Strong background in statistical analysis and data mining algorithms.