Haystack
← Back to Jobs
Employee
Technology

Data Scientist (Entry Level to SME) TS/SCI with Poly REQUIRED with Security Clearance

CGIArlington, VA🇺🇸United StatesPosted 13 Aug 2026

Quick Overview

Salary
$89.6k - $204k/yr
Work Type
Hybrid
Schedule
Employee
Level
Mid Senior

Job Description

Position Description: CGI Federal has an exciting opportunity for Data Scientists within our Intel sector advancing the national security mission through cutting edge technology. You must have a passion for keeping pace with rapidly evolving technology advancements and leveraging your knowledge on a highly collaborative team to deliver state-of-the-art capabilities. The Data Scientist supports the analytics team by cleaning messy data, running queries, and helping to identify business trends.

They focus on foundational tasks like Exploratory Data Analysis (EDA) and assist senior staff in preparing datasets for predictive models. CGI Federal is growing its high-performance team whose members share a passion for building high-quality, scalable, advanced IT solutions in a collaborative, fast-paced, outcome-driven mission. This position is located in our Arlington office; however, a hybrid working model is acceptable. Your future duties and responsibilities:

Key Responsibilities (Entry Level)

  • Data Preparation: Use SQL to extract raw data and Python or R to clean and format datasets.
  • Exploratory Analysis: Analyze data to find trends, patterns, and anomalies.
  • Visualization: Create basic charts and dashboards using tools like Tableau or Python libraries to present findings to the team.
  • Model Assistance: Help senior data scientists train, test, and evaluate basic machine learning algorithms.

Key Responsibilities (Junior Level)

  • Data Wrangling: Clean, process, and validate raw structured and unstructured data to ensure uniformity and accuracy.
  • Exploratory Data Analysis (EDA): Analyze data to identify trends, patterns, and anomalies.
  • Modeling: Assist in developing, testing, and updating basic machine learning models and statistical algorithms.
  • Visualization & Communication: Build dashboards and presentations to clearly communicate findings and recommendations to both technical and non-technical stakeholders.
  • Pipeline Maintenance: Collaborate with data engineers and senior data scientists to maintain and optimize data pipelines

Key Responsibilities (Mid-Level)

  • Model Development & Deployment: Design, train, evaluate, and deploy robust machine learning and deep learning models to solve ambiguous business problems.
  • Advanced Analytics: Conduct rigorous exploratory data analysis (EDA) and apply complex statistical techniques (e.g., A/B testing, regression analysis, clustering) to extract deep insights.
  • Data Engineering & Pipelines: Extract, clean, and manipulate unstructured datasets across distributed systems. Contribute to the design and optimization of data pipelines.
  • Stakeholder Collaboration: Translate high-level business goals into precise data science requirements. Present actionable recommendations to both technical and non-technical stakeholders.
  • Technical Leadership: Act as a subject matter expert and mentor junior analysts or entry-level data scientists on methodology and coding best practices.

Key Responsibilities (Senior Level)

  • Advanced Modeling: Architect and deploy Deep Learning (DL), Natural Language Processing (NLP), and Large Language Models (LLMs).
  • System Scalability: Build and optimize distributed data processing pipelines (e.g., using Spark) and automate reproducible workflows.
  • Translational Strategy: Convert ambiguous, high-dimensional business/mission requirements into strict technical requirements.
  • Leadership & Mentorship: Lead end-to-end projects autonomously and mentor junior data scientists and engineers

Key Responsibilities (SME Level)

  • Advanced AI & Model Architecture: Architect and operationalize complex Agentic AI systems and customized Retrieval-Augmented Generation (RAG) frameworks to support dynamic domain requirements.
  • Constrained Environment Optimization: Optimize resource-heavy machine learning models and deep neural networks to perform reliably at the edge or within restricted computing ecosystems.
  • Unstructured Data Synthesis: Develop reproducible analytical models and quantitative techniques to draw definitive conclusions from incomplete, noisy, or highly unstructured raw datasets.
  • Pipeline Automation: Design and automate scalable data pipelines, applying CI/CD principles to transition experimental prototypes seamlessly into enterprise production environments.
  • Statistical Rigor: Perform exhaustive statistical modeling, hypothesis testing, inference, and probabilistic forecasting using multi-variate domain knowledge. Required qualifications to be successful in this role: All levels require an active TS/SCI with Poly Entry Level:
  • Education: o High School Diploma/GED with 4 years of relevant experience, o Associates Degree with 2 years of relevant experience, o or Bachelor's Degree with 0 years of relevant experience
  • Programming: Working knowledge of Python, R, or SQL.
  • Math & Stats: Basic understanding of probability, descriptive statistics, and linear algebra.

Junior Level/Moderate:

  • Education: o High School Diploma/GED with 6 years of relevant experience, o Associates Degree with 4 years of relevant experience, o Bachelor's Degree with 2 years of relevant experience, o or Masters Degree with 0 years of relevant experience
  • Programming Skills: Strong proficiency in querying databases using SQL and coding in Python or R.
  • Statistical Knowledge: Foundational understanding of descriptive statistics, probability, and hypothesis testing.
  • Visualization Tools: Familiarity with BI and charting tools like Tableau, Power BI, or Python libraries (e.g., Matplotlib, Seaborn) Mid-Level/Complex:
  • Education: o High School Diploma/GED with 8 years of relevant experience, o Associates Degree with 6 years of relevant experience, o Bachelor's Degree with 4 years of relevant experience, o or Masters Degree with 2 years of relevant experience, o or PhD with 0 years of relevant experience
  • Programming: Advanced proficiency in Python or R, and mastery of SQL for data extraction and manipulation.
  • Machine Learning Libraries: Hands-on experience with core ML frameworks such as scikit-learn, TensorFlow, PyTorch, or XGBoost.
  • Data & Visualization: Experience using BI tools (e.g., Tableau, Power BI) and libraries like pandas, NumPy, matplotlib, or seaborn.
  • Foundations: Strong theoretical and practical understanding of statistical distributions, probability, and experimental design.

Senior Level/Exceptionally Complex:

  • Education: o High School Diploma/GED with 10 years of relevant experience, o Associates Degree with 8 years of relevant experience, o Bachelor's Degree with 6 years of relevant experience, o or Masters Degree with 4 years of relevant experience, o or PhD with 2 years of relevant experience
  • Technical Stack: Advanced proficiency in Python, R, and SQL. Hands-on experience with cloud computing platforms (AWS, Azure) and Big Data environments.

SME Level/Exceptionally Complex:

  • Education: o High School Diploma/GED with 12 years of relevant experience, o Associates Degree with 10 years of relevant experience, o Bachelor's Degree with 8 years of relevant experience, o or Masters Degree with 6 years of relevant experience, o or PhD with 4 years of relevant experience
  • Languages & Environments: Expert-level proficiency in Python (including advanced use of Pandas, NumPy, and PyTorch/TensorFlow) and experience working natively in environments like Jupyter Notebooks and Databricks.
  • Foundational Math: Deep understanding of multi-variable calculus, linear algebra, and advanced probability/statistics.

AI/ML Methodologies: Extensive hands-on background in constructing sophisticated algorithms, including Large Language Models (LLMs), natural language processing, and deep learning architectures.

  • Data Infrastructure: Advanced ability to clean, curate, and manipulate massive petabyte-scale datasets spanning multiple relational, NoSQL, and big data architectures.
  • Domain Context: Experience in highly specialized domains such as SIGINT/RF processing, quantitative finance, or complex biomedical analytics (dependent on industry). CGI is required by law in some jurisdictions to include a reasonable estimate of the compensation range for this role. The determination of this range includes various factors not limited to skill set, level, experience, relevant training, and licensure and certifications. To support the ability to reward for merit-based performance, CGI typically does not hire individuals at or near the top of the range for their role. Compensation decisions are dependent on the facts and circumstances of each case. A reasonable estimate of the current range for this role in the U.S. is $89,600.00 - $204,000.00. CGI Federal's benefits are offered to eligible professionals on their first day of employment to include: . Competitive compensation
. Comprehensive insurance options
. Matching contributions through the 401(k) plan and the share purchase plan
. Paid time off for vacation, holidays, and sick time
. Paid parental leave
. Learning opportunities and tuition assistance
. Wellness and Well-being programs #CGIFederalJob
#LI-LB1
#ClearanceJobs
#CGIInternationalSecurity What you can expect from us: Together, as owners, let's turn meaningful insights into action. Life at CGI is rooted in ownership, teamwork, respect and belonging. Here, you'll reach your full potential because... You are invited to be an owner from day 1 as we work together to bring our Dream to life. That's why we call ourselves CGI Partners rather than employees. We benefit from our collective success and actively shape our company's strategy and direction. Your work creates value. You'll develop innovative solutions and build relationships with teammates and clients while accessing global capabilities to scale your ideas, embrace new opportunities, and benefit from expansive industry and technology expertise. You'll shape your career by joining a company built to grow and last. You'll be supported by leaders who care about your health and well-being and provide you with opportuni

Skills

SQL
AWS
Linear
Machine Learning
NLP
NumPy
Scikit-learn
Tableau
Azure
Databricks
Deep Learning
Jupyter
Pandas
Power BI
PyTorch
Python
TensorFlow

Similar jobs