Haystack
← Back to Jobs
Technology

Data Analyst

Engineering SquareRichmond, VA🇺🇸United StatesPosted 7 Aug 2026

Quick Overview

Work Type
Hybrid
Level
Mid Senior

Job Description

Prefer Prior 1 cap or Strong Financial exp is needed

Role: Data Analyst
Job Location – Richmond VA – Hybrid – 3 days week(Mandatory)

Req - 104908-1

Need  strong on Pyspark

Need the skill summary for every candidate and also need to schedule a screening interview with me before submission on a teams video

Skills set



Skill

Experience

Self-Rating

PySpark

 

 

Python

 

 

Databricks

 

 

SQL

 

 

AWS (S3, Redshift)

 

 

Amazon QuickSight

 

 

Snowflake

 

 

 

Must Skills required

Pyspark- very strong – all the coding questions are on Pyspark , Python – Need to be very strong
Databricks
SQL  
AWS
Amazon Quick Sight dashboards.
Snowflake
 Interviews- 1 rounds of interview coding question on Pyspark & python , SQL , Amazon quicksight , Databricks

 Location:  Richmond VA-hybrid

Duration -3 months with possibility with extension as project will go up to more than 12 months

 This role is a part of the scanner mediation team which helps in converting large validated data from one infrastructure to another, scanning the HSA information with large amount of data .This  team has 2 developers /2 analysts and 3rd role will be reporting to the HM, partner with the partnership  team, daily stands up .


1. Data Ingestion & Missing File Reconciliation
Monitor automated enterprise notifications ensuring that structural migrations are captured and incoming datasets are mapped for assessment.
Analyze active AWS S3 storage buckets to calculate full counts of active data partition files.
Cross-reference S3 partition metrics against Snowflake metadata logging tables to run missing part-file reconciliation logic.
Identify missing scans or pipeline discrepancies and autonomously coordinate with the Enterprise SCAN team to initiate targeted ad-hoc data scanning requests.
2. Deep-Dive Vulnerability Analysis & Triage
Extract active security violation events from complex Snowflake analytical views
Build programmatic pivot metrics and classification profiles to map an exhaustive list of unique combinations of sensitive data types across multiple schema extensions.
Perform context-aware analysis inside Databricks platforms to evaluate unmasked validation values against underlying transactional metadata layouts.
Execute card data logic validation routines, validating raw payloads against the Luhn algorithm within our proprietary toolset.
Triage findings meticulously into defined structural buckets: True Positive, False Positive, or Unsure categories based on core risk profiles.
3. Compliance Tracking & Governance Engineering
Manage structural data dictionaries and trackers, including the DSF Backbook Migration Log and the Sensitive Data Classification Validation Log.
Isolate, format, and push verified false positives directly to data management platforms to run bulk suppression loads, systematically immunizing pipelines against redundant security alert fatigue.
Synthesize granular validation evidence sheets mapping schema properties, target data types, remediation rationale, and reference context profiles.
4. Remediation Validation & Dashboarding
Partner closely with data engineering operators to track the health of automated end-to-end data remediation routines executing via Databricks jobs.
Ensure strict risk isolation logic is sustained, confirming that downstream consumption layer systems remain locked until rescan validations return clean.
Perform end-to-end file auditing to verify that post-masked outputs generated inside targeted write paths precisely mirror input baseline file tallies.
Leverage Amazon QuickSight business intelligence environments to monitor visual tracking boards, ensuring zero data leakage or partition drops across integration scopes.
5. Cross-Functional Stakeholder Alignment
Prepare and frame analytical findings packets ahead of critical technical reviews.
Champion insights and drive live consensus calls within cross-functional operational meetings, including Daily Standups and Bi-Weekly SME Alignment Forums with Chief Data Office (CDO) leads and external integration points.
Technical Skills & Qualifications
Data Infrastructure Mastery: Extensive hand-on experience writing complex querying logic over structured and semi-structured architectures within Snowflake or equivalent enterprise cloud data platforms.
Advanced Data Processing Knowledge: Proven background working within Databricks compute frameworks to explore, extract, and inspect underlying enterprise code bases or dataset tables.
Cloud Architecture Fluency: Direct technical comfort querying, calculating, and inspecting cloud objects directly inside AWS S3 environments.
Data Visualization & Delivery: Experience configuring access and generating scannable metrics reporting within Amazon QuickSight dashboards.
Data Cleansing Logic: Deep comprehension of programmatic data filtration approaches, deduplication routines, and standard mathematical string validations (e.g., Luhn check patterns).
Agile Communications Delivery: Exceptional technical writing capacity to compile audit logs, create operational Markdown playbooks, and effectively drive multi-organizational daily tracking forums.

 

Skills

SQL
AWS
Snowflake
Agile
Databricks
Python
Redshift

Similar jobs