Haystack
← Back to Jobs
Technology
RI

Data Platform Engineer

Rivago infotech incCharlotte, NC🇺🇸United StatesPosted Sep 30, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Charlotte, NC, United States
Posted
18 hours ago
SonarQubeApacheBigQueryGitHub ActionsGoGoogle CloudHadoopHiveJenkinsKafkaLLMPythonREST

Job Description

Role : Data / Data Platform Engineer (4 Open position)

Location : Charlotte, NC (Hybrid)

Role Summary

We are seeking an experienced Data / Data Platform Engineer to join the Enterprise Security Transformation initiative. In this role, you will design, develop, and optimize high-throughput data processing pipelines (batch and real-time streaming) that ingest, normalize, enrich, and deduplicate vulnerability scan outputs generated by enterprise LLM security scanner. You will also build an extensible framework to rapidly onboard new scanning models and assist with migrating on-premise data pipelines to Google Cloud Platform (Google Cloud Platform).


Key Responsibilities

  • Pipeline Development & Optimization: Design and implement robust batch and real-time streaming pipelines using PySpark, Hadoop, Hive, and Apache Kafka to process large volumes of security vulnerability findings.
  • Data Cleansing & Deduplication: Implement advanced data quality, enrichment (severity, criticality scoring), and deduplication logic comparing multiple scanning feeds (e.g., LLM scanners vs. Checkmarx/Snow AR) to eliminate duplicate findings.
  • Cloud Migration & Data Lake Integration: Ingest and process data within Google Cloud Platform (Google Cloud Platform) leveraging Google Cloud Platform Data Lake, BigQuery, and modern cloud storage/compute frameworks, facilitating the transition from on-premise processing to Google Cloud Platform.
  • Extensible Scanning Framework: Build a modular, pluggable framework that enables the security team to integrate, evaluate, and compare output results from newly introduced LLM/security scanner models (e.g., Mythos, Codex, Astra) within days.
  • ServiceNow & Remediation Workflows: Format, validate, and deliver clean, actionable data to downstream systems.

Required Technical Skills & Qualifications

  • Programming Languages: Proficiency in Python and Go (Golang).
  • Big Data & Streaming Technologies: Hands-on experience with PySpark, Apache Hadoop, Apache Hive, and Apache Kafka (both batch and real-time/event-driven data pipelines).
  • Google Cloud Platform (Google Cloud Platform): Proven experience with BigQuery, Google Cloud Platform Data Lake architectures, and Google Cloud Platform data services.
  • Data Processing & Architecture: Demonstrated experience building data quality validation, cross-source deduplication, data enrichment, and schema standardization pipelines.
  • Tools & Integration: Experience integrating data pipelines with enterprise REST APIs, messaging systems, and security/ITSM platforms.

Preferred Skills & Nice-to-Haves

  • Prior domain experience in Vulnerability Management, Application Security, or Cyber Security Data Analytics.
  • Familiarity with Code Scanning / SAST / DAST tools (e.g., Checkmarx) and ServiceNow modules (Vulnerability Response / SecOps / Snow AR).
  • Experience handling enterprise-scale code repository metadata (GitHub/GitLab pipelines).
  • Financial services or highly regulated enterprise environment experience.

 

Preferred Qualifications

•     Certified Information Systems Security Professional (CISSP), Certified Application Security Engineer (CASE), or CSSLP.

•     Experience working in a large enterprise banking environment with multiple business units.

•     Familiarity with DevSecOps toolchains (Jenkins, GitHub Actions, SonarQube, Veracode, Checkmarx, Snyk).

Similar jobs