Haystack
← Back to Jobs
Technology
CI

Big Data Engineer

ComTec Information SystemsMcLean, VA🇺🇸United StatesPosted 14 Sept 2026

Why This Role Stands Out

This hybrid role offers a fantastic opportunity to leverage your expertise in Spark, Hadoop, and cloud technologies within a reputable company, fostering significant career growth. You'll thrive here if you're a skilled Big Data Engineer with a strong background in complex SQL, scripting, and cloud platforms, ready to contribute to innovative data solutions. Apply today to join a dynamic team and expand your technical capabilities!

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
McLean, VA, United States
Posted
Yesterday
SQLScalaAWSAgileData PipelineHadoopHivePerlPython

Job Description

Title: Big Data Engineer

Location: Rockville, MD or McLean, VA (Hybrid)

Contract: 6+ Months Contract

 

Only Local candidates who can take Assessment and only who are in DC/VA/MD who can got for F2F interview

 

Overview

  • Must haves: spark, Hadoop, scala, hive
  • scripting is a must- python or perl
  • must be expert level in Complex SQL- window functioning, complex multiple joins, cloud experience is mandatory-S3, glue, emr, athena
  • AI- How to use AI for prompt engineering
  • Github
  • Copiliot
  • Chjatgopt
  • Q

 

Need someone who is well versed in agile, test automations, CICD practices

Financial experience is preferred

ROLE FIT

  • 5+ years building enterprise-scale data solutions using Spark, Hadoop, Hive, and Scala
  • Strong scripting skills (Python or Perl) and expert-level complex SQL (window functions, multi-joins)
  • AWS cloud experience required (S3, EMR, Glue, Athena)
  • Experience with Agile delivery, CI/CD pipelines, automated testing, and GitHub workflows
  • Financial services or regulated industry experience preferred

OBJECTIVES

  • Design and maintain scalable, reliable big data pipelines
  • Optimize Spark/Hadoop workloads for performance, scalability, and cost efficiency
  • Implement automated testing and data quality validation
  • Enable analytics and data science teams with high-quality, accessible datasets
  • Leverage AI-assisted tools (Copilot, ChatGPT, Q Developer) to improve development productivity

PROBLEM-SOLVING

  • Diagnose and resolve Spark performance bottlenecks and data pipeline failures
  • Optimize complex SQL transformations and large-scale joins
  • Troubleshoot data quality, latency, and reliability issues in production
  • Improve AWS workload efficiency through tuning and resource optimization
  • Automate repetitive engineering tasks using AI-assisted development tools

EXPERIENCE VALIDATION

  • Delivered end-to-end pipelines using Spark and Hadoop ecosystem tools
  • Optimized SQL and pipeline performance with measurable improvements
  • Deployed and supported AWS data workloads (EMR, Glue, Athena, S3)
  • Implemented CI/CD and automated testing for data pipelines
  • Used AI coding assistants and GitHub workflows in team-based development

SKILLS TO PROBE
Spark internals, performance tuning, and partitioning strategies
Advanced SQL techniques and Hive/Trino optimization
AWS data architecture and cost/performance tuning practices
Prompt engineering and safe use of AI coding assistants
Collaboration and communication in fast-paced, cross-functional environments'

Similar jobs