Quick Overview
Job Description
Role Summary:
We are looking for an experienced AWS Data Engineer to design, build, and optimize scalable batch and real-time data pipelines that process high-volume data, preferably within the Communications domain (e.g., Call Detail Records, network telemetry). The ideal candidate has strong hands-on expertise across Spark, Java, Kafka, and the AWS/Cloudera data lakehouse ecosystem, and is comfortable owning pipelines end-to-end from ingestion and transformation to performance tuning and compliance.
Key Responsibilities
Design, build, and optimize scalable batch and real-time data pipelines using Apache Spark (Scala/Python), Core Java, and Apache Kafka to process high-volume data.
Manage and maintain enterprise data lakehouse architectures spanning AWS and Cloudera environments, leveraging S3, Apache Iceberg, and Hive for scalable storage and querying.
Implement Change Data Capture (CDC) and write complex SQL to integrate data from relational databases (Oracle, SQL Server) into dimensional data warehouses.
Automate and orchestrate pipeline workflows using Unix shell scripting and Python.
Lead and support migration of legacy Hadoop workloads to AWS cloud-native architectures.
Monitor, troubleshoot, and tune JVM and Spark job performance to ensure reliability and efficiency at scale.
Ensure all data pipelines and storage practices adhere to regulatory compliance and data governance standards.
Collaborate with cross-functional teams (data architects, analysts, platform engineers) to translate business requirements into robust data solutions.
Required Skills & Technology:
Strong programming experience in Scala/Python (Spark) and Core Java
Hands-on experience with Apache Kafka for streaming data pipelines
Working knowledge of AWS services (S3) and Cloudera platform tools (Hive, Iceberg)
Strong SQL skills, including CDC concepts, across Oracle and SQL Server
Experience with dimensional data warehouse modeling
Proficiency in Unix shell scripting and Python for automation
Experience migrating on-prem Hadoop workloads to cloud platforms
Understanding of JVM internals and Spark performance tuning techniques
Preferred Qualifications:
Prior experience in the Communications/Telecom domain (CDRs, network telemetry data)
Familiarity with data governance and regulatory compliance frameworks
Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience
Minimum years of experience
5-8 years
Similar jobs
- SG
Staff Data Engineer
NewSouthern Glazer's Wine&Spirits
Dallas, Texas🇺🇸Hybrid27 minutes agoSQLScalaETL+6Technology - PR
Lead Data Engineer
NewProtective
Nebraska🇺🇸$109.5k - $167.8k/yrHybrid28 minutes agoSQLMLOpsMLflow+8Technology - XP
Data engineer
NewXpertiz Inc
Saint Helena, NC🇺🇸Hybrid18 hours agoSQLMLflowAzure+5Technology - II
Senior Data Engineer
NewInnovative IT Solutions Inc
United States🇺🇸Hybrid18 hours agoSQLAWSPython+1Technology - LO
Data Engineer-AWS,Snowflake
NewLogicplanet, Inc.
Plano, TX🇺🇸On-site18 hours agoSQLAWSETL+7Technology - CI
Sr. Data Engineer- Elastic Search
NewCitiusTech
Irving, TX🇺🇸Hybrid18 hours agoDockerFastAPIMicroservices+13Technology