Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
McLean, VA, United States
Posted
20 hours ago
SQLScalaAWSETLScrumAgileApacheApache SparkHadoopHiveJavaKanbanPython
Job Description
Responsibilities:
- Design, develop, and maintain large-scale data processing pipelines using Big Data technologies (e.g., Hadoop, Spark, Python, Scala).
- Implement data ingestion, storage, transformation, and analysis of solutions that are scalable, efficient, and reliable.
- Stay current with industry trends and emerging Big Data technologies to continuously improve the data architecture
- Collaborate with cross-functional teams to understand business requirements and translate them into technical solutions.
- Optimize and enhance existing data pipelines for performance, scalability, and reliability.
- Develop automated testing frameworks and implement continuous testing for data quality assurance.
- Conduct unit, integration, and system testing to ensure the robustness and accuracy of data pipelines.
- Work with data scientists and analysts to support data-driven decision-making across the organization.
- Ability to write and maintain automated unit, integration, and end-to-end tests
- Monitor and troubleshoot data pipelines in production environments to identify and resolve issues.
Education/Experience Requirements:
- Bachelor's degree in Computer Science, Information Systems or related discipline with at least five (5) years of related experience, or equivalent training and/or work experience; Master's degree and past Financial Services industry experience preferred.
- Demonstrated technical expertise in Object Oriented and database technologies/concepts which resulted in deployment of enterprise quality solutions.
- Past experience with developing enterprise quality solutions in an iterative or Agile environment.
- Extensive knowledge of industry leading software engineering approaches including Test Automation, Build Automation and Configuration Management frameworks.
- Strong written and verbal technical communication skills.
- Demonstrated ability to develop effective working relationships that improved the quality of work products.
- Should be well organized, thorough, and able to handle competing priorities.
- Ability to maintain focus and develop proficiency in new skills rapidly.
- Ability to work in a fast paced environment.
- Experience with object oriented programming languages such as Java, Scala or Python.
Essential Technical Skills:
- AI Tool Proficiency: Hands-on experience with AI development tools (GitHub Copilot, Q Developer, ChatGPT, Claude, etc.)
- Technical Background: Strong software development background with ability to contribute to technical discussions
- Agile Methodology: Extensive experience with Scrum, Kanban, and continuous improvement practices
Big Data technologies
- Experience with Big data technologies such as Hadoop, Spark, Hive & Trino
- Evaluate understanding of common issues like:
o Data skew and strategies to mitigate it.
o Working with massive data volumes in PetaBytes.
o Troublehsooting job failures due to resource limitations, bad data, scalability challenged.
- Look for real-world debugging and mitigation stories.
AI Skills
- Prompt Engineering: Proficiency in crafting effective prompts for AI coding assistants and analysis tools
- AI Workflow Design: Experience redesigning development processes to leverage AI capabilities
- Data Analysis: Ability to interpret AI-generated insights and translate them into actionable team improvements
- Change Management: Experience leading teams through AI adoption and workflow transformation
SQL Skills (Window Functions, Joins, Complex Queries)
- Assess comfort with SQL window functions, multi-table joins, aggregations.
- Provide examples or ask them to write/optimize SQL queries on the spot.
- Probe how they handle edge cases like NULLs, duplicates, ordering, etc.
Apache Spark (Development, Internals & Tuning)
- Test their understanding of Spark s core architecture executors, tasks, stages, DAG.
- Focus on Spark performance tuning techniques: partitioning, caching, broadcast joins, etc.
- Ask scenario-based questions on troubleshooting slow running/stuck jobs or resource issues in Spark.
- Explore their experience optimizing Spark jobs for large-scale datasets.
Cloud Technologies
- Check exposure to AWS services like S3, EMR, Glue, Lambda, Athena, etc.
- Ask how they ve used S3 with Spark (e.g., dealing with file formats, consistency issues).
- EKS, Serverless knowledge, etc
Programming - Python or Scala
- Assess ability to write clean, modular, and performant code.
- Look for experience in functional programming concepts (e.g., immutability, higher-order functions).
- Ask about real-world use cases where they wrote scalable data processing code.
- Evaluate understanding of collections, concurrency, and memory management.
Good to have:
- Experience with managing production data pipelines/ETL systems
- Experience with CI/CD
- Experience writing test cases
- AWS certifications
Similar jobs
- AI
Service Now Developer SIR / VR (SME)
NewAasritha infotech Inc
United States🇺🇸Hybrid20 hours agoRESTTechnology - MU
Full Stack Infrastructure Engineer | NJ/NY – 3 days hybrid
NewMomento USA LLC
New York, NY🇺🇸Hybrid20 hours agoShellAWSSSO+11Technology - LI
Fullstack Engineer- Personalization Platform
LinkedIn
Mountain View, CA🇺🇸Hybrid1 week agoJavaScriptPythonJava+9Technology - TA
Senior Software Engineer, Full Stack
NewTrue Anomaly
Denver🇺🇸23 hours ago401kElixirFiber+19Technology - ST
Senior Software Engineer (Full Stack)
NewStubHub
Aliso Viejo🇺🇸5 hours agoMicroservicesSQLSQL Server+5Technology - BU
FIS Developer
BuzzClan LLC
New York, NY🇺🇸Hybrid3 weeks agoSQLAgileJira+1Technology