Haystack
← Back to Jobs
Remote
Technology
GR

Senior Data Engineer - Full Time Only - Remote

GD Resources LLCUnited States🇺🇸United StatesPosted Sep 30, 2026

Quick Overview

Salary
$140k/yr
Seniority
Mid Senior
Work mode
Remote
Location
United States
Posted
Yesterday
AWSETLAirflowCDKData PipelineDatabricksJiraLLMRedshiftTerraformdbt

Job Description

Senior Data Engineer

Salary: $140,000.00 Annually

Remote

 Only Full Time consultants will consider for this role..Profile has to match all the required skills.

Key Responsibilities

Data Pipelines & Platforms

  • Lead pipeline architecture decisions for assigned workstreams, including batch and real-time ingestion, transformation, and delivery to analytics and AI layers.
  • Design, build, and maintain scalable ETL/ELT workflows with strong data quality, lineage, monitoring, and observability built in from the start.
  • Establish data contracts and pipeline standards for the projects you own, and ensure those standards are followed by other contributors on the workstream.
  • Ensure data reliability, performance, and scalability across platforms; proactively surface bottlenecks and recommend improvements before they affect delivery.
  • Support both batch and streaming ingestion patterns, including real-time data pipeline design and implementation.

LLM Enablement & AI Data Foundations

  • Design and implement data pipelines that support LLM and AI use cases, including:
  • Document and unstructured data ingestion
  • Data preprocessing, enrichment, and embedding generation
  • Vector store integration and retrieval-optimized data structures
  • Lead embedding pipeline architecture and vector store configuration in collaboration with developers building LLM features.
  • Ensure data freshness, lineage, and governance for AI-powered systems.
  • Optimize data structures and retrieval patterns to support efficient LLM context usage.
  • Contribute to RAG pipeline design and maintain awareness of foundation model data requirements across active engagements.

Cloud Architecture (AWS)

  • Architect and provision AWS services to support data and AI workloads, including S3, Glue, Glue Catalog, Redshift, Athena, EMR, Kinesis, MSK, Lambda, Step Functions, and EventBridge.
  • Lead FedRAMP-compliant architecture design for data environments where required; apply security and access control patterns using Lake Formation and IAM.
  • Contribute to reusable reference architectures for data lakes, warehouses, streaming systems, and AI-ready platforms.
  • Partner with platform and DevOps teams to ensure secure, cost-effective, and scalable cloud deployments.
  • Apply infrastructure-as-code and automation practices to data platform provisioning and maintenance.

Client & Stakeholder Engagement

  • Participate in client discovery and requirements-gathering sessions, translating operational needs into concrete data architecture recommendations.
  • Communicate clearly with both technical and non-technical stakeholders, adapting depth and language to the audience without losing precision.
  • Assess and document client data readiness for analytics and AI adoption; identify gaps and propose remediation paths to the lead architect or engagement manager.
  • Support the technical narrative during delivery, contributing to solution design documents, architecture diagrams, and client-facing documentation.
  • Build working trust with client counterparts through consistency, follow-through, and clear expectation-setting.

Technical Delivery Leadership

  • Own the data engineering workstream within a project delivery plan, including task decomposition, estimation, and sequencing in Jira.
  • Understand how individual tickets connect to the larger delivery arc, and surface dependencies or risks before they block progress.
  • Facilitate technical planning and review ceremonies for the data workstream; provide clear updates to project managers and technical leads.
  • Partner with project managers and product owners to keep the data engineering track aligned with contractual and delivery constraints.
  • Coordinate across engineering, platform, analytics, and ML teams to ensure data pipelines meet downstream requirements.

Product & Data Readiness Support

  • Support the evolving data architecture behind product capabilities, including predictive and real-time ML systems.
  • Assess and improve internal and client data readiness for analytics and AI adoption.
  • Translate business, product, and client needs into scalable data architectures that can be maintained and extended by the broader team.
  • Document data architectures, pipeline designs, and integration patterns to support transparency and reuse across engagements.

Staff Development & Knowledge Sharing

  • Mentor junior and core-level data engineers through code reviews, architecture critiques, and hands-on guidance on active projects.
  • Identify skill gaps in team members and work with technical leadership to address them through structured coaching or pairing.
  • Contribute to internal playbooks, onboarding materials, and engineering standards that reduce tribal knowledge and improve team consistency.
  • Help team members understand the business context behind their work, connecting individual tasks to client outcomes and technical strategy.

Collaboration & Communication

  • Collaborate closely with developers building LLM features to ensure data pipelines meet AI requirements.
  • Work with product, analytics, and technical leadership to align data strategy with organizational and project goals.
  • Communicate findings, risks, and architectural decisions clearly in both written and verbal form, including client-facing documentation.
  • Contribute to proposal efforts and technical volume sections as a subject matter contributor.

Qualifications

Required

  • Bachelor's degree or equivalent experience in data engineering, computer science, or a related field.
  • 5+ years of hands-on data engineering experience, with demonstrated progression into senior or lead responsibilities.
  • Strong hands-on experience with AWS data services, including S3, Glue, Redshift, Athena, Kinesis or MSK, Lambda, and Lake Formation.
  • Proven ability to design and deliver production-grade ETL/ELT pipelines and data warehousing solutions.
  • Solid understanding of streaming architectures and real-time data pipeline design.
  • Experience supporting AI or LLM-adjacent data workflows, including embedding pipelines and vector store integration.
  • Ability to communicate clearly with both technical and non-technical stakeholders, including direct client interaction.
  • Experience decomposing and tracking delivery work in Jira or equivalent tooling.
  • Strong problem-solving skills and architectural reasoning at the workstream level.

Preferred

  • AWS certifications: Solutions Architect Associate (SAA-C03), Data Engineer Associate (DEA-C01), or equivalent.
  • Databricks Certified Data Engineer Associate and/or dbt Certified Developer.
  • Experience with infrastructure-as-code tools such as Terraform or CDK.
  • Background in consulting, professional services, or multi-client delivery environments.
  • Familiarity with data governance frameworks, data cataloging, and lineage tooling.
  • Experience with Databricks, dbt, and Airflow or Prefect orchestration.
  • Exposure to vector databases such as Amazon OpenSearch, Pinecone, or pgvector.
  • Exposure to DoD, federal, or regulated-sector data environments; FedRAMP-compliant architecture experience a plus.

Similar jobs