Haystack
← Back to Jobs
Remote
Technology

AWS Data Engineer (Exp- 14+ Years) with AI Exp-Full time-Remote

Visionary Innovative Technology SolutionsUnited States🇺🇸United StatesPosted 4 Aug 2026

Quick Overview

Work Type
Remote
Level
Mid Senior

Job Description

We are seeking a highly experienced Senior AWS Data Engineer with AI/ML experience to design, develop, and optimize enterprise-scale data platforms and pipelines on AWS. The ideal candidate should have strong hands-on expertise in Python, AWS data services, data engineering, ETL/ELT, data lakes, data warehouses, and AI/ML integration.

Key Responsibilities

  • Design and develop scalable data pipelines and data engineering solutions on AWS.
  • Build and maintain robust ETL/ELT pipelines using Python and AWS services.
  • Develop data ingestion, transformation, validation, and processing frameworks.
  • Design and implement enterprise AWS Data Lake and Data Warehouse architectures.
  • Work with AWS services such as S3, Glue, Lambda, EMR, Redshift, Athena, Kinesis, Step Functions, and ECS/EKS.
  • Develop complex data processing solutions using Python, PySpark, and SQL.
  • Design data models for structured, semi-structured, and unstructured datasets.
  • Optimize data pipelines for performance, scalability, reliability, and cost.
  • Implement data quality, governance, lineage, validation, and monitoring processes.
  • Develop orchestration workflows using Apache Airflow / AWS MWAA.
  • Build CI/CD pipelines for data engineering applications and infrastructure.
  • Collaborate with architects, data scientists, ML engineers, application developers, and business stakeholders.

AI / ML Responsibilities

  • Integrate AI/ML capabilities into enterprise data platforms.
  • Build data pipelines supporting Machine Learning and Generative AI use cases.
  • Prepare, clean, transform, and engineer datasets for AI/ML models.
  • Work with LLMs, embeddings, vector databases, and RAG pipelines.
  • Experience with Amazon Bedrock and foundation models is highly desirable.
  • Develop AI-enabled data processing and automation solutions.
  • Integrate pre-trained AI/ML models through APIs.
  • Support model training, evaluation, deployment, and monitoring workflows.
  • Use Python, Pandas, NumPy, Scikit-learn, and other AI/ML libraries as needed.
  • Work with AI coding assistants such as GitHub Copilot, Amazon Q Developer, or Cursor.

Required Skills

  • 14+ years of overall IT experience with strong Data Engineering background.
  • Extensive hands-on experience with AWS Data Engineering.
  • Strong Python programming skills.
  • Advanced SQL skills including complex queries, CTEs, window functions, and optimization.
  • Strong experience with AWS S3, Glue, Redshift, Lambda, EMR, and Athena.
  • Strong experience designing ETL/ELT pipelines.
  • Hands-on experience with PySpark / Apache Spark.
  • Experience with data lake and data warehouse architecture.
  • Experience with Airflow / AWS MWAA.
  • Strong understanding of data modeling and database concepts.
  • Experience with Git and CI/CD.
  • Strong understanding of cloud security, IAM, encryption, and AWS best practices.

AI/GenAI Skills – Preferred

  • Generative AI / LLM experience.
  • Amazon Bedrock.
  • RAG architecture.
  • Prompt Engineering.
  • Embeddings and Vector Databases.
  • LangChain / LlamaIndex.
  • OpenAI / Azure OpenAI / Anthropic.
  • Pinecone / OpenSearch / Elasticsearch / FAISS.
  • Machine Learning and model integration.
  • AI/ML API integration.

Skills

SQL
AWS
ETL
Encryption
Machine Learning
NumPy
Scikit-learn
Airflow
Apache
Apache Spark
Azure
Generative AI
Git
LLM
Pandas
Python
Redshift

Similar jobs