Haystack
← Back to Jobs
Technology
HT

Data Engineer with Gen AI

Hexaware Technologies, IncAtlanta, GA🇺🇸United StatesPosted Sep 16, 2026

Why This Role Stands Out

This hybrid role offers a fantastic opportunity to shape the future of Generative AI solutions, developing cutting-edge data platforms and working with leading technologies. You'll thrive here if you are a skilled Data Engineer passionate about building robust data pipelines and eager to contribute to innovative AI initiatives within a reputable company.

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Atlanta, GA, United States
Posted
23 hours ago
DockerSQLAWSETLMachine LearningSnowflakeAirflowApacheApache SparkAzureBigQueryDatabricksGenerative AIGitGoogle CloudKubernetesLLMPythonRedshift

Job Description

Data Engineer Generative AI

We are seeking a skilled Data Engineer with Generative AI experience to design, build, and optimize scalable data platforms that support AI and analytics initiatives. This role will focus on developing reliable data pipelines, preparing high-quality datasets for large language model applications, and enabling secure, production-ready GenAI solutions.

Key Responsibilities

  • Design, develop, and maintain scalable batch and real-time data pipelines.
  • Build and optimize data ingestion, transformation, validation, and orchestration processes.
  • Develop data models and curated datasets for analytics, machine learning, and Generative AI use cases.
  • Support GenAI applications by preparing, chunking, embedding, indexing, and retrieving enterprise data for Retrieval-Augmented Generation (RAG) workflows.
  • Integrate data sources such as relational databases, APIs, data lakes, document repositories, and streaming platforms.
  • Work with vector databases and embedding models to enable semantic search and GenAI knowledge retrieval.
  • Implement data quality checks, metadata management, lineage, monitoring, and alerting.
  • Partner with data scientists, AI engineers, architects, and business stakeholders to translate requirements into scalable data solutions.
  • Ensure data security, privacy, governance, and access controls are applied across pipelines and AI datasets.
  • Optimize pipeline performance, storage costs, and query efficiency.
  • Document data architecture, pipeline designs, data mappings, and operational procedures.

Required Qualifications

  • Bachelor s degree in Computer Science, Engineering, Data Science, or a related field, or equivalent practical experience.
  • Strong experience in data engineering, ETL/ELT development, and data warehousing.
  • Proficiency in Python and SQL.
  • Experience with data processing frameworks such as Apache Spark, PySpark, Databricks, or similar tools.
  • Experience with cloud data platforms such as AWS, Azure, or Google Cloud Platform.
  • Hands-on experience with data lakes, lakehouses, or cloud warehouses such as Snowflake, Databricks, BigQuery, Redshift, or Synapse.
  • Familiarity with workflow orchestration tools such as Apache Airflow, Azure Data Factory, AWS Glue, or similar platforms.
  • Understanding of Generative AI concepts, including LLMs, embeddings, prompt engineering, vector search, and RAG architectures.
  • Experience integrating APIs and working with semi-structured and unstructured data, including JSON, PDFs, documents, and text files.
  • Strong problem-solving, communication, and collaboration skills.

Preferred Qualifications

  • Experience with vector databases such as Pinecone, Weaviate, Chroma, FAISS, Azure AI Search, or OpenSearch.
  • Experience with GenAI frameworks such as LangChain, LlamaIndex, Semantic Kernel, or similar tools.
  • Familiarity with LLM platforms and services such as Azure OpenAI, Amazon Bedrock, Google Vertex AI, or open-source models.
  • Experience with data governance, cataloging, master data management, and data quality tools.
  • Knowledge of DevOps and CI/CD practices, including Git, Docker, Kubernetes, and automated deployment pipelines.
  • Experience in implementing data masking, PII detection, access controls, and responsible AI practices.
  • Domain experience in financial services, healthcare, retail, or another regulated industry.

Key Skills

  • Python, SQL, PySpark, Spark
  • ETL/ELT, Data Modeling, Data Warehousing
  • Data Lakes, Lakehouse Architecture, Data Quality
  • Apache Airflow, Databricks, Snowflake
  • AWS, Azure, or Google Cloud Platform
  • LLMs, RAG, Embeddings, Vector Databases
  • API Integration, Semantic Search, Unstructured Data Processing
  • Data Governance, Security, and Privacy

Similar jobs