Why This Role Stands Out
This hybrid role offers a fantastic opportunity to shape the future of Generative AI solutions, developing cutting-edge data platforms and working with leading technologies. You'll thrive here if you are a skilled Data Engineer passionate about building robust data pipelines and eager to contribute to innovative AI initiatives within a reputable company.
Quick Overview
Job Description
Data Engineer Generative AI
We are seeking a skilled Data Engineer with Generative AI experience to design, build, and optimize scalable data platforms that support AI and analytics initiatives. This role will focus on developing reliable data pipelines, preparing high-quality datasets for large language model applications, and enabling secure, production-ready GenAI solutions.
Key Responsibilities
- Design, develop, and maintain scalable batch and real-time data pipelines.
- Build and optimize data ingestion, transformation, validation, and orchestration processes.
- Develop data models and curated datasets for analytics, machine learning, and Generative AI use cases.
- Support GenAI applications by preparing, chunking, embedding, indexing, and retrieving enterprise data for Retrieval-Augmented Generation (RAG) workflows.
- Integrate data sources such as relational databases, APIs, data lakes, document repositories, and streaming platforms.
- Work with vector databases and embedding models to enable semantic search and GenAI knowledge retrieval.
- Implement data quality checks, metadata management, lineage, monitoring, and alerting.
- Partner with data scientists, AI engineers, architects, and business stakeholders to translate requirements into scalable data solutions.
- Ensure data security, privacy, governance, and access controls are applied across pipelines and AI datasets.
- Optimize pipeline performance, storage costs, and query efficiency.
- Document data architecture, pipeline designs, data mappings, and operational procedures.
Required Qualifications
- Bachelor s degree in Computer Science, Engineering, Data Science, or a related field, or equivalent practical experience.
- Strong experience in data engineering, ETL/ELT development, and data warehousing.
- Proficiency in Python and SQL.
- Experience with data processing frameworks such as Apache Spark, PySpark, Databricks, or similar tools.
- Experience with cloud data platforms such as AWS, Azure, or Google Cloud Platform.
- Hands-on experience with data lakes, lakehouses, or cloud warehouses such as Snowflake, Databricks, BigQuery, Redshift, or Synapse.
- Familiarity with workflow orchestration tools such as Apache Airflow, Azure Data Factory, AWS Glue, or similar platforms.
- Understanding of Generative AI concepts, including LLMs, embeddings, prompt engineering, vector search, and RAG architectures.
- Experience integrating APIs and working with semi-structured and unstructured data, including JSON, PDFs, documents, and text files.
- Strong problem-solving, communication, and collaboration skills.
Preferred Qualifications
- Experience with vector databases such as Pinecone, Weaviate, Chroma, FAISS, Azure AI Search, or OpenSearch.
- Experience with GenAI frameworks such as LangChain, LlamaIndex, Semantic Kernel, or similar tools.
- Familiarity with LLM platforms and services such as Azure OpenAI, Amazon Bedrock, Google Vertex AI, or open-source models.
- Experience with data governance, cataloging, master data management, and data quality tools.
- Knowledge of DevOps and CI/CD practices, including Git, Docker, Kubernetes, and automated deployment pipelines.
- Experience in implementing data masking, PII detection, access controls, and responsible AI practices.
- Domain experience in financial services, healthcare, retail, or another regulated industry.
Key Skills
- Python, SQL, PySpark, Spark
- ETL/ELT, Data Modeling, Data Warehousing
- Data Lakes, Lakehouse Architecture, Data Quality
- Apache Airflow, Databricks, Snowflake
- AWS, Azure, or Google Cloud Platform
- LLMs, RAG, Embeddings, Vector Databases
- API Integration, Semantic Search, Unstructured Data Processing
Data Governance, Security, and Privacy
Similar jobs
- MS
Data Engineer with Mortgage - Onsite
NewMSYS Inc.
Dallas, TX🇺🇸On-site23 hours agoSQLETLEncryption+5Technology - GI
Data Engineer
NewGurus Infotech, Inc.
New York, NY🇺🇸Hybrid23 hours agoDockerSQLAWS+6Technology - LT
Jr. Data Engineer
NewLedgent Technology
Los Angeles, CA🇺🇸Hybrid23 hours agoSQLETLLooker+10Technology - AK
Data Engineer IV
NewAkidev Corporation
Oakland, CA🇺🇸Hybrid23 hours agoETLPythonTechnology - CI
Azure Data Engineer
NewCitadel Information Services Inc
Charlotte, NC🇺🇸Hybrid23 hours agoSQLETLAzure+5Technology - ES
Systems Analyst- Informatica ETL/Data Engineer (State exp) in Austin, TX- only Local Austin
NewEsolvit, Inc.
Austin, TX🇺🇸Hybrid23 hours agoOracleSQLETL+1Technology