Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
Tampa, FL, United States
Posted
Yesterday
DockerAWSETLApacheGenerative AIKubernetesLLMPandasPython
Job Description
OVERVIEW: We are seeking a Data Engineer with strong hands-on experience designing, developing, and managing large-scale data workflows across structured and unstructured datasets. This role focuses heavily on building reliable RAG (Retrieval-Augmented Generation) pipelines, orchestrating ETL/ELT processes, and deploying scalable data systems in AWS.
GENERAL DUTIES
- Design, build, and maintain RAG pipelines, including document ingestion, indexing, embedding workflows, and model retrieval flows.
- Develop and manage structured and unstructured data pipelines supporting analytics, ML, and application workloads. Build and optimize ETL/ELT pipelines in AWS using services such as S3, Lambda, Step Functions, EMR, Glue, ECS/EKS, and IAM best practices. Implement and operate NiFi flows for high-throughput, low-latency data ingestion and transformation.
- Develop, orchestrate, and schedule workflows using Prefect, ensuring reliability, observability, and proper error handling. Implement indexing, search, and retrieval patterns using ElasticSearch, including schema design, cluster management, and query optimization.
- Collaborate closely with architecture, ML, and application teams to support scalable data solutions.
- Ensure data quality, lineage, governance, and security across all pipelines.
- Monitor system performance and troubleshoot issues across distributed data systems.
REQUIRED QUALIFICATIONS
- Solid understanding of ETL/ELT processes and data modeling best practices.
- Hands-on experience implementing workflows in Prefect (Prefect 2.0 preferred). In-depth knowledge of ElasticSearch indexing, cluster management, and search optimization.
- Proficiency in Python and familiarity with common data libraries (Pandas, PySpark, requests, etc.).
DESIRED QUALIFICATIONS
- Strong experience with Apache NiFi for data flow management and real-time ingestion.
- Experience building or maintaining RAG pipelines (e.g., vector databases, embeddings, document chunking strategies, retrieval optimization).
- Proven ability to manage structured and unstructured data pipelines at scale.
- Experience with AWS cloud services for data engineering.
- Strong version control and CI/CD experience.
- Experience with vector databases (OpenSearch, Pinecone, Weaviate, etc.)
- Familiarity with containerized workflows (Docker, Kubernetes)
- Experience supporting LLM or generative AI production systems
- Background in distributed systems, streaming platforms, or data mesh architectures CLEARANCE:
- Active TS/SCI clearance minimum required
Similar jobs
- BI
AI/ML Data Engineer with Security Clearance
NewBigBear.ai
Reston, VA🇺🇸HybridYesterdayGCPSQLAWS+5Technology - CS
Journeyman Data Manager in Arlington, Virginia with Security Clearance
CIS Secure
Arlington, VA🇺🇸On-site2 weeks agoSQLMicrosoft OfficePower BI+1Technology - GI
Database Administrator with Security Clearance
NewGridiron IT Solutions
Washington, DC🇺🇸$120k - $130k/yrHybridYesterdayOracleAWSGrafana+2Technology - PS
4531 Data Engineer with Security Clearance
NewProcession Systems
Suitland, MD🇺🇸HybridYesterdayGCPSQLAWS+20Technology - ME
Delivery Engineer, Frontier Data Products
NewMercor
San Francisco🇺🇸12 hours agoEngineering - LA
Sr. Analytics Engineer/ Data Engineer
NewLaurel
San Francisco🇺🇸Remote5 hours agoGCPSQLAWS+17Technology