Haystack
← Back to Jobs
Technology
PS

4521 Data Engineer with Security Clearance

Procession SystemsTampa, FL🇺🇸United StatesPosted 15 Sept 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
Tampa, FL, United States
Posted
Yesterday
DockerAWSETLApacheGenerative AIKubernetesLLMPandasPython

Job Description

OVERVIEW: We are seeking a Data Engineer with strong hands-on experience designing, developing, and managing large-scale data workflows across structured and unstructured datasets. This role focuses heavily on building reliable RAG (Retrieval-Augmented Generation) pipelines, orchestrating ETL/ELT processes, and deploying scalable data systems in AWS.

GENERAL DUTIES

  • Design, build, and maintain RAG pipelines, including document ingestion, indexing, embedding workflows, and model retrieval flows.
  • Develop and manage structured and unstructured data pipelines supporting analytics, ML, and application workloads. Build and optimize ETL/ELT pipelines in AWS using services such as S3, Lambda, Step Functions, EMR, Glue, ECS/EKS, and IAM best practices. Implement and operate NiFi flows for high-throughput, low-latency data ingestion and transformation.
  • Develop, orchestrate, and schedule workflows using Prefect, ensuring reliability, observability, and proper error handling. Implement indexing, search, and retrieval patterns using ElasticSearch, including schema design, cluster management, and query optimization.
  • Collaborate closely with architecture, ML, and application teams to support scalable data solutions.
  • Ensure data quality, lineage, governance, and security across all pipelines.
  • Monitor system performance and troubleshoot issues across distributed data systems.

REQUIRED QUALIFICATIONS

  • Solid understanding of ETL/ELT processes and data modeling best practices.
  • Hands-on experience implementing workflows in Prefect (Prefect 2.0 preferred). In-depth knowledge of ElasticSearch indexing, cluster management, and search optimization.
  • Proficiency in Python and familiarity with common data libraries (Pandas, PySpark, requests, etc.).

DESIRED QUALIFICATIONS

  • Strong experience with Apache NiFi for data flow management and real-time ingestion.
  • Experience building or maintaining RAG pipelines (e.g., vector databases, embeddings, document chunking strategies, retrieval optimization).
  • Proven ability to manage structured and unstructured data pipelines at scale.
  • Experience with AWS cloud services for data engineering.
  • Strong version control and CI/CD experience.
  • Experience with vector databases (OpenSearch, Pinecone, Weaviate, etc.)
  • Familiarity with containerized workflows (Docker, Kubernetes)
  • Experience supporting LLM or generative AI production systems
  • Background in distributed systems, streaming platforms, or data mesh architectures CLEARANCE:
  • Active TS/SCI clearance minimum required

Similar jobs