Haystack
← Back to Jobs
Technology
MC

AWS Data Engineer

Maven CompaniesUnited States🇺🇸United StatesPosted 9 Sept 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
Yesterday
DynamoDBAPI GatewayAWSKafkaPythonRedshiftTerraform

Job Description

Role Summary

We are looking for an AWS-focused software and data engineer to build student intelligence, predictive risk, personalization, knowledge retrieval, and AI context services. The work centers on creating a reliable, longitudinal understanding of a user and making that context seamlessly available to downstream applications and AI experiences.

The ideal candidate builds reusable, highly available production capabilities. This role is not intended for a pure pipeline developer, a notebook-focused data scientist, or an engineer centered only on experimentation.

 

What You'll Do & Technical Expectations

Your responsibilities and impact are aligned directly with our core engineering priorities:

  • Python Engineering (Must-Have): Act as a core developer using Python to build robust APIs, data processing pipelines, automation, testing frameworks, ML/AI integrations, and production-grade services.
  • Infrastructure as Code (Must-Have): Maintain full hands-on ownership of our AWS environments using Terraform, IAM, networking, and CI/CD pipelines to create repeatable, governed, and version-controlled infrastructure.
  • Data Modeling & Context Engineering (Priority 1): Leverage DynamoDB, Aurora, Redshift, Neptune, and S3 to build longitudinal user profiles, reconcile current vs. historical states, and support behavioral profiles and memory.
  • Applied ML & Predictive Systems (Priority 2): Productionize predictive risk models, reusable features, scoring, explainability, and model monitoring utilizing SageMaker, Bedrock, Feature Store, and Model Monitor.
  • Retrieval / RAG / Knowledge Systems (Priority 3): Support grounded policy and knowledge retrieval with traceable source material for AI experiences using Bedrock Knowledge Bases, OpenSearch, S3, and Lambda.
  • Real-Time Data & Integration (Priority 4): Transform behavioral and institutional events into updated profile and context attributes in near real-time using MSK/Kafka, Kinesis, EventBridge, Lambda, API Gateway, and Glue.
  • Production AI & Agent Engineering (Priority 5): Deliver governed context to AI applications and autonomous agents with observability, security, tool calling, and controlled access through Bedrock Agents/Runtime, Lambda, ECS/EKS, CloudWatch, IAM, and Secrets Manager.

Similar jobs