Haystack
← Back to Jobs
Technology
SE

AI/ML Data Platform Engineer (Senior) (Onsite)

SerigorLinthicum Heights, MD🇺🇸United StatesPosted Sep 22, 2026

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
Linthicum Heights, MD, United States
Posted
Yesterday
AWSMLOpsMLflowMachine LearningScikit-learnApacheApache SparkCDKCloudFormationDatabricksGenerative AIHIPAAKafkaLLMPyTorchPythonTensorFlowTerraform

Job Description

Job Title: AI/ML Data Platform Engineer (Senior) (Onsite)

Location: Linthicum, MD, 21090

Duration: Upto 7 Years

 

Position Description: The Senior AI/ML Data Platform Engineer will be expected to design, build, and operationalize machine learning infrastructure and AI-driven data solutions on AWS. The successful candidate will bridge data engineering and MLOps, ensuring scalable, secure, and compliant AI/ML pipelines that support advanced analytics and decision-making across programs.

The AI/ML Data Platform Engineer creates and/or maintains operating systems, communications software, database packages, compilers, repositories, and utility and assembler programs. This position is responsible for modifying existing software and developing special-purpose software to ensure efficiency and integrity between systems and applications.

 

Key Responsibilities

        Design and implement end-to-end ML pipelines including data ingestion, feature engineering, model training, validation, and deployment on AWS (SageMaker, Glue, Lambda, Step Functions).

        Build and maintain MLOps infrastructure: model registries, CI/CD for ML models, experiment tracking (MLflow or SageMaker Experiments), and automated retraining pipelines.

        Develop scalable feature stores and data preprocessing pipelines using AWS Glue, EMR, or Spark on Databricks.

        Partner with data scientists and analysts to productionize models and translate research prototypes into robust, maintainable systems.

        Ensure ML systems meet government data security, privacy (PII/PHI handling), and compliance requirements (FISMA, NIST).

        Implement data versioning, lineage tracking, and model explainability/audit capabilities.

        Monitor deployed models for drift, performance degradation, and data quality issues; implement alerting and automated remediation.

        Contribute to data platform architecture decisions, including lakehouse design, data cataloging (AWS Glue Data Catalog / Apache Atlas), and access controls.

        Document architectures, pipeline designs, and operational runbooks to government documentation standards.

Education: This position requires a Bachelor’s degree from an accredited college or university with a major in computer science, information systems, engineering, business, or other related scientific or technical discipline. Three (3) years of equivalent experience in a related field may be substituted for the Bachelor’s degree. (Note: A Master’s degree is preferred.)

 

General Experience: The proposed candidate must have twelve (12) years of computer experience in information systems design. Other experience required:

        6+ years of data engineering or software engineering experience, with 2+ years focused on ML platform or MLOps engineering.

        Deep expertise with AWS ML/AI services: SageMaker, Glue, EMR, Lambda, Step Functions, Kinesis, S3.

        Strong Python skills; experience with ML frameworks (scikit-learn, TensorFlow, PyTorch, XGBoost).

        Experience building and maintaining MLOps pipelines and CI/CD for ML workflows.

        Solid understanding of distributed computing (Spark/EMR) and large-scale data processing.

        Familiarity with government data security requirements and experience operating in FedRAMP-authorized AWS environments.

        Experience with infrastructure-as-code tools (Terraform, AWS CDK, CloudFormation).

 

Specialized Experience: The proposed candidate must have at least ten (10) years of experience in IT systems analysis and programming including the following:

        Experience with generative AI / LLM integration (AWS Bedrock, LangChain) in enterprise or government contexts.

        Knowledge of responsible AI, model governance, and bias detection frameworks.

        Experience with real-time/streaming ML pipelines using Kinesis or Kafka.

        Familiarity with data mesh or federated data architectures.

        Background in healthcare, public health, or social services data (HIPAA/42 CFR Part 2 awareness).

 

Preferred Certifications

        AWS Certified Machine Learning — Specialty

        AWS Certified Data Analytics — Specialty

        AWS Certified Solutions Architect — Associate or Professional

        Databricks Certified Associate Developer for Apache Spark

 

Similar jobs