Quick Overview
Job Description
Senior AI/ML Platform Engineer
We're seeking a Senior AI/ML Platform Engineer to build and scale enterprise AI/ML platforms that support production-grade machine learning and generative AI workloads. This is a hands-on engineering role focused on creating secure, reliable, observable, and scalable AI/ML infrastructure across cloud and data platforms.
Rather than developing one-off AI solutions, you'll build the foundational systems, automation, and governance that enable AI/ML teams to deploy and operate models at scale.
Key Responsibilities
- Design, build, and support enterprise AI/ML platforms across Databricks, AWS, MLflow, model registries, model serving, feature stores, and related technologies.
- Develop reusable patterns for model development, deployment, monitoring, security, and production support.
- Implement CI/CD pipelines, infrastructure automation, secrets management, access controls, and deployment frameworks.
- Support batch, streaming, real-time, and API-based model deployment architectures.
- Define engineering standards for experimentation, model promotion, observability, governance, and operational support.
- Partner with data, security, infrastructure, and architecture teams to deliver production-ready AI/ML capabilities.
- Build reference architectures, templates, and platform enablement resources to support enterprise adoption.
Required Qualifications
- Bachelor's degree in Computer Science, Engineering, Information Systems, or related field (or equivalent experience).
- Experience building, operating, or supporting production AI/ML platforms within cloud environments.
- Hands-on experience with at least one AI/ML platform such as Databricks, AWS SageMaker, MLflow, Azure ML, or Vertex AI.
- Strong background with CI/CD, Infrastructure as Code, environment management, secrets management, access controls, and deployment automation.
- Experience supporting model development and deployment beyond experimentation and notebook-based workflows.
- Solid understanding of cloud-native architecture, APIs, containers, compute, storage, and observability.
- Experience creating reusable engineering frameworks, templates, and platform standards.
- Proven ability to support AI/ML workloads in governed, production environments.
- Strong troubleshooting skills across platform, deployment, performance, and integration challenges.
Preferred Qualifications
- Deep Databricks experience including Unity Catalog, MLflow, Model Serving, Jobs/Workflows, Clusters, Permissions, and Cost Optimization.
- AWS expertise including IAM, S3, Lambda, ECS/EKS, API Gateway, SageMaker, Bedrock, Networking, and Security.
- Experience in highly regulated or mission-critical environments such as energy, industrial, healthcare, finance, or manufacturing.
- Background in platform cost management and workload optimization.
- Experience creating enablement materials for engineers and data science teams.
Ideal Background
Candidates with experience in MLOps, ModelOps, AI Platform Engineering, Machine Learning Infrastructure Engineering, or AI/ML Enablement Platforms will be particularly successful in this role.
Similar jobs
- NG
Embedded DevOps Engineer Level 4 (AHT)
NewNorthrop Grumman
Los Angeles, California🇺🇸$142.2k - $213.4k/yrOn-site8 minutes agoAWSSonarQubeAgile+8Technology - BA
Operations Domain Systems Engineer with Security Clearance
Barbaricum
Fort Benning, GA🇺🇸Hybrid5 days agoTechnology - SU
DevOps Software Engineer - TS/SCI with Security Clearance
Sunayu, LLC
Bethesda, MD🇺🇸Remote1 week agoMicroservicesLogstashAgile+12Technology - SO
Software Engineer (DevOps Focused) with Security Clearance
NewSet of X
Ft Meade, MD🇺🇸Hybrid2 days agoDjangoSQLAWS+7Technology - KT
Site Reliability Engineer - Data & Caching Systems
NewKforce Technology Staffing
Foster City, CA🇺🇸Hybrid7 hours agoRustAnsibleBamboo+8Technology - CI
DevOps Engineer
NewCelestica International LP
United States🇺🇸Remote7 hours agoAzureConfluenceGit+3Technology