Staff Platform Engineer (AI/ML Infrastructure)
Quick Overview
Job Description
Job Title: Staff Platform Engineer – AI/ML Infrastructure
Location: USA (Remote)
Job Type: Full-Time
Job Overview:
We are seeking a Staff Platform Engineer, AI/ML Infrastructure to provide technical leadership for the cloud platforms, deployment systems, and operational foundations that power enterprise-scale generative AI applications. This role will define and evolve the infrastructure architecture for AI/ML platforms running across AWS, Kubernetes, serverless, and containerized environments. The engineer will lead platform standards for reliability, scalability, observability, CI/CD, security, and developer enablement, while partnering closely with software engineering, AI engineering, security, and operations teams.
Cloud engineering experience with staff-level technical influence. They are comfortable designing infrastructure patterns, writing infrastructure-as-code, improving delivery pipelines, mentoring engineers, and making architectural decisions that raise the operational maturity of AI platforms across multiple teams
Key Responsibilities
- Define and drive the technical strategy for AI/ML platform infrastructure supporting generative AI applications, LLM integrations, model routing, and enterprise AI services.
- Architect, build, and operate scalable cloud platforms using AWS services such as EKS, ECS Fargate, Lambda, DynamoDB, S3, OpenSearch, Secrets Manager, CloudWatch, ALB, and MWAA.
- Establish reusable infrastructure patterns using CloudFormation, Helm, and Terraform to support reliable multi-environment and multi-region deployments.
- Lead CI/CD architecture using GitHub Actions, reusable workflows, OIDC-based AWS authentication, automated quality gates, deployment promotion, and environment approvals.
- Design and improve observability across AI platforms, including CloudWatch dashboards, logs, alarms, PrometheGrafana, OpenSearch, Langfuse, and LLM-specific operational metrics.
- Build platform capabilities for GenAI workloads, including model availability monitoring.
- Partner with software engineering teams to improve deployment reliability, rollback strategies, health checks, autoscaling, load testing, and runtime performance.
- Define and enforce security and compliance practices for infrastructure, including IAM permission boundaries, Secrets Manager usage, secret scanning, audit logging, tagging standards, and change-management controls.
- Provide technical leadership for cost optimization, capacity planning, environment standardization, and operational resilience across development, test, production, and sandbox environments.
- Mentor engineers, review architecture and infrastructure designs, and influence platform engineering practices across teams.
- Troubleshoot complex production issues across cloud infrastructure, networking, containers, serverless workloads, CI/CD systems, and observability platforms.
- Translate enterprise requirements for security, compliance, reliability, and governance into pragmatic engineering standards and automation
Basic Qualifications
- 7+ years of experience in DevOps, platform engineering, cloud infrastructure, site reliability engineering, or software engineering roles.
- Strong hands-on experience with AWS/Azure/Google Cloud Platform infrastructure and services, including container, serverless, networking, storage, observability, and security services.
- Experience designing and operating production systems on Kubernetes, ECS/Fargate, or comparable container orchestration platforms.
- Proficiency with infrastructure-as-code, especially CloudFormation, Terraform, Helm, or similar tooling.
- Strong CI/CD experience with GitHub Actions or similar platforms, including reusable workflows, automated testing, deployment gates, and cloud authentication.
- Experience building and operating observability solutions using CloudWatch, PrometheGrafana, OpenSearch, or similar tools.
- Strong understanding of cloud security practices, IAM, secrets management, least-privilege access, audit logging, and compliance requirements.
- Experience supporting distributed systems, microservices, APIs, asynchronous workloads, and multi-environment deployments.
- Demonstrated ability to lead technical design, mentor engineers, and influence engineering practices across teams
Preferred Qualifications
- Experience supporting AI/ML or generative AI platforms, including LLM gateways, model routing, prompt observability, token metering, or model failover.
- Experience operating platforms in regulated enterprise environments, ideally healthcare, pharmaceutical, finance, or life sciences.
- Experience with multi-account, multi-region AWS architectures and enterprise governance patterns.
- Experience with cost optimization, autoscaling strategies, capacity planning, and cloud budget monitoring.
- Experience with load testing and performance validation using tools such as Locust or comparable frameworks.
- Strong Python or scripting skills for platform automation, operational tooling, and CI/CD extensions.
- Ability to communicate complex technical decisions clearly to engineering, security, operations, and leadership audiences.
Technical Environment
This role works across a modern AI platform ecosystem including:
- Cloud: AWS EKS, ECS Fargate, Lambda, DynamoDB, S3, OpenSearch, CloudWatch, Secrets Manager, ALB, VPC, IAM
- Infrastructure-as-Code: CloudFormation, Helm, Terraform
- CI/CD: GitHub Actions, reusable workflows, OIDC federation, environment approvals, automated release promotion
- AI/ML Platform: AWS Bedrock, Azure OpenAI, LiteLLM, Langfuse
- Observability: CloudWatch dashboards and alarms, Prometheus, Grafana, OpenSearch, Langfuse, custom metrics
- Security & Governance: IAM permission boundaries, secret scanning, audit logging, tagging compliance, change-management automation
- Engineering Practices: Docker, Python, pre-commit, automated testing, load testing, code quality gates, monorepo service standards
Leadership Expectations
As a J090 Staff-level engineer, this role is expected to operate beyond individual delivery. The engineer will identify systemic platform gaps, define technical direction, create reusable standards, and raise engineering maturity across multiple teams. Success in this role requires strong judgment, ownership, and communication. The engineer should be able to balance hands-on implementation with architectural leadership, guide teams through ambiguous technical decisions, and build platform capabilities that make AI product teams faster, safer, and more reliable.
Skills
Similar jobs
Platform Engineer III
Apex Systems · Cincinnati, United States
8 minutes ago$60 - $68/hrDevOps Engineer - Madison, WI, Columbus, OH, Chicago, IL, Minneapolis, MN, Detroit, MI.
TechniPros, LLC · Madison, United States
10 minutes agoDevOps Engineer - Dallas, TX, Austin, TX, Houston, TX, San Antonio, TX.
TechniPros, LLC · Dallas, United States
12 minutes agoDevOps Engineer - Washington, DC, Silver Spring, MD, Philadelphia, PA, Wilmington, DE, Fairfax, VA.
TechniPros, LLC · Washington, United States
30 minutes agoSalesforce DevOps Engineer
Tixy Services LLC · Dallas, United States
30 minutes agoDevOps Engineer (Java Environment)
Wise Skulls Corp. · Austin, United States
30 minutes ago