Haystack
← Back to Jobs
Technology
TS

Senior Cloud & DevOps Engineer (ML Infrastructure)

Tanisha Systems, Inc.East Brunswick, NJ🇺🇸United StatesPosted 6 Sept 2026

Why This Role Stands Out

This role offers a fantastic opportunity to build and manage cutting-edge ML infrastructure, driving significant impact within a reputable tech company. You'll thrive here if you have a strong development background and expertise in Kubernetes and cloud platforms, eager to contribute to a dynamic team. Apply now to advance your career in a challenging and rewarding environment.

Quick Overview

Seniority
Mid Senior
Work mode
On Site
Location
East Brunswick, NJ, United States
Posted
10 hours ago
DockerAWSELKFlinkMLOpsMachine LearningApacheApache SparkArgoCDAzureGitHub ActionsGoogle CloudGrafanaHelmJavaJenkinsKafkaKubernetesPrometheusTerraform

Job Description

Senior Cloud & DevOps Engineer (ML Infrastructure)
Location: Sunnyvale, CA/Austin, TX (Onsite) – preference is local
FTE/Contract
Salary: Markert - based on exp
Rate: Markert - based on exp


WE NEED SOME FROM STRONG DEVELOPMENT BACKGROUND EXPERIENCE

Job Description
We are seeking a highly skilled Senior Cloud & DevOps Engineer to design, build, and operate scalable cloud-native platforms that support modern data, machine learning, and application workloads. The ideal candidate will have strong expertise in Kubernetes, cloud infrastructure, Linux systems, and distributed data platforms, along with hands-on experience building secure and reliable CI/CD and platform automation solutions.
Key Responsibilities
  • Design, deploy, and manage cloud-native infrastructure on AWS or other public cloud platforms.
  • Build and maintain Kubernetes-based platforms for scalable application and ML workload deployments.
  • Develop infrastructure automation and operational tooling using Go or Java.
  • Administer and troubleshoot Linux systems, networking, and distributed environments.
  • Implement and manage cloud security controls, IAM policies, and platform governance.
  • Support machine learning teams by enabling scalable ML infrastructure and deployment pipelines.
  • Deploy and operate big data and streaming technologies such as Kafka, Spark, and Flink.
  • Build and enhance CI/CD pipelines, Infrastructure-as-Code (IaC), and platform observability solutions.
  • Monitor platform reliability, performance, security, and cost optimization.
  • Collaborate with engineering, data, and ML teams to improve developer productivity and operational excellence.
Required Qualifications
  • Bachelor’s degree in computer science, Engineering, or a related technical field.
  • 5-7 years of experience in Software Engineering, Platform Engineering, SRE, or DevOps roles.
  • 3-4 years of hands-on experience with AWS or other cloud platforms (Azure/Google Cloud Platform).
  • 3-4 years of experience deploying and operating Kubernetes in production environments.
  • Strong programming experience in Go and/or Java.
  • Deep understanding of Linux internals, system administration, troubleshooting, and networking concepts.
  • Experience with cloud IAM, security best practices, secrets management, and access control.
  • Hands-on experience with at least one big data technology such as Apache Spark, Apache Flink, and messaging platforms like Apache Kafka.
  • Experience with containerization technologies such as Docker.
  • Strong understanding of CI/CD pipelines, automation, and Infrastructure as Code.
Preferred Qualifications
  • Exposure to Machine Learning platforms, MLOps, or AI infrastructure.
  • Experience with Terraform, Helm, ArgoCD, GitHub Actions, or Jenkins.
  • Knowledge of monitoring and observability tools such as Prometheus, Grafana, ELK, or OpenTelemetry.
  • Experience supporting large-scale distributed systems in production.
  • Cloud certifications (AWS, Kubernetes, or equivalent) are a plus.
Key Skills
  • Cloud: AWS, Azure, Google Cloud Platform
  • Containers & Orchestration: Kubernetes, Docker, Helm
  • Languages: Go, Java
  • Operating Systems: Linux Administration & Internals
  • Data & Streaming: Kafka, Spark, Flink
  • Security: IAM, Cloud Security, Secrets Management
  • DevOps: CI/CD, IaC, Automation, Monitoring
  • ML Exposure: MLOps, Model Deployment, ML Infrastructure

Similar jobs