Devops Enigneer/Devops运维
Quick Overview
Job Description
Core Responsibilities
Design, build, and maintain AWS cloud infrastructure and private data center environments using Infrastructure as Code (IaC).
Manage large-scale Kubernetes (EKS) clusters, including cluster deployment, upgrades, scaling, networking (CNI), storage management, and operational maintenance.
Develop internal DevOps platforms, automation tools, and command-line utilities using Python or Go to improve engineering productivity and operational efficiency.
Build and maintain end-to-end monitoring and observability platforms based on Prometheus, Grafana, and ELK Stack to ensure system reliability and rapid troubleshooting.
Manage AI infrastructure, including GPU servers (NVIDIA A100, T4, etc.), CUDA environments, driver versions, and GPU resource allocation.
Deploy, maintain, and optimize containerized AI applications such as ComfyUI and other Generative AI services for high concurrency and production environments.
Requirements
Education
Bachelor''s degree or above in Computer Science, Information Technology, or related disciplines.
Experience
Minimum 3 years of experience in DevOps, Site Reliability Engineering (SRE), or Infrastructure Engineering.
Languages
Fluent in Mandarin (spoken and written).
Technical Skills
Cloud & Containers
Strong experience with AWS services, including EC2, EKS, S3, VPC, IAM, and related cloud infrastructure.
Deep understanding of Kubernetes architecture, scheduling, networking, storage, and container orchestration.
Programming
Strong programming skills in Python or Go.
Experience developing backend services, automation tools, or internal DevOps platforms.
System Administration
Strong knowledge of Linux operating systems.
Familiarity with TCP/IP, HTTP, DNS, Shell scripting, and system troubleshooting.
CI/CD
Hands-on experience with Jenkins, GitLab CI/CD, GitHub Actions, or similar continuous integration and deployment platforms.
Preferred Qualifications
GPU Infrastructure
Experience managing large-scale GPU clusters.
Knowledge of GPU monitoring, resource scheduling, memory optimization, and Spot Instance cost optimization.
Generative AI Infrastructure
Hands-on experience deploying and maintaining ComfyUI, Stable Diffusion WebUI, or similar AI inference platforms.
Experience with dependency management, multi-user concurrency optimization, and AI model loading acceleration.
MLOps
Familiarity with Kubeflow, MLflow, Triton Inference Server, or similar MLOps platforms.
High Performance Computing
Experience with RDMA networking, distributed computing, and large-scale parallel processing environments.
Skills
Similar jobs
DevOps - W2 Only
ASCII Group LLC · Charlotte, United States
30 minutes ago$65/hrAWS DevOps/Platform Engineer
QUANTUM TECHNOLOGIES LLC · Houston, United States
30 minutes ago$88/hrSenior DevOps Engineer
Raas Infotek LLC · New York, United States
30 minutes agoW2 :: Need Site Reliability Engineer at Pennington, NJ
SVARA SOFTWARE SOLUTIONS INC · Pennington, United States
33 minutes agoLead Platform Engineer/Sr. Endpoint Engineer
Irvine Technology Corporation (ITC) · United States
40 minutes ago€85/hrDevOps & Site Reliability Engineer
DTEL Engineering & Consultants Inc · Deerfield, United States
41 minutes ago