Why This Role Stands Out
This hybrid DevOps Engineer role at LuminaryTech offers a dynamic environment to architect and manage cutting-edge AWS and Kubernetes infrastructure, including specialized AI hardware, fostering significant career growth. You'll thrive here if you're passionate about building robust platforms, automating complex systems with Python or Go, and ensuring high availability for critical AI services, all within a collaborative and forward-thinking team. Apply now to leverage your expertise and make a tangible impact in the rapidly evolving tech landscape!
Quick Overview
Job Description
Core Responsibilities
Design, build, and maintain AWS cloud infrastructure and private data center environments using Infrastructure as Code (IaC).
Manage large-scale Kubernetes (EKS) clusters, including cluster deployment, upgrades, scaling, networking (CNI), storage management, and operational maintenance.
Develop internal DevOps platforms, automation tools, and command-line utilities using Python or Go to improve engineering productivity and operational efficiency.
Build and maintain end-to-end monitoring and observability platforms based on Prometheus, Grafana, and ELK Stack to ensure system reliability and rapid troubleshooting.
Manage AI infrastructure, including GPU servers (NVIDIA A100, T4, etc.), CUDA environments, driver versions, and GPU resource allocation.
Deploy, maintain, and optimize containerized AI applications such as ComfyUI and other Generative AI services for high concurrency and production environments.
Requirements
Education
Bachelor''s degree or above in Computer Science, Information Technology, or related disciplines.
Experience
Minimum 3 years of experience in DevOps, Site Reliability Engineering (SRE), or Infrastructure Engineering.
Languages
Fluent in Mandarin (spoken and written).
Technical Skills
Cloud & Containers
Strong experience with AWS services, including EC2, EKS, S3, VPC, IAM, and related cloud infrastructure.
Deep understanding of Kubernetes architecture, scheduling, networking, storage, and container orchestration.
Programming
Strong programming skills in Python or Go.
Experience developing backend services, automation tools, or internal DevOps platforms.
System Administration
Strong knowledge of Linux operating systems.
Familiarity with TCP/IP, HTTP, DNS, Shell scripting, and system troubleshooting.
CI/CD
Hands-on experience with Jenkins, GitLab CI/CD, GitHub Actions, or similar continuous integration and deployment platforms.
Preferred Qualifications
GPU Infrastructure
Experience managing large-scale GPU clusters.
Knowledge of GPU monitoring, resource scheduling, memory optimization, and Spot Instance cost optimization.
Generative AI Infrastructure
Hands-on experience deploying and maintaining ComfyUI, Stable Diffusion WebUI, or similar AI inference platforms.
Experience with dependency management, multi-user concurrency optimization, and AI model loading acceleration.
MLOps
Familiarity with Kubeflow, MLflow, Triton Inference Server, or similar MLOps platforms.
High Performance Computing
Experience with RDMA networking, distributed computing, and large-scale parallel processing environments.
Similar jobs
- SI
Sr Site Reliability Engineer
NewSpar Information Systems
Frisco, TX🇺🇸Hybrid10 hours agoDockerAzureBash+8Technology - MS
Sr. DevSecOps Engineer I
NewM9 Solutions
Reston, VA🇺🇸$150k - $170k/yrOn-site10 hours agoEngineering - SA
DevSecOps Engineer
SAIC
Panama City, FL🇺🇸Remote7 weeks agoEngineering - BA
Operations Domain Systems Engineer with Security Clearance
NewBarbaricum
Fort Benning, GA🇺🇸Hybrid2 days agoTechnology - LE
DevOps Engineer
Leidos
Kissimmee, FL🇺🇸$87.1k - $157.4k/yrHybrid1 week agoSAFeDockerMicroservices+12Technology - LE
DevOps Engineer
Leidos
Winter Garden, FL🇺🇸$87.1k - $157.4k/yrHybrid1 week agoSAFeDockerMicroservices+12Technology