Quick Overview
Job Description
Need consultant local to Plano, TX to attend Onsite interview.
Job Title: Core Machine Learning Engineer (Python and SRE Focus)
Location: Plano, TX - Onsite
Duration: Long Term
Job Summary:
Experienced Machine Learning Engineer with a strong Site Reliability Engineering (SRE) mindset to join our team. The candidate will have hands-on experience maintaining applications on both Windows and Linux environments, managing on-premises servers, and working with Kubernetes clusters. This role requires solid Python programming skills, a good understanding of machine learning concepts, and practical knowledge of ML model deployment, monitoring, and debugging.
Key Responsibilities:
- Maintain and support machine learning applications running on Windows and Linux servers in on-premises environments.
- Manage and troubleshoot Kubernetes clusters hosting ML workloads.
- Collaborate with data scientists and engineers to deploy machine learning models reliably and efficiently.
- Implement and maintain monitoring and alerting solutions using DataDog to ensure system health and performance.
- Debug and resolve issues in production environments using Python and monitoring tools.
- Automate operational tasks to improve system reliability and scalability.
- Ensure best practices in security, performance, and availability for ML applications.
- Document system architecture, deployment processes, and troubleshooting guides.
Required Qualifications:
- Proven experience working with Windows and Linux operating systems in production environments.
- Hands-on experience managing on-premises servers and Kubernetes clusters and Docker containers
- Strong proficiency in Python programming.
- Solid understanding of machine learning concepts and workflows.
- Experience with machine learning model deployment and lifecycle management.
- Familiarity with monitoring and debugging tools, e.g. DataDog.
- Ability to troubleshoot complex issues in distributed systems.
- Experience with CI/CD pipelines for ML applications.
- Familiarity with AWS cloud platforms
- Background in Site Reliability Engineering or DevOps practices.
- Strong problem-solving skills and attention to detail.
- Excellent communication and collaboration skills.
- We need an engineer who is also familiar with model development
Similar jobs
- PP
Senior Snowflake Platform Engineer
NewPraxis Precision Medicines, Inc.
United States - Remote🇺🇸Remote4 hours agoSQLAWSSnowflake+2Technology - PP
Senior Data Platform Engineer, Commercial
NewPraxis Precision Medicines, Inc.
United States - Remote🇺🇸Remote4 hours agoSQLETLSnowflake+5Technology - OC
Lead DevOps Engineer
NewOctus
Remote - US🇺🇸Remote6 hours agoDockerMicroservicesAWS+5Technology - ON
Site Reliability Engineering Lead
NewOneapp
United States (Remote)🇺🇸Remote3 hours agoProcurementPythonTechnology - MY
Site Reliability Engineer
NewMyFitnessPal
Remote - US🇺🇸$120k - $165k/yrRemote5 hours agoAWSPCI DSSSOC 2+8Technology - KR
AI Platform Engineer, Enablement and Governance Operations
NewKraken.com
United States🇺🇸Remote4 hours agoOAuthSSOCompliance+7Technology