Senior AI ( Forward Deployment Engineer ) FDE - With ML Experience/ FDE with hands-on experience building ML solutions for observability, application performance monitoring, capacity planning and infrastructure optimization- Onsite work Denver CO
Quick Overview
Job Description
Senior AI ( Forward Deployment Engineer ) FDE - With ML Experience/ FDE with hands-on experience building ML solutions for observability, application performance monitoring, capacity planning and infrastructure optimization- Onsite work Denver CO
Requisition Name: Senior AI FDE - With ML Experience
Start Date: 9/28/2026
Duration: 14 Weeks
Services Location: CO/Denver
Mandatory qualifying question:
- Do you have strong hands-on experience with Python and data/ML libraries such as Pandas, NumPy, Scikit-learn, XGBoost, TensorFlow, or PyTorch?
- Do you have strong experience in Machine Learning and Predictive Analytics?
Description Of Services:
Senior AI FDE - With ML Experience
Mandatory: Python, Pandas, NumPy, Scikit-learn, Machine Learning, Time-Series Forecasting, Anomaly Detection, Regression, SQL, Observability/Monitoring
We are looking for an experienced Machine Learning Engineer / Data Scientist to build intelligent, data-driven solutions for application performance monitoring, capacity planning, and infrastructure resource optimization.
The ideal candidate will have strong hands-on experience in Python, Machine Learning, Time-Series Forecasting, Anomaly Detection, Observability, and large-scale telemetry data processing. This role will focus on analyzing application/API usage patterns and infrastructure metrics to develop ML models that predict resource requirements and provide intelligent recommendations for CPU, memory, capacity, and scaling.
You will work closely with DevOps, SRE, Cloud, Infrastructure, and Application Engineering teams to develop and operationalize ML-driven capacity optimization solutions.
Required Qualifications & Skills
Strong hands-on experience with Python and data/ML libraries such as Pandas, NumPy, Scikit-learn, XGBoost, TensorFlow, or PyTorch.
Strong experience in Machine Learning and Predictive Analytics.
Hands-on experience with Time-Series Forecasting and analysis of time-dependent data.
Experience with Anomaly Detection, Regression, Classification, and Clustering techniques.
Strong understanding of feature engineering, model evaluation, model optimization, and statistical analysis.
Experience building and deploying production-grade ML models.
Strong experience working with observability, monitoring, logs, metrics, and telemetry data.
Experience with observability/monitoring platforms such as Prometheus, Grafana, Splunk, ELK/Elastic, Datadog, or equivalent technologies.
Experience processing large volumes of logs, metrics, and telemetry data.
Strong SQL skills and experience working with large datasets and/or data warehouses.
Experience designing or working with batch and/or streaming data pipelines.
Experience with MLOps, including model deployment, monitoring, versioning, retraining, and drift detection.
Strong problem-solving and analytical skills with the ability to translate complex data into actionable recommendations.
Preferred Qualifications
Experience with Docker and Kubernetes.
Experience with cloud platforms such as AWS, Azure, or Google Cloud Platform.
Experience with CI/CD and cloud-based ML platforms.
Experience with streaming technologies such as Kafka, Spark Streaming, or equivalent.
Experience with distributed data processing and large-scale analytics.
Experience in application performance management, infrastructure optimization, capacity planning, or FinOps.
Experience building ML-powered dashboards, APIs, or recommendation systems
Deliverables:
Observability & Data Analytics
Analyze historical and real-time application and backend API usage patterns, including request volume, throughput, latency, concurrency, peak traffic, and usage trends.
Analyze infrastructure utilization metrics such as CPU, memory, disk, network, and application performance metrics.
Process and derive insights from large volumes of application, system, API logs, metrics, and telemetry data.
Identify trends, seasonality, usage patterns, anomalies, and peak-load behavior.
Correlate API traffic and application workload with infrastructure consumption to understand resource utilization per transaction/request.
Develop scalable data pipelines for collecting, transforming, aggregating, and processing telemetry from multiple applications.
Machine Learning & Predictive Analytics
Develop ML and statistical models to forecast application traffic, resource utilization, and future capacity requirements.
Build models to determine optimal CPU and memory sizing based on historical and projected workloads.
Develop anomaly detection models to identify unusual resource consumption, traffic patterns, and application behavior.
Apply regression, classification, clustering, and other ML techniques where appropriate.
Perform feature engineering, model selection, validation, optimization, and performance evaluation.
Establish confidence levels and explain the rationale behind ML-driven sizing recommendations.
Capacity Optimization & Recommendations
Identify over-provisioned and under-provisioned applications based on workload and resource utilization patterns.
Develop intelligent recommendations for:
Recommended CPU allocation
Recommended memory allocation
Minimum and maximum capacity
Expected peak resource requirements
Scaling thresholds
Future capacity requirements based on projected traffic growth
Continuously evaluate recommendations against actual production performance and improve the models based on feedback.
Partner with DevOps, SRE, Cloud, and Application teams to validate and implement ML-driven recommendations.
MLOps & Productionization
Deploy and monitor ML models in production environments.
Implement ML lifecycle practices including model versioning, monitoring, retraining, model validation, and drift detection.
Build production-grade ML pipelines and inference solutions.
Develop dashboards, reports, APIs, or other interfaces to expose capacity recommendations and application performance insights.
Ensure data quality, feature consistency, model reliability, and production monitoring.
Similar jobs
- TA
Desktop Support Technician Level II with Security Clearance
NewTyto Athene
Fort Worth, TX🇺🇸$25 - $30/hrOn-site2 days agoLESSTechnology - QS
Information Security Analyst - SME with Security Clearance
NewQuantum Sky Engineering LLC
Camp Lejeune, NC🇺🇸$155k - $165k/yrHybrid2 days agoPenetration TestingTechnology - WY
System Administrator 2/3 with Security Clearance
NewWyetech, LLC
Annapolis Junction, MD🇺🇸$42 - $118/hrOn-site2 days agoAnsibleBashDNS+5Technology - SG
Oracle APEX Developer
NewSkywalk Global
Richmond, VA🇺🇸HybridYesterdayOraclePL/SQLSOAP+5Technology - MC
Sr. Platform Engineer - DevSecOps, AI with Security Clearance
NewMindbank Consulting Group
San Diego, CA🇺🇸$165k - $220k/yrOn-siteYesterdayDockerAWSTCP/IP+11Technology - RD
Help Desk Support
NewRandstad Digital
PA🇺🇸$15 - $17/hrHybridYesterdayTechnology