Why This Role Stands Out
This remote DevOps opportunity offers a highly competitive hourly rate and the chance to leverage your NetApp expertise to drive innovation in storage automation and AI. You'll thrive in this role if you possess strong Infrastructure-as-Code skills and are eager to apply AI/ML to enhance storage operations and predictive management. Apply now to make a significant impact within a forward-thinking technology environment.
Quick Overview
Job Description
Title: Storage Engineer (Resident)
Location: Remote
Duration: 6+ Months Contract
Pay rate : $85/hr on W2
Job Description:
Client is seeking a senior NetApp Resident Engineer to embed with the Storage Systems Group (SSG) and accelerate automation and AI-enabled operations across a large multi-site ONTAP fleet. This role pairs deep NetApp ONTAP expertise with hands-on Infrastructure-as-Code (Terraform/Ansible) skills and practical experience applying AI/ML tooling to storage operations — including anomaly detection, capacity forecasting, and AI-assisted troubleshooting. The resident engineer will work directly with SSG leadership and engineers on active initiatives spanning fleet lifecycle management and ServiceNow ITOM/AIOps integration.
Key Responsibilities:
- Design, build, and maintain Terraform modules and Ansible playbooks/roles for ONTAP provisioning, SVM/volume lifecycle management, SnapMirror/SnapCenter operations, and fleet-wide configuration drift remediation.
- Partner with SSG engineers to extend existing automation (e.g., self-service backup/restore workflows, capacity reclamation scripting) into standardized, version-controlled IaC pipelines.
- Apply AI/ML capabilities — including NetApp BlueXP/AIOps tooling, anomaly detection, and LLM-assisted diagnostics — to reduce time-to-resolution and to support predictive capacity and health management across the fleet.
- Support integration work between ONTAP telemetry, Metabase/Snowflake reporting, and ServiceNow ITOM Event Management webhook pipelines.
- Support NetApp snapshot and clone technology (Snapshot, FlexClone) automation and NFS datastore lifecycle management across the fleet.
- Document runbooks, automation architecture, and operational procedures; provide knowledge transfer and upskilling to SSG engineers on IaC and AI-assisted operations practices.
Required Qualifications:
- 8+ years of experience with NetApp ONTAP administration in enterprise, multi-cluster environments (SVMs, FlexVols/FlexGroups, SnapMirror, SnapCenter, CIFS/NFS).
- 3+ years hands-on experience writing and maintaining Terraform for infrastructure provisioning, including NetApp/ONTAP or adjacent storage/infrastructure providers.
- 3+ years hands-on experience with Ansible for configuration management and operational automation (playbooks, roles, idempotent task design).
- Demonstrated experience applying AI/ML or LLM-based tooling to infrastructure or storage operations (e.g., predictive analytics, anomaly detection, AI-assisted scripting or troubleshooting, AI-output verification practices).
- Strong scripting ability in Python and/or PowerShell for automation glue code, API integration, and reporting.
- Working knowledge of REST API-based automation against ONTAP and adjacent platforms.
- Experience with version control and CI/CD practices (Git-based workflows) for infrastructure code.
- Excellent written communication skills for runbook and architecture documentation.
Preferred Qualifications:
- Experience with NFS-backed datastore performance tuning (NFSv3 vs. NFSv4.1/4.2) across virtualized environments.
- Familiarity with Dell PPDM/Data Domain, StorageGRID, or other backup/DR platforms.
- Experience integrating storage/infrastructure telemetry into ITSM/ITOM platforms (ServiceNow Event Management, PagerDuty).
- NetApp certifications (NCDA, NCIE) and/or HashiCorp Terraform Associate certification.
- Prior experience in a vendor-resident or embedded consulting engagement model.
Engagement Details:
- Contract engagement placed by NetApp to work embedded within Client''s SSG team.
- Remote/hybrid; occasional coordination across U.S. data center sites as fleet work requires.
- Success will be measured by automation coverage delivered (Terraform/Ansible modules in production use), reduction in manual operational toil, and measurable AI-assisted improvements to incident response or capacity planning.
Similar jobs
- SE
Senior Staff Data Platform Engineer - Kafka - Apache Iceberg - Apache Spark
NewServiceNow
San Diego, CALIFORNIA🇺🇸Hybrid16 hours agoJavaPostgreSQLMySQL+6Technology - SE
Senior Staff Software Engineer – SRE & AIOps
NewServiceNow
Santa Clara, CALIFORNIA🇺🇸HybridYesterdayGCPMicroservicesSQL+11Technology - IN
Principal Platform Engineer
NewInterSystems
Boston🇺🇸YesterdayBashHTTPSJava+2Technology - NE
AI Platform Engineer-Anthropic-US West
NewNewRocket
USA - Remote🇺🇸Remote15 hours agoDockerMicroservicesMongoDB+32Technology - NE
AI Platform Engineer-Anthropic-US East
NewNewRocket
USA - Remote🇺🇸Remote15 hours agoDockerMicroservicesMongoDB+32Technology - AN
FedRAMP Site Reliability Engineer (FedSRE) - CloudVision
Arista Networks
Remote🇺🇸Remote9 months agoPythonGoBash+6Technology