Site Reliability / Environment Support Lead
Quick Overview
Job Description
Role: Site Reliability / Environment Support Lead
Location: Remote
Duration: 12 months
Note: candidates must meet the residency requirements (Last 5 years in the US with no more than 6 months out of the country).
JOB SUMMARY:
REQUIRED SKILLS:
• Strong SQL skills (complex queries/joins; performance awareness) across Oracle/Cloudera/DataStax or similar platforms.
• Hands-on Linux/UNIX operations and shell scripting for troubleshooting and maintenance/automation.
• Experience with Kafka/Confluent and common integration/data movement patterns (topics, consumers, validation).
• Experience supporting multi-environment application operations (DEV/SIT/CAT/PROD) across on-prem and cloud.
• Proficiency with ServiceNow (incidents/requests) and ability to manage vendor support escalations effectively.
• Monitoring/log analysis experience with Splunk Enterprise; familiarity with AppDynamics/Neustar and cloud monitoring a plus.
• Working knowledge of Agile delivery and tools such as VersionOne and ALM platforms.
• Ability to produce/maintain technical documentation (architecture diagrams, network specs, NCRB artifacts).
• Strong communication/facilitation skills; able to lead incident bridges and coordinate stakeholders under pressure.
RESTRICTIONS:
• Cannot be a member of the USPS eAccess Corporate Developer Registration (CDR) to develop new code due to segregation of duties requirements.
Environment Management Lead partners with business and technical teams to ensure DEV, SIT, and CAT environments are available, secure, and ready for stakeholder use. The role coordinates environment maintenance, including patching, upgrades, and planned outages—while maintaining key technical artifacts (architecture diagrams, NCRB documentation, and system/network specifications). The lead uses GitHub/Copilot and DevSecOps practices to drive automation and AI-assisted analysis, improve reliability, and proactively identify operational efficiencies and cost-reduction opportunities across cloud usage and supplier services.
KEY RESPONSIBILITIES:
• Lead cross-team coordination to keep DEV/SIT/CAT environments available, stable, and release/test ready.
• Manage environment change activities (configuration, refreshes, access, deployments) following governance/compliance processes.
• Create and manage ServiceNow requests/incidents through closure; open/track vendor support cases (e.g., Google, Oracle) and drive timely resolution.
• Coordinate patching, vulnerability remediation, and system upgrades with security/infrastructure teams; track schedules, testing, and risks/escalations.
• Own critical incident management for high-severity outages (triage, communications, restoration, post-incident actions).
• Maintain key technical artifacts: architecture diagrams, system/network specs, and NCRB documentation; ensure accuracy and currency.
• Ensure effective monitoring/alerting (ESM coordination; Splunk dashboards/alerts), analyze trends, and improve reliability.
• Identify and implement automation/AI-assisted improvements, and drive cloud/supplier cost optimization recommendations.
Skills
Similar jobs
DevOps Engineer - Linux (Top Secret with agreement to obtain CI Poly)
North Point Technology · Herndon, United States
16 minutes agoDevOps Engineer (Top Secret with agreement to obtain CI Poly)
North Point Technology · Herndon, United States
16 minutes agoSalesforce Platform Engineer - Hybrid
Genesis10 · Milwaukee, United States
16 minutes ago$70 - $75/hrCloud Devops Engineer
SDH Systems · Dallas, United States
23 minutes agoSenior SRE / DevOps Engineer Python, Jenkins, Linux, F5 & HAProxy -Hybrid NJ
StoneGate-Technologies LLC · Parsippany-Troy Hills, United States
24 minutes agoSite Reliability Engineer (SRE) Cloud Migration & Operational Excellence
Infinite Computer Solutions (ICS) · United States
24 minutes ago