Why This Role Stands Out
This role offers a significant opportunity to deepen your expertise in complex cloud systems and web applications, driving critical incident resolution and preventative solutions for a reputable company. You'll thrive here if you possess strong analytical skills, enjoy problem-solving, and are eager to collaborate with diverse technical teams to enhance application reliability. Apply now to make a substantial impact in a challenging and rewarding environment.
Quick Overview
Job Description
Your Impact /
Key Responsibilities
Tier 4 Escalation & Advanced Troubleshooting: Serve as the highest escalation point for critical and complex web application incidents, performing in-depth analysis, root cause identification, and implementing effective resolutions to minimize downtime and prevent recurrence.
Web Application Subject Matter Expertise: Develop and maintain comprehensive expertise in the web application's architecture, components, dependencies, and operational behavior within the deployment environment.
Problem Management & Prevention: Lead efforts in post-incident reviews, identifying underlying issues, and driving the implementation of permanent solutions and preventative measures to enhance application reliability and stability.
Collaboration & Knowledge Transfer: Work closely with Operations, On-site DevOps, and Development teams to facilitate efficient incident response, share expert knowledge, and contribute to the continuous improvement of operational processes and application design.
Deployment Environment Support: Troubleshoot and resolve issues related to the web application's deployment within its environment, encompassing infrastructure, networking, and related services.
Kubernetes Troubleshooting: Utilize existing Kubernetes knowledge or rapidly acquire proficiency to diagnose and resolve application issues within containerized environments, understand deployment strategies, and assist with related infrastructure challenges.
Documentation & Best Practices: Create and maintain detailed troubleshooting guides, runbooks, knowledge base articles, and technical documentation to empower other support tiers and foster a culture of shared knowledge.
Performance Monitoring & Optimization: Proactively monitor application performance, identify bottlenecks, and recommend architectural or configuration changes to optimize efficiency, scalability, and resilience.
Continuous Improvement: Champion initiatives to enhance the application's operational excellence, including automation, improved monitoring, and streamlined deployment practices. Minimum Qualifications
Proven experience as a Systems Engineer, Application Engineer, or similar role with a strong focus on supporting and troubleshooting complex web applications. Demonstrated Subject Matter Expertise in a critical web application, including its full stack (front-end, back-end, database, APIs). Extensive experience in incident management and problem resolution at a Senior escalation level.
Strong understanding of web application architecture, performance tuning, and security best practices. Experience working with and troubleshooting applications in modern deployment environments. Familiarity with containerization technologies, specifically Kubernetes, or a strong aptitude and willingness to learn and apply it for troubleshooting and operational tasks. Proficiency with monitoring, logging, and alerting tools (e.g., Prometheus, Grafana, ELK Stack, Splunk). Exceptional analytical, problem-solving, and critical thinking skills with a methodical approach to complex issues.
Excellent verbal and written communication skills, with the ability to articulate complex technical concepts to diverse audiences. Ability to work effectively under pressure, prioritize tasks, and manage multiple concurrent issues. Nice to Have Qualifications
Experience with specific web application frameworks or technologies relevant to our stack. Certifications in cloud platforms (e.g., AWS, Azure, GCP) or Kubernetes (e.g., CKA, CKAD). Experience with scripting and automation (e.g., Python, Bash, Ansible). Prior experience in a government or highly regulated environment.
Similar jobs
- TE
Senior Network Engineer - Application Delivery & Security
Auto ApplyTeraswitch
Pittsburgh🇺🇸Hybrid1 month agoLoad BalancingNginxTCP/IP+11Technology - OR
Senior Core Infrastructure Engineer
NewOracle Corporation
United States🇺🇸$79.2k - $209.5k/yrHybrid2 days agoOracleRustEdge Computing+12Technology - OR
Senior Engineer, Core Infrastructure
Oracle Corporation
Seattle, WA🇺🇸$79.2k - $209.5k/yrHybrid4 days agoOracleRustEncryption+6 - SA
IT Systems Engineer
NewSAIC
Santa Maria, CA🇺🇸$140.0k - $180k/yrRemote17 hours agoBashPerlPowerShell+1Technology - SC
Senior Storage Engineer with Security Clearance
NewShield Consulting Solutions
Laurel, MD🇺🇸$190k - $200k/yrRemoteYesterdayFiberOpenStackTCP/IP+3Technology - TS
Senior AWS Cloud Engineer with Security Clearance
NewThe Swift Group
Herndon, VA🇺🇸$5k/moOn-siteYesterdaySwiftAPI GatewayAWS+7Technology - SY
Mid-Level Cloud Software Engineer [$320k/yr+] TS/SCI-FS Poly with Security Clearance
NewSYSTOLIC
Annapolis Junction, MD🇺🇸$320k/yrHybrid17 hours agoDockerAWSGrafana+5Technology - VT
Cloud Software Engineer lv 4 (BP) with Security Clearance
NewVisionary Technologies Inc
Annapolis Junction, MD🇺🇸HybridYesterdayDjangoDockerAWS+6Technology - AK
Network Engineer SME with Security Clearance
NewAkima
Alexandria, VA🇺🇸$156k - $166k/yrHybridYesterdayTCP/IPDNSTechnology - MI
Cloud Engineering Consultant - CTJ - TS/SCI with Security Clearance
NewMicrosoft Corporation
Springfield, VA🇺🇸$101.8k - $193.8k/yrHybridYesterdayScrumAgileAdministrative - CH
Senior Network Engineer with Security Clearance
NewChenega Corporation
Arlington, VA🇺🇸$144.8k/yrHybridYesterdaySplunkTCP/IPAnsible+3Technology - SY
Entry-Level Cloud Software Engineer [$186k/yr+] TS/SCI-FS Poly with Security Clearance
NewSYSTOLIC
Annapolis Junction, MD🇺🇸$186k/yrHybrid17 hours agoC++JavaTechnology