Haystack
← Back to Jobs
Technology
IN

Lead Platform Engineer (SRE)

InfoVision, Inc.United States🇺🇸United StatesPosted Sep 22, 2026

Quick Overview

Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
22 hours ago
MicroservicesELKLoad BalancingTCP/IPTDDAgileAnsibleDNSGrafanaHTTPTerraform

Job Description

Technology Engineering - Lead Platform Engineer (SRE) (IC3)
Location: Evansville IN

Job Description:
OneMain Financial is the country's largest lending-exclusive financial company, proudly serving millions of customers with safe, affordable, and transparent installment loans. Our customers turn to us every day - online and at over 1,400 branches in 44 states - to help them take control and improve their financial lives. It's all about doing the right thing - a mission that hasn't changed for more than 100 years.

We are seeking a Lead Platform Engineer to:-       
Able to collaboratively establish availability and performance objectives, measure progress towards those objectives, and implement necessary changes.- Able to interpret changing technical & process needs of product teams and adjust platform and tooling meets those needs.

The Lead Monitoring and Observability Engineer (IC3) serves as a senior technical contributor responsible for ensuring monitoring reliability, telemetry quality, automation maturity, and operability across OMF’s infrastructure and application ecosystem. This role acts as a technical mentor, monitoring and networking SME, and hands-on engineer who guides monitoring-platform evolution, improves service quality, and collaborates with product and engineering teams to deliver scalable, stable, and observable systems.

Key Responsibilities
Monitoring Reliability & Performance
•       Establish platform SLOs, availability goals, latency/error budgets, reliability metrics, and monitoring coverage expectations in partnership with teams.
•       Continuously measure service and platform health and implement changes to improve reliability, performance, alert quality, and operational stability.
•       Define and maintain standards for actionable alerting, dashboards, logs, metrics, traces, and service-health reporting.

Architecture & Technical Leadership
•       Lead observability design efforts for Elastic/ELK, telemetry pipelines, monitoring platforms, dashboards, alerting, distributed tracing, synthetic monitoring, microservices platforms, cloud infrastructure, or network monitoring, depending on assignment.
•       Provide “out-of-the-box” technical solutions that balance velocity, reliability, operational visibility, and cost.
•       Evaluate monitoring and observability tools, patterns, integrations, and emerging requirements.
•       Define reusable observability patterns, reference architectures, onboarding guidance, and validation practices for product and platform teams.

Hands-On Engineering & Automation
•       Perform advanced configuration, IaC development (Terraform/Ansible), monitoring-as-code development, CI/CD pipeline engineering, and cloud platform automation.
•       Build and maintain scalable, resilient monitoring and observability components that support product teams.
•       Implement and maintain monitoring, alerting, data visualization, logging, metrics, tracing, and telemetry collection capabilities as noted in the IC3 role reference.
•       Improve automation for monitoring onboarding, dashboard creation, alert configuration, telemetry collection, and operational workflows.

Operational Excellence
•       Reduce manual operations through automation and self-service observability tooling.
•       Review, optimize, and maintain observability capabilities including logs, metrics, traces, dashboards, alerting, and service-health reporting.
•       Improve alert signal quality, reduce alert noise, and ensure alerts have clear ownership, escalation paths, and actionable runbook guidance.
•       Participate in and lead high-severity incident response for monitoring and observability-owned domains.
•       Support root-cause analysis by correlating telemetry, infrastructure conditions, application behavior, and operational events.

Network Monitoring
•       Define and maintain network-monitoring standards for network availability, reachability, latency, packet loss, interface health, capacity, device health, routing, and dependency-related service impact.
•       Partner with Networking, Cloud, AppDev, Security, and SRE teams to establish monitoring coverage for critical network devices, services, and customer journeys.
•       Build and maintain network-monitoring dashboards, alerting patterns, service-health views, and operational workflows.
•       Establish validation practices for network-device onboarding, telemetry collection, alert quality, dashboard completeness, and operational readiness.
•       Support the migration of network monitoring from OpsRamp to the selected network-monitoring tool, including requirements definition, technical design, migration planning, testing, validation, and operational handoff.
•       Evaluate network-monitoring tools, integrations, automation opportunities, and emerging network observability requirements.
•       Connect network telemetry and alerts to application, infrastructure, and customer-impact signals to improve triage and incident response.

Cross-Team Collaboration
•       Partner with AppDev, Security, Observability, Networking, Cloud, and SRE teams to deliver cohesive monitoring services and shared solutions.
•       Adjust monitoring-platform design, observability standards, and processes based on evolving product team needs.
•       Collaborate with service owners to prioritize observability improvements for critical services.

Mentorship & Influence
•       Mentor IC1–IC2 engineers in engineering best practices, TDD, agile ceremonies, monitoring and observability fundamentals, and operational workflows.
•       Provide code reviews, design feedback, and internal technical enablement.
•       Guide partner teams in adopting agreed observability standards and reusable implementation patterns.
•       Provide technical leadership on monitoring architecture, telemetry quality, alerting practices, and operational readiness.

Qualifications

Required
•       Demonstrated expertise in monitoring and observability engineering, including Elastic/ELK, dashboards, alerting, logging, metrics, tracing, telemetry pipelines, cloud infrastructure, networking, or equivalent.
•       Proven ability to define availability and performance objectives and implement monitoring-driven improvements.
•       Hands-on experience with automation, including Terraform, Ansible, scripts, monitoring-as-code, or cloud-platform automation.
•       Experience designing, implementing, and maintaining monitoring, alerting, dashboards, and service-health reporting.
•       Experience with observability practices across logs, metrics, traces, dashboards, alerting, and incident response.
•       Ability to evaluate alternatives, provide thought leadership, and guide technical direction.
•       Proficiency with TDD, agile development, and iterative delivery.
•       Strong communication and mentoring capability.

Network Monitoring Requirements
•       Experience with enterprise network monitoring, network operations, network observability, or a related infrastructure discipline.
•       Working knowledge of network technologies and protocols, including TCP/IP, DNS, HTTP/S, SNMP, ICMP, routing, switching, firewalls, load balancing, VPN, and WAN connectivity.
•       Experience monitoring network devices, interfaces, availability, latency, packet loss, bandwidth utilization, capacity, and reachability.
•       Ability to design network-focused dashboards, alerts, escalation workflows, and service-health views.
•       Experience collaborating with Networking and incident-response teams during complex service-impacting events.
•       Familiarity with network-monitoring migrations, platform consolidations, tool evaluations, or proof-of-concepts.

Preferred
•       Experience with distributed tracing, observability platforms, monitoring analytics, or OpenTelemetry-aligned telemetry practices.
•       Familiarity with DevOps/SRE practices, including SLOs, runbooks, deployments, incident response, and post-incident improvement.
•       Experience with multi-cloud environments or hybrid architectures.
•       Experience with OpsRamp, SolarWinds, Elastic/ELK, Grafana, ServiceNow, or related monitoring, alerting, and incident-management integrations.
•       Experience with synthetic monitoring, customer-journey monitoring, or service-level reporting.
•       Experience leading a network-monitoring platform migration or enterprise monitoring-tool evaluation

If interested, Please share below details with update resume:

Full Name: 
Phone:     
E-mail:  
Rate:
Location: 
Visa Status: 
Availability:
SSN (Last 4 digit): 
Date of Birth: 
LinkedIn Profile:
Availability for the interview:
Availability for the project:

Similar jobs