Quick Overview
Job Description
About the job
The world’s most critical--and at-risk--business applications have been neglected for far too long. Onapsis eliminates this blind spot by providing cybersecurity solutions dedicated to business-critical applications. Onapsis helps nearly 30% of the Forbes Global 100 understand the threats and risks across their SAP and Oracle landscapes, whether running on-premises, in the cloud, or in a hybrid environment.
We are looking for a Site Reliability Engineer II to join our global engineering team. In this role, you will apply software engineering principles to operations, ensuring our cloud platform and distributed security products remain highly available, scalable, and resilient. You will be directly accountable for monitoring, inspecting, troubleshooting, and resolving service and product issues while continuously working with engineering partners to improve telemetry and related operation automations.
Rather than performing manual operational maintenance, you will write code, build automation, and design observability frameworks that eliminate toil and prevent system failures. You will work side-by-side with product development teams to embed reliability into the software lifecycle from day one.
What you will be doing, your legacy:
As a member of our global SRE team, you will take shared full-stack ownership of Onapsis production environments, balancing active operational response with modern software engineering principles. You will develop a deep, end-to-end understanding of our system architecture, technical dependencies, and service behaviors to maximize the performance, scalability, and resilience of our enterprise security platforms.
Your time will be split between managing live production environments and driving engineering initiatives that ensure long-term system stability. When troubleshooting live incidents, you will diagnose complex issues across distributed cloud services and stateful infrastructure. Between operational cycles, you will shift into a software engineering mindset—designing, developing, and maintaining custom automation tooling to eliminate repetitive toil, optimize monitoring telemetry, and increase operational efficiency.
The ideal Site Reliability Engineer II brings a strong foundational toolkit spanning the software development lifecycle (SDLC), Linux systems administration, core networking protocols, and cloud computing (AWS preferred). You excel at viewing operational friction through a software lens, leveraging coding and automation to transform system vulnerabilities and outages into permanently solved engineering problems.
Requirements:
- Bachelor's or Master’s degree in Computer Science or related fields or equivalent experience.
- 2-4 years experience in SRE, DevOps, or Cloud Engineering role supporting production distributed systems in cloud environments
- Beginner knowledge of programming skills such as Java and Python
- Intermediate Knowledge of the SRE Observability practices such as Application Performance Monitoring, Error Budget definition and tracking, and ensuring monitoring, logging and tracing are connected.
- Understanding of SLI, SLO, SLA methodologies
- Beginner knowledge of Infrastructure as Code and Configuration Management using Terraform
- Beginner knowledge of Containers and related orchestration platforms such as Kubernetes
- Intermediate knowledge of orchestrating technical operations through scripting and automation using Bash or Python.
- Experience using AI coding assistants for daily tasks, debugging, and rapid prototyping
- Experience managing Linux OS (Debian/Ubuntu/OpenSuse), Queue servers (RabbitMQ), Database instances (PostgreSQL) and AWS cloud computing infrastructure such as EC2 instances or EKS
- Intermediate knowledge in troubleshooting complex software and networking issues
- Proven ability to quickly learn new technical domains and then train others
- Great verbal and written communication skills
Desired skills or interests in:
- Ability to own end-to-end features, service integrations, database schema design, and operational telemetry.
- Secure software development best practices knowledge
- Knowledge of test-driven development (TDD), CI / CD tooling, and Agile methodologies
- Knowledge of professional software engineering practices & best practices for the full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations
- Experience with distributed computing and enterprise-wide systems
- Experience in communicating with users, other technical teams, and senior management to collect requirements, describe software product features, technical designs, and product strategy
What we offer:
- A role in shaping the future of protecting the most critical applications that run the world's business and a career that grows as the company grows.
- A unique culture of high achievement and teamwork.
- Supportive and humble colleagues are some of the space's top problem solvers and innovators.
- Financial security through competitive compensation and incentives.
- A comprehensive benefits plan, including medical, dental, vision, disability, life insurance, and a 401K.
- Flexible work options: OnaFlex PTO
Location: Dallas Mid-Cities - Hybrid - 2 days per week in office, so candidates must be commutable to Mid-Cities / Dallas, Texas - This is not a remote-only opportunity.
About Onapsis:
Onapsis protects the business applications that run the global economy. The Onapsis Platform delivers vulnerability management, change assurance, and continuous compliance for business applications from leading vendors such as SAP, Oracle, and others. The Onapsis Platform is powered by the Onapsis Research Labs, the team responsible for the discovery and mitigation of more than 1,000 zero-day vulnerabilities in business applications.
Onapsis is headquartered in Boston, MA, with offices in Dallas, TX, Heidelberg, Germany, Bucharest, Romania, and Buenos Aires, Argentina, and proudly serves hundreds of the world’s leading brands, including close to 30% of the Forbes Global 100, six of the top 10 automotive companies, five of the top 10 chemical companies, four of the top 10 technology companies, and three of the top 10 oil and gas companies.
For more information, connect with Onapsis on LinkedIn or visit https://www.onapsis.com.
#LI-AC1
#LI-Hybrid
Similar jobs
- NE
Site Reliability Engineer
NewAuto ApplyNebius
Remote - United States🇺🇸$130k - $180k/yrRemote4 hours agoBashPythonTechnology - EL
Senior DevOps Engineer (NOAA badge required)
NewAuto ApplyElement84
Alexandria HQ (remote)🇺🇸Remote11 hours agoDockerDynamoDBSQL+20Technology - CM
Senior AI Platform Engineer
NewAuto ApplyCode Metal
Boston Hub🇺🇸Hybrid2 hours agoAPI GatewayMLflowSalesforce+7Technology - CL
Platform Engineer (Kubernetes), Mid-Level
NewAuto ApplyClera
San Francisco🇺🇸On-site12 hours agoService MeshArgoCDGrafana+4Technology - WE
Site Reliability Engineer, Cloud Infrastructure
NewAuto ApplyWeave
Weave - Headquarters (Lehi🇺🇸Remote11 hours agoDockerGCPAnsible+11Technology - CO
Senior Site Reliability Engineer
NewAuto ApplyCoalition, Inc.
Any location🇺🇸Hybrid6 hours agoMicroservicesAWSCapacity Planning+7Technology - CI
DevSecOps Engineer
NewAuto ApplyCHAOS Industries
Washington🇺🇸Hybrid7 hours agoDockerAWSOWASP+14Engineering - SA
Senior Infrastructure & Reliability Engineer
NewAuto ApplyStuut Ai
New York City🇺🇸Hybrid8 hours agoSOC 2AuditingERP+2Technology - CA
Associate Site Reliability Engineer/Site Reliability Engineer
NewAuto ApplyC3 AI
Redwood City🇺🇸Hybrid13 hours agoGCPAWSAnsible+8Technology - AI
Contractor: DevOps Engineer
NewAuto ApplyAbacus Insights
United States🇺🇸Hybrid13 hours agoDockerAWSSplunk+10Technology - ED
DevOps Engineer III
NewAuto ApplyEnable Dental
Austin, Texas🇺🇸Remote13 hours agoDockerAWSEncryption+17Technology - TH
DevOps Engineer
NewAuto ApplyTheIncLab
Colorado Springs, Colorado🇺🇸Hybrid7 hours agoDockerShellAWS+22Technology