← Back to Jobs
Technology
Site Reliability Engineer (SRE)
Xona Space Systems, IncBurlingame, CA🇺🇸United StatesPosted 14 Aug 2026
Quick Overview
Work Type
Hybrid
Level
Mid Senior
Job Description
Xona is the navigational intelligence company bringing real-time, centimeter-level certainty to any device, anywhere on Earth.
With Pulsar - the world's most advanced PNT satellite infrastructure in Low Earth Orbit - Xona will offer a future-proof, backwards-compatible global positioning system optimized for absolute precision, superior power, and robust protection.
We are seeking a Site Reliability Engineer (SRE) to architect and manage the critical ground infrastructure for our satellite constellation. This role is responsible for the "last mile" of mission success: ensuring that the software controlling our orbital assets is highly available, scalable, and seamlessly integrated with Mission Operations.
You will own the lifecycle of our production environments, from automating deployments via Infrastructure as Code (IaC) to managing the core data systems that track constellation health and user activity.
Required Qualifications
For U.K. Roles: To comply with U.K. regulations, this role requires Baseline Personnel Security Standard (BPSS) checks, and successful candidates must be eligible to obtain UK Security Clearance (SC).
For Canada Roles: Successful candidates must obtain and hold a security clearance at the reliability status level, and pass security assessment for the Canadian Controlled Goods Program (CGP) and ITAR.
We celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or veteran status.
With Pulsar - the world's most advanced PNT satellite infrastructure in Low Earth Orbit - Xona will offer a future-proof, backwards-compatible global positioning system optimized for absolute precision, superior power, and robust protection.
We are seeking a Site Reliability Engineer (SRE) to architect and manage the critical ground infrastructure for our satellite constellation. This role is responsible for the "last mile" of mission success: ensuring that the software controlling our orbital assets is highly available, scalable, and seamlessly integrated with Mission Operations.
You will own the lifecycle of our production environments, from automating deployments via Infrastructure as Code (IaC) to managing the core data systems that track constellation health and user activity.
Required Qualifications
- Infrastructure as Code (IaC): Design and maintain scalable, repeatable cloud infrastructure (AWS) using tools like Terraform or CloudFormation.
- Mission Ops Integration: Build and optimize the interfaces between core data management systems and Mission Operations software, ensuring reliable telemetry and command flows.
- User & Data Management: Architect and maintain high-availability identity providers (IdP) and distributed databases to support global user access and real-time data processing.
- Automated Deployment Pipelines: Create and manage robust CI/CD pipelines to deploy containerized applications into production with a focus on zero-downtime and rollback capabilities.
- Observability & Reliability: Implement comprehensive monitoring, alerting, and logging (e.g., Prometheus, Grafana, ELK) to ensure 99.99% uptime for ground segment services.
- Scalability Engineering: Perform capacity planning and performance tuning to handle the high-throughput data requirements of a growing satellite constellation.
- Cloud Operations: 4+ years of experience managing production-grade environments in AWS, Google Cloud Platform, or Azure.
- Orchestration: Expert-level proficiency with Kubernetes (EKS), including networking, ingress controllers, and service mesh management.
- Automation: Strong experience with configuration management and IaC (e.g., Terraform, Ansible, Helm).
- Data Systems: Deep knowledge of SQL and NoSQL database administration, focusing on replication, backup, and disaster recovery.
- Programming: Proficiency in Python and C++ for developing internal tooling and automating complex operational workflows.
- Systems Internals: Strong understanding of Linux networking, storage, and kernel tuning.
- Prior experience in Aerospace, Defense, or high-reliability sectors.
- Familiarity with CCSDS standards or satellite ground station software.
- Experience with secure, air-gapped, or hybrid-cloud deployments.
For U.K. Roles: To comply with U.K. regulations, this role requires Baseline Personnel Security Standard (BPSS) checks, and successful candidates must be eligible to obtain UK Security Clearance (SC).
For Canada Roles: Successful candidates must obtain and hold a security clearance at the reliability status level, and pass security assessment for the Canadian Controlled Goods Program (CGP) and ITAR.
We celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or veteran status.
Skills
SQL
AWS
ELK
Service Mesh
Ansible
Azure
CloudFormation
C++
Google Cloud
Grafana
Helm
Kubernetes
Prometheus
Python
Terraform
Similar jobs
Senior Cloud & Platform Engineer
M&T BANK CORPORATION · Buffalo, United States
36 minutes ago$139.7k - $232.9k/yrPlatform Engineer
QuantumScape · United States
37 minutes ago$125.2k - $181.6k/yrStaff Engineer, XBAT DevOps (R4542) (Dallas, TX)
Shield AI Inc · Dallas, United States
37 minutes ago$158k - $215k/yrStaff Site Reliability Engineer
Visa Inc. · Austin, United States
37 minutes ago$131.6k - $210.3k/yrSenior DevOps Engineer, K8
NxT Level · San Francisco, United States
40 minutes agoAI Platform Engineer (PPA - 008)
SageCor Solutions · Columbia, United States
40 minutes ago