Why This Role Stands Out
You'll have a significant impact leading cloud platform architecture for critical aviation systems, driving innovation and scaling for global operations within a highly reputable company. This role is perfect for a hands-on, mid-senior engineer passionate about building resilient infrastructure and mentoring a team, offering immense growth potential in a dynamic and safety-critical domain.
Quick Overview
Job Description
About ASI
ASI's mission-critical technology powers decision-making across aviation, defense, energy, and other critical infrastructure domains. Backed by top-tier investors including Andreessen Horowitz, Spark Capital, and Renegade Partners, ASI delivers operational decision superiority—compressing days of analysis into seconds of action. ASI is leading the way and pushing the boundaries of what’s possible.
About the Civil Aviation Team
We are modernizing America’s air traffic control system on an accelerated timeline, tackling one of the world’s most complex real-time optimization challenges: safely coordinating tens of thousands of daily flights in a dynamic, safety-critical environment. Our team is building next-generation trajectory prediction and decision-support systems while preparing the airspace for emerging technologies like commercial space operations and advanced air mobility. This mission-critical work demands exceptional precision, resilience, and reliability to support the safety and efficiency of the national airspace.
What You Will Do:
As a Lead Cloud Platform Engineer for the Civil Aviation team, you will own the architecture, reliability, and evolution of a cloud platform supporting mission-critical aviation systems. You’ll lead a small team while remaining deeply hands-on, building and operating infrastructure that must scale globally, maintain ~99.99% availability, and support safe, continuous delivery in a 24/7 environment.
You will define and drive platform architecture, establishing best practices for infrastructure-as-code, Kubernetes-based systems, deployment pipelines, and production operations. You’ll own reliability end-to-end, designing for failure, implementing observability (metrics, logging, tracing), and leading incident response, postmortems, and continuous improvement efforts.
In addition to technical leadership, you will manage and develop a small team of engineers—setting clear direction, establishing high standards, and creating the systems (metrics, processes, feedback loops) that enable teams to operate autonomously while maintaining a strong culture of reliability and ownership.
What We Value:
Strong experience building and operating highly reliable cloud platforms (e.g., ~99.99% uptime).
Deep expertise in AWS, with a strong understanding of multi-region and highly available architectures.
Proficiency with Kubernetes and containerized environments, along with infrastructure-as-code tools (e.g., Terraform, Helm).
Experience designing and maintaining CI/CD pipelines that support safe, frequent production releases.
Hands-on experience with progressive delivery techniques (e.g., canary deployments, safe rollback strategies).
Strong understanding of observability and monitoring systems (e.g., Grafana, Prometheus).
Experience leading small teams (1-2+ years) while remaining hands-on.
Ability to own platform architecture end-to-end, from design through production operations.
Proven track record of driving reliability culture and production excellence across teams.
Experience operating and supporting 24/7 production systems.
Strong ability to lead design reviews and make architectural decisions in distributed environments.
Experience with FedRAMP High environments or other regulated cloud systems is a plus.
Exposure to high-availability systems (e.g., “four nines” reliability) is a strong plus.
Proficient in leveraging modern LLM tools to accelerate development workflows and enhance code quality.
How We Hire:
We look at the interview process not as a screening or test, but rather as an opportunity to simulate what it would look like working together. We build the interview process around you.
ASI works with export-controlled technology and restricted U.S. Government data, including on contracts mandating U.S. immigration status and location restrictions for performing personnel. Employment offers are contingent on ability to timely obtain all required authorizations for contemplated job duties.
Similar jobs
- VA
Site Reliability Engineer
NewVannevar
Remote; San Diego🇺🇸Remote6 hours agoDockerSQLAWS+9Technology - RS
Site Reliability Engineer
NewRedwood Software
United States (Remote)🇺🇸Remote3 hours agoDockerAWSBash+13Technology - PU
Senior Site Reliability Engineer
NewPrecisely US Jobs
United States🇺🇸3 hours agoDockerOracleAWS+13Technology - AK
Core Production Engineer OR SRE Engineer
NewAkshaya Inc
San Francisco, CA🇺🇸Hybrid20 hours agoRoot Cause AnalysisManufacturing - ST
Salesforce DevOps Architect
NewShiro Technologies
United States🇺🇸Hybrid20 hours agoAWSSnowflakeGitTechnology - CS
Senior DevOps Engineer (Automation & Middleware)
NewCynet Systems
Oakland, CA🇺🇸$45 - $50/hrHybrid20 hours agoDockerShellSOC 2+8Technology