Quick Overview
Job Description
Lead Platform Engineer
Location: South San Francisco / San Mateo, CA
Work Arrangement: Hybrid — 3 days onsite per week
Employment Type: Full-Time
Openings: 2
Position Overview
We are seeking two hands-on Lead Platform Engineers to help build and scale the backend and cloud infrastructure supporting a global DNA-sequencing operation.
This is a backend/platform engineering leadership role, not a traditional DevOps, IT infrastructure, data engineering, bioinformatics, or ML model-training position.
The ideal candidate is a Senior or Staff-level engineer with strong Python, AWS, distributed-systems, and cloud-orchestration experience, who has worked in an early-stage startup environment and personally built systems from the ground up.
The role will initially be highly hands-on, with the engineer learning the existing systems and taking ownership of major technical areas. Over time, the position will transition into approximately:
- 70% hands-on architecture and software engineering
- 30% technical leadership, mentoring, and team leadership
What You'll Build
The platform supports a globally distributed sequencing operation involving:
- Laboratory robots and DNA sequencers
- On-premises Linux infrastructure
- AWS cloud environments
- Terabytes of sequencing data movement
- Hundreds of thousands to millions of bioinformatics jobs
- Distributed queues and asynchronous workloads
- Workflow orchestration and scheduling
- Retry and failure-handling mechanisms
- Distributed state management
- Fault-tolerant production systems
- Large-scale compute infrastructure
- AI agents that interact with internal tools, APIs, operational data, and physical workflows
The infrastructure currently supports approximately 5,000 CPUs, 12.5 TB of RAM, and 100+ GPUs, with significant production AWS usage.
Key Responsibilities
- Design, build, and operate high-scale distributed backend and platform systems.
- Develop production backend services primarily using Python.
- Build and maintain cloud-native services and infrastructure on AWS.
- Design systems for asynchronous processing, distributed queues, workflow orchestration, scheduling, retries, state management, and fault tolerance.
- Build reliable systems for moving large volumes of data between laboratory equipment, on-premises infrastructure, and AWS.
- Develop services that orchestrate large numbers of compute and bioinformatics workloads.
- Design and implement scalable APIs, microservices, workers, and event-driven services.
- Work across application code, cloud infrastructure, Linux systems, data movement, and operational tooling.
- Architect and implement production systems from concept through deployment and ongoing operation.
- Troubleshoot complex production issues and improve system reliability, performance, and scalability.
- Lead major technical projects from design through production.
- Establish engineering patterns, standards, and best practices for a growing platform team.
- Mentor and provide technical guidance to engineers while remaining deeply hands-on.
- Collaborate with engineering, infrastructure, security, and scientific teams.
- Explore and implement practical applications of AI agents and modern AI development tools within production systems.
Required Qualifications
- 6+ years of professional experience in backend engineering, platform engineering, infrastructure engineering, distributed systems, or closely related software engineering roles.
- Strong professional experience with Python backend development.
- Strong hands-on experience with AWS cloud infrastructure and services.
- Significant experience designing and building distributed production systems.
- Hands-on experience with task queues, asynchronous processing, event-driven systems, workflow orchestration, scheduling, retries, state management, and fault tolerance.
- Experience building systems from scratch or from early-stage prototypes through production.
- Meaningful experience working in an early-stage startup or rapidly scaling engineering environment.
- Demonstrated ownership of major technical systems or projects.
- Experience scaling systems, infrastructure, workloads, or engineering platforms as a company grows.
- Experience providing technical leadership and mentoring engineers.
- Strong Linux and cloud-infrastructure fundamentals.
- Ability and willingness to remain approximately 70% hands-on with architecture, coding, debugging, and production engineering.
Required Technical Skills
Candidates should have strong experience with the following core technologies and concepts:
Backend & Programming
- Python
- FastAPI, Django, or comparable Python backend frameworks
- REST APIs
- Microservices
- Event-driven architecture
- Asynchronous processing
- Backend service development
AWS & Cloud
- Amazon Web Services (AWS)
- ECS
- AWS Batch
- AWS Step Functions
- SQS
- Lambda
- Cloud-native architecture
- Infrastructure as Code
- Production cloud environments
Distributed Systems & Orchestration
- Distributed systems
- Distributed computing
- Task queues
- Job queues
- Message queues
- Workflow orchestration
- Job orchestration
- Scheduling
- Retry mechanisms
- State management
- Fault tolerance
- Failure recovery
- High-volume workload processing
- Event-driven services
Messaging & Distributed Processing
Experience with one or more of:
- Celery
- Amazon SQS
- Kafka
- RabbitMQ
- gRPC
- Similar distributed messaging or task-processing technologies
Infrastructure
- Linux
- Docker / containerized environments
- Cloud infrastructure
- Infrastructure automation
- Production monitoring and troubleshooting
Similar jobs
- CA
Test Automation & DevOps Engineer with Security Clearance
CACI
Chantilly, VA🇺🇸$94.4k - $198.2k/yrHybrid5 weeks agoDockerMongoDBRust+22Technology - NT
DevOps Applications Developer with Security Clearance
NewNewGen Technologies
Herndon, VA🇺🇸Hybrid19 hours agoDockerRubyShell+8Technology - TS
W2 - Infrastructure/DevOps Engineer Sr. - GJ2026
NewTechLink Systems, Inc.
Marysville, OH🇺🇸On-site19 hours agoAWSAnsibleAzure+7Technology - TE
Sr. Devops Engineer with Salesforce and Copado
NewTECHProjects
Trenton, NJ🇺🇸On-site19 hours agoDockerAWSAnsible+5Technology - LT
Google Cloud Platform DATA Platform Engineer in Charlotte, NC (Onsite)-Contract
NewLorven Technologies, Inc.
Charlotte, NC🇺🇸On-site19 hours agoApacheGoogle CloudRESTTechnology - PE
HPC Software Deployment Configuration Manager, Lead with Security Clearance
Peraton
Fort Meade, MD🇺🇸$104k - $166k/yrHybrid3 weeks agoHTTPSTechnology