Principal Observability Platform Engineer
Quick Overview
Job Description
ABOUT THE ROLE
We''re hiring a Principal/Staff Observability Platform Engineer to own the technical direction of our observability platform — the systems that give us deep visibility into GPU clusters, AI workloads, and the infrastructure running them.
This is a "define, build, and lead" role, not a "maintain and operate" role. You''ll set the architectural roadmap, raise the engineering bar across teams, and make sure the platform scales ahead of the business, not behind it. We have a strong bias toward simplicity — the systems you build should be easy to operate and self-evidently correct when something goes wrong.
RESPONSIBILITIES
- Own the technical strategy and architecture for observability across metrics, logs, traces, and alerting at scale
- Drive platform decisions with multi-year impact: tooling, data models, ingestion patterns, retention, cardinality management
- Identify systemic gaps before they become incidents; design platforms that make failure visible and fast to diagnose
- Partner with SRE, infrastructure, and AI/ML teams to embed observability natively into how we build and operate
- Define standards and patterns that other engineers adopt because they''re clearly better, not by mandate
- Mentor and technically grow the observability team
- Lead incident postmortems and drive durable platform improvements
- Evaluate and introduce tooling that improves signal quality, operational efficiency, or scalability — and retire what doesn''t
REQUIRED SKILLS & EXPERIENCE
- 8+ years in SRE, infrastructure engineering, platform engineering, or observability-focused roles
- Proven experience operating observability infrastructure at serious scale
- Deep hands-on experience with a significant subset of: Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, Elastic
- Strong engineering fundamentals in Python, Go, or similar
- Kubernetes at scale
- Infrastructure-as-Code as default practice (Terraform, Ansible, or equivalent)
- Demonstrated ability to architect systems, write code, review others'' work, and clearly explain tradeoffs
- Track record of influencing engineering direction across teams without formal authority
PREFERRED SKILLS & EXPERIENCE
- Experience with high-volume streaming pipelines for observability data (Kafka, Vector, Fluent Bit, etc.)
- Background in AI/ML infrastructure observability: GPU utilization, training job visibility, inference latency
- Familiarity with GPU infrastructure or HPC environments (Slurm)
- Prior experience defining observability strategy at an organizational level
EQUAL OPPORTUNITY
We strongly encourage applications from people of color, the LGBTQ+ community, people with disabilities, neurodivergent individuals, parents, carers, and people from lower socio-economic backgrounds. If there''s anything we can do to accommodate your specific situation, please let us know.
Skills
Similar jobs
Lead AWS Cloud Platform Engineer
nTech Solutions · Reston, United States
22 minutes agoDevSecOps Engineer
Vantor · Herndon, United States
23 minutes agoSystems Engineer
TEKsystems c/o Allegis Group · Grand Haven, United States
23 minutes ago$90k - $115k/yrDevOps Engineer
Kforce Technology Staffing · Irving, United States
23 minutes agoPCF (Pivotal Cloud Foundry) Platform Engineer
HCLTech · Alpharetta, United States
25 minutes agoDevOps Engineer
M9 Solutions · Chantilly, United States
25 minutes ago$60k - $180k/yr