Quick Overview
Job Description
Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies. With more than 3,000 customers and ARR that has grown over 250 percent year over year, ClickHouse leads the market in real-time analytics, data warehousing, observability, and AI workloads.
The company's sustained, accelerating momentum was recently validated by a $400M Series D financing round. Over the past three months, customers including Capital One, Lovable, Decagon, Polymarket, and Airwallex have adopted the platform or expanded existing deployments. These customers join an established base of AI innovators and global brands such as Meta, Cursor, Sony, and Tesla.
We're on a mission to transform how companies use data. Come be a part of our journey!
Note: This position can be based remotely in any country where ClickHouse has a hiring presence.
We are committed to providing our customers with reliable and secure services at ClickHouse. To continue this, we are building out our Site Reliability Engineering team in ClickHouse Core. As one of the first members of our Reliability Engineering Team at Core, you will be responsible for building and leading processes to ensure and improve the reliability, availability, scalability, and performance of ClickHouse. You will collaborate with different teams like Control Plane, Dataplane,Security, Support and Operations and guide them to implement ClickHouse in the best way for our customers. You will also own the areas of managing engineering escalation management and response, investigations, post-mortem analysis including running blameless postmortems, and continuous improvement of how Clickhouse is run and optimized in the cloud. This role is a unique opportunity to make a significant impact on our elastic, limitless scale, high-performance ClickHouse in ClickHouse Cloud.
What will you do?- Continuously improve the reliability and performance of ClickHouse core.
- Improve and create metrics and alerts for ClickHouse to be able to identify and prevent problems in production before they affect customers.
- Dig deeper into the most common problems encountered by customers in Clickhouse Core to identify the root cause of problems and submit bug fixes, issue reports and suggest improvements.
- Enhance and refine incident response processes and post-mortem analysis for ClickHouse core related outages including working with support and Cloud teams to communicate to the impacted customers.
- Plan, enable, and drive Chaos initiatives across Engineering teams, based upon internal priorities.
- Manage on-call processes to respond to performance and reliability issues, and establish best practices for coordinating escalation to resolve issues and minimize customer impact.
- Bachelor's or Master's degree in Co
Similar jobs
- VH
Lead Software Engineer
NewVisa Hunt
Australia🇦🇺Remote54 minutes agoTechnology - KL
Salesforce Engineer
NewKing Living
Sydney🇦🇺On-site54 minutes agoSOAPGitRESTTechnology - RE
Senior DevOps Engineer
NewReapit
Brisbane, Queensland🇦🇺Hybrid54 minutes agoAWSAnsibleCDK+3Technology - CA
SAP BTP IS Integration Consultant
NewCapgemini
Melbourne, Victoria🇦🇺Hybrid54 minutes agoSOAPGroovyHTTP+1Technology - VA
Data Scientist, Ai
NewVirgin Australia Airlines
Brisbane, Queensland🇦🇺Hybrid54 minutes agoMachine LearningNLPScikit-learn+4Technology - RM
Hydraulic Sales & System Engineer
NewRCR Mining Technologies
Western Australia🇦🇺Hybrid54 minutes agoTechnology