Haystack
← Back to Jobs
Engineering
II

Druid Engineer

Innovative IT Solutions IncDallas, TX🇺🇸United StatesPosted 31 Aug 2026

Why This Role Stands Out

This hybrid Druid Engineer role at Innovative IT Solutions Inc. offers a fantastic opportunity to architect and optimize large-scale data platforms, fostering significant career growth in a dynamic environment. If you excel at problem-solving and enjoy building robust, high-performance data solutions, this position is perfect for you to leverage your expertise and make a real impact. You'll thrive here by contributing to cutting-edge projects and collaborating with a supportive team, all while enjoying competitive compensation.

Quick Overview

Salary
$50/hr
Seniority
Mid Senior
Work mode
Hybrid
Location
Dallas, TX, United States
Posted
Yesterday

Job Description

Druid Engineer

Location: Dallas, TX
Rate: $55/hr C2C
Employment Type: Contract

Job Description

We are seeking an experienced Druid Engineer to design, deploy, administer, optimize, and support large-scale Apache Druid environments supporting PB-scale datasets. The ideal candidate will have strong expertise in Druid cluster administration, performance tuning, ingestion, monitoring, automation, and production support.

Key Responsibilities

  • Design, deploy, configure, administer, and optimize large-scale Apache Druid clusters supporting PB-scale analytical workloads.
  • Manage, monitor, and troubleshoot Apache Druid services, clusters, and distributed components.
  • Perform Druid cluster upgrades, patching, capacity planning, scaling, and platform modernization.
  • Monitor cluster health, ingestion performance, query latency, segment distribution, resource utilization, and overall platform performance.
  • Troubleshoot ingestion failures, stuck tasks, compaction issues, retention policies, segment management, and indexing problems.
  • Optimize Druid queries, partitioning strategies, indexing specifications, segment allocation, and compaction configurations.
  • Configure and maintain Druid metadata stores using MySQL, including metadata management and database connectivity.
  • Develop operational automation using Ansible, Infrastructure as Code (IaC), and configuration management practices.
  • Build and maintain Grafana dashboards, Prometheus monitoring, alerts, metrics, and centralized logging solutions.
  • Implement and support High Availability (HA), Disaster Recovery (DR), backup/recovery, security, and compliance requirements.
  • Participate in production support, incident management, root cause analysis (RCA), troubleshooting, and performance tuning.
  • Collaborate with DevOps, SRE, platform engineering, architects, data engineering, and business teams to translate analytical requirements into scalable Druid solutions.
  • Support Druid performance optimization, scalability, reliability, availability, and operational excellence across enterprise environments.

Required Skills

  • Strong hands-on experience with Apache Druid / Druid Engineering.
  • Experience administering large-scale distributed Druid clusters and PB-scale datasets.
  • Strong knowledge of Druid ingestion, indexing, segments, partitioning, compaction, retention, and query optimization.
  • Experience with Druid MiddleManager/Indexer, Historical, Broker, Coordinator, Overlord, Router, and Metadata Store components.
  • Strong experience with MySQL for Druid metadata management.
  • Experience with Ansible and Infrastructure as Code (IaC).
  • Hands-on experience with Grafana, Prometheus, monitoring, alerting, metrics, and logging.
  • Strong troubleshooting, performance tuning, capacity planning, and production support experience.
  • Knowledge of High Availability, Disaster Recovery, backup/recovery, security, and compliance.

 

Apache Druid, Druid Engineer, Druid Administrator, Druid Cluster Administration, Druid Cluster Management, Druid Performance Tuning, Druid Optimization, Druid Ingestion, Druid Indexing, Druid Segments, Druid Compaction, Druid Partitioning, Druid Query Optimization, Druid Metadata Store, MySQL, PB-Scale Data, Distributed Systems, Big Data, Data Analytics, Ansible, Infrastructure as Code, IaC, Grafana, Prometheus, Monitoring, Alerting, Logging, High Availability, Disaster Recovery, Backup & Recovery, Capacity Planning, Cluster Scaling, Production Support, Incident Management, Root Cause Analysis, Performance Engineering, DevOps, SRE, Cloud Infrastructure.

Similar jobs