Quick Overview
Job Description
Data Engineer - Level 2
Onsite, Cincinnati, OH
Your role will involve collaboration with cross-functional teams, leveraging cutting-edge technologies, and ensuring scalable, efficient, and secure data engineering practices. You’ll be working with data across 2 cloud platforms (Azure and Google Cloud Platform) and thus will be working with a wide variety of technologies. A strong emphasis will be placed on expertise in python, Vertex AI, and advanced feature engineering techniques.
Responsibilities Take ownership of systems, processes, and the tech stack while driving features to completion through all phases of the entire 84.51° SDLC. This includes internal and external facing applications as well as process improvement activities:
· Build and Maintain Data Pipelines: Design, build, and maintain scalable, efficient, and reliable data pipelines to support data ingestion, transformation, and integration across diverse sources and destinations, using tools such as Kafka, Databricks, and similar toolsets.
· Drive Digital Innovation: Leverage innovative technologies and approaches to modernize and extend core data assets, including SQL-based, NoSQL-based, cloud-based, and real-time streaming data platforms.
· Implement Feature Engineering: Develop and manage feature engineering pipelines for machine learning workflows, utilizing tools like Vertex AI, BigQuery ML, and custom Python libraries.
· Implement Automated Testing: Design and implement automated unit, integration, and performance testing frameworks to ensure data quality, reliability, and compliance with organizational standards.
· Optimize Data Workflows: Optimize data workflows for performance, cost efficiency, and scalability across large datasets and complex environments.
· Draft and Review Documentation: Draft and review architectural diagrams, interface specifications, and other design documents to ensure clear communication of data solutions and technical requirements.
Key Responsibilities
Requirements:
Bachelor’s degree typically in Computer Science, Management Information Systems, Mathematics, Business Analytics or another STEM degree.
4+ years of professional Data Development experience.
4+ years of experience with SQL and NoSQL technologies.
3+ years of experience building and maintaining data pipelines and workflows.
2+ years of experience developing with Python.
Experience in feature engineering for machine learning pipelines.
Experience with CI/CD pipelines and processes.
Experience with automated unit, integration, and performance testing.
Experience with version control software such as Git.
Full understanding of ETL and Data Warehousing concepts.
Strong understanding of Agile principles (Scrum).
Preferred Qualifications
· Knowledge of Structured Streaming (Spark, Kafka, EventHub, or similar technologies).
· Experience with Google Cloud Platform services (or Databricks equivalent) such as BigQuery, Vertex AI Platform, Cloud Storage, AutoMLOps, and Dataflow.
· Experience with GitHub SaaS/GitHub Actions.
· Experience understanding Databricks concepts.
· Experience with PySpark and Spark development.
· Experience with Service Oriented Architecture.
Similar jobs
- HT
Senior Data Engineer- Snowflake+ DBT-
NewHeadway Tek Inc
United States🇺🇸Remote19 hours agoSQLT-SQLAWS+6Technology - AM
Sr. Data (cProbe) Engineer with Security Clearance
NewAmyx
Fort Belvoir, VA🇺🇸Hybrid19 hours agoEngineering - DN
Database Engineer DE3 [D.26.0078] with Security Clearance
NewDover Networks LLC
Southern Md Facility, MD🇺🇸$174k - $190k/yrHybrid19 hours agoOracleBashTechnology - AP
Oracle DBA
NewApptad Inc
New York, NY🇺🇸On-site19 hours agoOracleSQLShellTechnology - KT
Data Engineer (PL/SQL)
NewKforce Technology Staffing
Durham, NC🇺🇸HybridYesterdayMongoDBOraclePL/SQL+10Technology - PT
Data Engineer
NewPrudent Technologies and Consulting
Plano, TX🇺🇸Hybrid19 hours agoSQLSnowflakePython+1Technology