Haystack
← Back to Jobs
Remote
Engineering

Senior Knowledge Graph Engineer

Empower ProfessionalsUnited States🇺🇸United StatesPosted 30 Jul 2026

Why This Role Stands Out

This remote Senior Knowledge Graph Engineer role offers a fantastic opportunity to leverage your deep expertise in life sciences ontologies and semantic web technologies to build impactful knowledge graphs within the innovative life sciences domain. If you thrive on solving complex data challenges and shaping the future of information, this position is an excellent fit for your advanced skills and career aspirations.

Quick Overview

Work Type
Remote
Level
Mid Senior

Job Description

Role: Senior Knowledge Graph Engineer

Duration: 12 Months

Location: Remote (For USA)

Domain Life Sciences

Basic Qualifications:

  • We are looking for professionals with the required skills to achieve our goals:
  • Masters degree in Biosciences, Biomedical Science, Biomedical Engineering, Biotechnology (with a life science/pharma application focus)
  • 6+ years of relevant knowledge graph work experience
  • Specific hands-on experience contributing to Knowledge Graph development efforts, including entity modeling, relationship design, r2rml, and schema governance
  • Hands-on experience with open-source ontology tools and languages: Protg, SPARQL, OWL, SKOS, SHACL, RML, RDF-start
  • Working knowledge of major life sciences ontologies: Gene Ontology (GO), OBO Foundry ontologies (CL, UBERON, HPO, MONDO, CHEBI, EFO, CLO), MeSH, SNOMED CT, UMLS
  • Familiarity with linked data principles and semantic web technologies
  • Demonstrated experience with industry-standard tools for building data serialization protocols (e.g., JSON Schema, LinkML)
  • Proficiency in at least one programming language preferably Python and LLM for scripting vocabulary mappings, building data models, automating QC, and prototyping pipelines

Preferred Qualifications:

  • If you have the following characteristics, it would be a plus:
  • Experience with data governance and data quality tooling (e.g., Ataccama, Informatica, Talend, OpenRefine, Great Expectations, dbt)
  • Experience with at least one programming language e.g. Python for scripting vocabulary mappings, building data models, etc
  • Experience supporting LLM integration or AI-readiness workflows including metadata enrichment, entity linking, embedding pipelines, or retrieval-augmented generation (RAG) architectures
  • Understanding of vector databases and their role in semantic search and knowledge retrieval (e.g., Weaviate, Chroma)
  • Familiarity with cloud data platforms and infrastructure relevant to large-scale biological data (e.g., AWS, Google Cloud Platform, Azure)
  • Familiarity with graph database technologies (e.g., Neo4j, Amazon Neptune, Stardog, GraphDB, TigerGraph)

Similar jobs