Haystack
← Back to Jobs
Technology
MT

Junior Graph Data Engineer with Security Clearance

Marathon TS IncArlington, VA🇺🇸United StatesPosted Sep 29, 2026

Quick Overview

Salary
$85k - $105k/yr
Seniority
Junior
Work mode
Hybrid
Location
Arlington, VA, United States
Posted
Yesterday
SQLETLJavaLLMPython

Job Description

Junior Graph Data Engineer Pentagon or Reston VA Clearance: Active TS/SCI Required ($85,000 - $105,000) Marathon TS is seeking a motivated Junior Graph Data Engineer to support a mission-focused data and AI initiative within a Department of Defense (DoD) environment. This is an excellent opportunity for an early-career engineer to gain hands-on experience working with graph databases, data integrations, APIs, semantic data, and emerging AI technologies.

The Junior Graph Data Engineer will support the development and maintenance of an enterprise semantic data environment designed to make complex information more discoverable, understandable, trusted, and accessible. Working alongside experienced graph, data, and software engineers, this role will help integrate disparate data sources, configure graph data structures, and support automated workflows that enable advanced search, analytics, and AI-driven research.

Key Responsibilities 1. Automated Source Discovery && Metadata Ingestion (Technical Metadata) *

  • Supplying the "Raw Ingredients" for the Semantic Knowledge Graph: Assist in designing and deploying automated pipelines that programmatically Client enterprise data assets and interface with existing data catalogs. Scan, catalog, and ingest technical metadata - including schemas, tables, columns, and API endpoints - from legacy, cloud, and distributed environments to establish baseline assets for alignment to the Enterprise Core Ontology.
  • Scaling the Semantic Map: Use automated pipelines and orchestrated workflows to ingest metadata at scale rather than relying on manual, field-by-field mapping. Help keep the ontology current as a dynamic, living "semantic contr ol plane" rather than a static document.
  • Establishing the Entry Point for Lineage: Register the technical origin of ingested data and capture metadata at the point of ingestion. This creates the foundation for automated provenance chains that track where data originated and how it changes over time. 2. Semantic && Provenance Mapping (Semantic && Lineage Metadata) *
  • Ontological Alignment Support: Collaborate with senior engineers to align discovered data elements from local systems to the shared Enterprise Core Ontology and specialized Domain Ontologies, with particular attention to compatibility with established institutional frameworks (e.g., DIA's DIKEM). Preserve local naming conventions while establishing standardized, shared meaning.
  • Lineage Tracking Support: Help engineering teams construct and maintain data lineage chains within the Provenance Layer, following applicable industry lineage standards to document where data originates, how it is transformed, and who governs it.
  • Graph Querying Support (Growth Area): As your technical skills develop, write and test basic graph queries to support metadata retrieval, logical validation, and graph manipulation. 3. Enterprise Systems Thinking && Alignment *
  • Big-Picture Integration: Evaluate how newly integrated data sources and automated pipelines affect the broader Enterprise Semantic Map, selected use cases, downstream consumers, and enterprise search and discovery.
  • Downstream Awareness: Connect data assets to relevant mission metadata so technical capabilities can be clearly linked to the mission workflows they support.
  • Governance Compliance: Help ensure enterprise assets are associated with appropriate governance metadata, including ownership, classifications, handling rules, and access constraints. Support the translation of complex data policies into machine-readable semantic structures. 4. Smart Search && Agent Enablement *
  • Semantic Contr ol Plane Maintenance: Help maintain and optimize the Enterprise Semantic Map within enterprise graph database platforms so human analysts, applications, and autonomous AI agents can efficiently search, navigate, and Client resources.
  • Agent Integration Support: Collaborate with AI engineers to help planning, research, and tool agents dynamically query the graph and build grounded, trustworthy reasoning and retrieval strategies. Required Experience/Clearance
  • Bachelor's Degree with 1 of relevant professional experience or equivalent.
  • Active TS SCI Clearance.
  • Core Technical Skills: Foundational proficiency across the following areas, demonstrated in any comparable technology:
  • Programming and scripting for automation (e.g., Python, Java, or a comparable general-purpose language)
  • Relational database querying (e.g., SQL)
  • Structured and semi-structured data formats (e.g., JSON, XML, YAML)
  • Knowledge graph concepts, including nodes, edges, relationships, and metadata schemas
  • Foundational Data Engineering: Basic understanding of data structures, databases, and how data moves through pipelines or ETL (Extract, Transform, Load) processes.
  • Systems-Thinking Mindset: Ability to understand how individual data pipelines connect to and support a broader enterprise ecosystem.
  • Attention to Detail: Precision in aligning metadata terms, formatting data endpoints, and maintaining technical schemas.
  • Collaboration && Communication: Ability to take direction from senior engineers, document work clearly, and explain technical decisions to non-specialist stakeholders.

Preferred Qualifications

  • Graph Query Languages: Exposure to - or willingness to learn - graph query languages for metadata retrieval and validation (e.g., Cypher for property graphs, SPARQL for RDF/triple stores).
  • Graph Database Platforms: Conceptual familiarity with modern enterprise graph database platforms.
  • Agentic AI && AI Frameworks: Basic conceptual understanding of, coursework in, or project experience with LLM orchestration or agentic workflows.
  • Data Lineage && Metadata Standards: Exposure to open lineage specifications or metadata management frameworks.
  • Standard Ontologies && Semantic Models: Conceptual familiarity with established government- or defense-related semantic models that support standardized enterprise data integration.
  • Data Catalogs: Familiarity with metadata catalog environments and data stewardship systems.
  • Workflow Orchestration: Exposure to pipeline scheduling and orchestration tooling.

Marathon TS is committed to the development of a creative, diverse and inclusive work environment. In order to provide equal employment and advancement opportunities to all individuals, employment decisions at Marathon TS will be based on merit, qualifications, and abilities. Marathon TS does not discriminate against any person because of race, color, creed, religion, sex, national origin, disability, age or any other characteristic protected by law (referred to as "protected status"). #CJJOBS

Similar jobs