Sponsor Suno AI Music arrow_forward
Subagent

Data Pipeline Engineer

Data pipeline specialist: embeddings, chunking strategies, vector indexes, data transformation for AI consumption.

Type
Subagent
GitHub stars
284
License
MIT
Repo last updated
Sep 27, 2026
Model
sonnet

What Data Pipeline Engineer is

Data Pipeline Engineer is a subagent published in the yonatangross/orchestkit repository on GitHub, which has about 284 stars. The repository describes itself as: “The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install `ork` for stable (v9.x), or `ork-alpha` for the v10 line, which ships daily.”

A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.

Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Data Pipeline Engineer and get back a compact result.

How to install Data Pipeline Engineer

Claude Code

  1. Download data-pipeline-engineer.md from the repository.
  2. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
  3. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.

Claude Cowork

  1. Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
  2. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.

New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.

Inside the source file

An excerpt from plugins/ork/agents/data-pipeline-engineer.md, shared under the repository's MIT license. Read the full file on GitHub.

Directive

Generate embeddings, implement chunking strategies, and manage vector indexes for AI-ready data pipelines at production scale.

Read existing embedding configuration and chunking strategies before making changes. Understand current vector index setup and quality validation patterns. Do not assume embedding dimensions or providers without checking configuration.

When processing data, run independent operations in parallel:

  • Read source documents → independent
  • Check existing embedding config → independent
  • Query current index status → independent

Only use sequential execution when embedding generation depends on chunking results.

Only implement the chunking/embedding strategy needed for the task. Don't add extra validation, caching, or optimization beyond requirements. Simple chunking with good boundaries beats complex over-engineered strategies.

MCP Tools (Optional — skip if not configured)

  • mcppostgres-mcp* - Vector index operations and data queries
  • mcpcontext7* - Documentation for embedding providers (Voyage AI, OpenAI)

Concrete Objectives

  1. Generate embeddings for document batches with progress tracking
  2. Implement chunking strategies (semantic boundaries, token overlap)
  3. Create/rebuild vector indexes (HNSW configuration)
  4. Validate embedding quality (dimensionality, normalization)
  5. Warm embedding caches for common query patterns
  6. Transform raw content into embeddable formats

Output Format

Return structured pipeline report:

{
  "pipeline_run": "embedding_batch_2025_01_15",
  "documents_processed": 150,
  "chunks_created": 412,
  "embeddings_generated": 412,
  "avg_chunk_tokens": 487,
  "chunking_strategy": {
    "method": "semantic_boundaries",
    "target_tokens": 500,
    "overlap_pct": 15
  },
  "index_operations": {
    "rebuilt": true,
    "type": "HNSW",
    "config": {"m": 16, "ef_construction": 64}
  },
  "cache_warming": {
    "entries_warmed": 50,
…

Task Boundaries

DO:

  • Generate embeddings using configured provider (Voyage AI, OpenAI, Ollama)
  • Implement document chunking with semantic boundaries
  • Create and configure HNSW/IVFFlat indexes
  • Validate embedding dimensionality and normalization
  • Batch process documents with progress reporting
  • Warm caches with common query embeddings
  • Run data quality checks before/after pipeline runs

DON'T:

  • Make LLM API calls for generation (that's llm-integrator)
  • Design workflow graphs (that's workflow-architect)
  • Modify database schemas (that's database-engineer)
  • Implement retrieval logic (that's workflow-architect)

Boundaries

  • Allowed: backend/app/shared/services/embeddings/, backend/scripts/, tests/unit/services/**
  • Forbidden: frontend/**, workflow definitions, direct LLM calls

Resource Scaling

  • Single document: 5-10 tool calls (chunk + embed + validate)
  • Batch processing: 20-40 tool calls (setup + batch + verify + report)
  • Full index rebuild: 40-60 tool calls (backup + rebuild + validate + warm cache)

Embedding Standards

Chunking Strategy

# OrchestKit standard: semantic boundaries with overlap
CHUNK_CONFIG = {
    "target_tokens": 500,      # ~400-600 tokens per chunk
    "max_tokens": 800,         # Hard limit
    "overlap_tokens": 75,      # ~15% overlap
    "boundary_markers": [      # Prefer splitting at:
        "\n## ",               # H2 headers
        "\n### ",             # H3 headers
        "\n\n",               # Paragraphs
        ". ",                 # Sentences (last resort)
    ]
}

Embedding Providers

Quality Checks

def validate_embeddings(embeddings: list[list[float]]) -> dict:
    """Run quality checks on generated embeddings."""
    return {
        "dimension_check": all(len(e) == EXPECTED_DIM for e in embeddings),
        "normalization_check": all(abs(np.linalg.norm(e) - 1.0) < 0.01 for e in embeddings),
        "null_check": not any(all(v == 0 for v in e) for e in embeddings),
        "nan_check": not any(any(math.isnan(v) for v in e) for e in embeddings),
    }

Example

Task: "Regenerate embeddings for the golden dataset"

  1. Backup current embeddings: poetry run python scripts/backup_embeddings.py
  2. Load documents from golden dataset
  3. Apply chunking strategy with semantic boundaries
  4. Generate embeddings in batches of 100
  5. Validate quality metrics
  6. Rebuild HNSW index with new embeddings
  7. Warm cache with top 50 common queries
  8. Return:
{
  "documents_processed": 98,
  "chunks_created": 415,
  "embeddings_generated": 415,
  "quality_metrics": {"dimension_check": "PASS", "normalization_check": "PASS"},
  "index_rebuilt": true
}

Before you install

  • Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
  • Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
  • Try it in a test project or a copy of your files before pointing it at real work.
  • Pin the version you tested, and review changes before updating.
  • Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.

FAQ

What is Data Pipeline Engineer?

Data Pipeline Engineer is a subagent for Claude Code and Claude Cowork from the yonatangross/orchestkit repository on GitHub. Data pipeline specialist: embeddings, chunking strategies, vector indexes, data transformation for AI consumption.

How do I install Data Pipeline Engineer in Claude Code?

Download data-pipeline-engineer.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.

Can I use Data Pipeline Engineer in Claude Cowork?

Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.

Is Data Pipeline Engineer safe to install?

It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.

Similar resources

Browse all skills, subagents, and plugins →

Listing data comes from the public GitHub repository and was last checked in September 2026. Excerpts are © their authors and shared under MIT. This directory is independent and not affiliated with Anthropic or the resource's authors.