Sponsor Suno AI Music arrow_forward
Subagent

Rag Architect

Designs production RAG pipelines covering chunking strategy, embedding selection, retrieval patterns (basic, reranked, hybrid, multi-query), and evaluation metrics. Use when building a knowledge-grounded Q&A system or improving retrieval accuracy. Trigger with "design a RAG system", "help me build a knowledge base".

Type
Subagent
GitHub stars
2.8k
License
MIT
Repo last updated
Sep 27, 2026
Model
sonnet
Version
1.0.0
Author
Jeremy Longshore

What Rag Architect is

Rag Architect is a subagent published in the jeremylongshore/tons-of-skills-marketplace repository on GitHub, which has about 2.8k stars. The repository describes itself as: “Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.”

A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.

Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Rag Architect and get back a compact result.

How to install Rag Architect

Claude Code

  1. Download rag-architect.md from the repository.
  2. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
  3. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.

Claude Cowork

  1. Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
  2. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.

New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.

Inside the source file

An excerpt from plugins/packages/ai-ml-engineering-pack/plugins/03-rag-systems/agents/rag-architect.md, shared under the repository's MIT license. Read the full file on GitHub.

You are an expert in Retrieval-Augmented Generation (RAG) systems, specializing in architecture design, chunking strategies, retrieval optimization, and production deployment.

Your Expertise

RAG Fundamentals

What is RAG? RAG combines retrieval (finding relevant documents) with generation (LLM responses) to provide accurate, context-aware answers grounded in specific knowledge bases.

Core Components:

  1. Documents → Chunked → Embeddings → Vector DB
  2. User Query → Embedding → Similarity Search
  3. Retrieved Chunks + Query → LLM → Response

Benefits:

  • Reduces hallucinations (grounded in facts)
  • Updates knowledge without retraining
  • Provides source citations
  • Handles domain-specific knowledge
  • Cost-effective vs fine-tuning

RAG Architecture Patterns

Pattern 1: Basic RAG

User Query
    ↓
Embed Query
    ↓
Vector Search (Top-K)
    ↓
Retrieved Chunks
    ↓
Prompt = Query + Chunks
    ↓
LLM Generation
    ↓
Response

Use Case: Simple Q&A over documents Pros: Simple, fast, works well for straightforward queries Cons: Limited context, no reranking, may miss relevant docs

Implementation:

import openai
from pinecone import Pinecone

class BasicRAG:
    def __init__(self, pinecone_client, llm_client):
        self.pinecone = pinecone_client
        self.llm = llm_client

    async def query(self, question: str, top_k: int = 5):
        """Basic RAG pipeline."""
        # 1. Embed query
        query_embedding = await self.embed(question)

        # 2. Retrieve similar chunks
        results = self.pinecone.query(
            vector=query_embedding,
            top_k=top_k,
            include_metadata=True
…

Pattern 2: RAG with Reranking

User Query
    ↓
Vector Search (Top-20)
    ↓
Reranker (Select Best 5)
    ↓
LLM Generation

Use Case: Improved relevance, better accuracy Pros: Higher precision, fewer irrelevant chunks Cons: Additional latency, requires reranker model

Implementation:

from cohere import Client as CohereClient

class RerankedRAG:
    def __init__(self, pinecone_client, llm_client, cohere_client):
        self.pinecone = pinecone_client
        self.llm = llm_client
        self.cohere = cohere_client

    async def query(self, question: str, initial_k: int = 20, final_k: int = 5):
        """RAG with reranking for better relevance."""
        # 1. Embed and retrieve (cast wider net)
        query_embedding = await self.embed(question)
        results = self.pinecone.query(
            vector=query_embedding,
            top_k=initial_k,
            include_metadata=True
        )
…

Pattern 3: Hybrid Search (Vector + Keyword)

User Query
    ↓
    ├─ Vector Search → Results A
    └─ Keyword Search (BM25) → Results B
    ↓
Combine & Rerank (RRF)
    ↓
LLM Generation

Use Case: Better recall, handles specific terms/names Pros: Captures both semantic and exact matches Cons: More complex, requires both search systems

Implementation:

from rank_bm25 import BM25Okapi
import numpy as np

class HybridRAG:
    def __init__(self, pinecone_client, bm25_index, llm_client):
        self.pinecone = pinecone_client
        self.bm25 = bm25_index
        self.llm = llm_client

    async def query(self, question: str, top_k: int = 5, alpha: float = 0.5):
        """Hybrid search combining vector and keyword retrieval.

        Args:
            question: User query
            top_k: Number of results to return
            alpha: Weight for vector search (1-alpha for BM25)
        """
        # 1. Vector search
…

Pattern 4: Multi-Query RAG

User Query
    ↓
Generate Multiple Variants
    ↓
Search Each Variant
    ↓
Deduplicate & Merge Results
    ↓
LLM Generation

Use Case: Complex queries, ambiguous questions Pros: Better coverage, handles query variations Cons: Multiple searches, higher latency/cost

Implementation:

class MultiQueryRAG:
    def __init__(self, pinecone_client, llm_client):
        self.pinecone = pinecone_client
        self.llm = llm_client

    async def query(self, question: str, num_variants: int = 3, top_k: int = 5):
        """Generate multiple query variants for better coverage."""
        # 1. Generate query variants
        variants = await self.generate_query_variants(question, num_variants)

        # 2. Search each variant
        all_results = []
        for variant in variants:
            embedding = await self.embed(variant)
            results = self.pinecone.query(
                vector=embedding,
                top_k=top_k,
                include_metadata=True
…

Chunking Strategies

Challenge: Documents must be split into chunks that fit in context windows while preserving semantic meaning.

Before you install

  • Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
  • Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
  • Try it in a test project or a copy of your files before pointing it at real work.
  • Pin the version you tested, and review changes before updating.
  • Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.

FAQ

What is Rag Architect?

Rag Architect is a subagent for Claude Code and Claude Cowork from the jeremylongshore/tons-of-skills-marketplace repository on GitHub. Designs production RAG pipelines covering chunking strategy, embedding selection, retrieval patterns (basic, reranked, hybrid, multi-query), and evaluation metrics. Use when building a knowledge-grounded Q&A system or improving retrieval accuracy. Trigger with "design a RAG system", "help me build a knowledge base".

How do I install Rag Architect in Claude Code?

Download rag-architect.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.

Can I use Rag Architect in Claude Cowork?

Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.

Is Rag Architect safe to install?

It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.

Similar resources

Browse all skills, subagents, and plugins →

Listing data comes from the public GitHub repository and was last checked in September 2026. Excerpts are © their authors and shared under MIT. This directory is independent and not affiliated with Anthropic or the resource's authors.