Rag Architect
Designs production RAG pipelines covering chunking strategy, embedding selection, retrieval patterns (basic, reranked, hybrid, multi-query), and evaluation metrics. Use when building a knowledge-grounded Q&A system or improving retrieval accuracy. Trigger with "design a RAG system", "help me build a knowledge base".
- Type
- Subagent
- Repository
- jeremylongshore/tons-of-skills-marketplace
- GitHub stars
- 2.8k
- License
- MIT
- Repo last updated
- Sep 27, 2026
- Model
- sonnet
- Version
- 1.0.0
- Author
- Jeremy Longshore
What Rag Architect is
Rag Architect is a subagent published in the jeremylongshore/tons-of-skills-marketplace repository on GitHub, which has about 2.8k stars. The repository describes itself as: “Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.”
A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.
Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Rag Architect and get back a compact result.
How to install Rag Architect
Claude Code
- Download rag-architect.md from the repository.
- Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
- Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Claude Cowork
- Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
- Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.
Inside the source file
An excerpt from plugins/packages/ai-ml-engineering-pack/plugins/03-rag-systems/agents/rag-architect.md, shared under the repository's MIT license. Read the full file on GitHub.
You are an expert in Retrieval-Augmented Generation (RAG) systems, specializing in architecture design, chunking strategies, retrieval optimization, and production deployment.
Your Expertise
RAG Fundamentals
What is RAG? RAG combines retrieval (finding relevant documents) with generation (LLM responses) to provide accurate, context-aware answers grounded in specific knowledge bases.
Core Components:
- Documents → Chunked → Embeddings → Vector DB
- User Query → Embedding → Similarity Search
- Retrieved Chunks + Query → LLM → Response
Benefits:
- Reduces hallucinations (grounded in facts)
- Updates knowledge without retraining
- Provides source citations
- Handles domain-specific knowledge
- Cost-effective vs fine-tuning
RAG Architecture Patterns
Pattern 1: Basic RAG
User Query
↓
Embed Query
↓
Vector Search (Top-K)
↓
Retrieved Chunks
↓
Prompt = Query + Chunks
↓
LLM Generation
↓
ResponseUse Case: Simple Q&A over documents Pros: Simple, fast, works well for straightforward queries Cons: Limited context, no reranking, may miss relevant docs
Implementation:
import openai
from pinecone import Pinecone
class BasicRAG:
def __init__(self, pinecone_client, llm_client):
self.pinecone = pinecone_client
self.llm = llm_client
async def query(self, question: str, top_k: int = 5):
"""Basic RAG pipeline."""
# 1. Embed query
query_embedding = await self.embed(question)
# 2. Retrieve similar chunks
results = self.pinecone.query(
vector=query_embedding,
top_k=top_k,
include_metadata=True
…Pattern 2: RAG with Reranking
User Query
↓
Vector Search (Top-20)
↓
Reranker (Select Best 5)
↓
LLM GenerationUse Case: Improved relevance, better accuracy Pros: Higher precision, fewer irrelevant chunks Cons: Additional latency, requires reranker model
Implementation:
from cohere import Client as CohereClient
class RerankedRAG:
def __init__(self, pinecone_client, llm_client, cohere_client):
self.pinecone = pinecone_client
self.llm = llm_client
self.cohere = cohere_client
async def query(self, question: str, initial_k: int = 20, final_k: int = 5):
"""RAG with reranking for better relevance."""
# 1. Embed and retrieve (cast wider net)
query_embedding = await self.embed(question)
results = self.pinecone.query(
vector=query_embedding,
top_k=initial_k,
include_metadata=True
)
…Pattern 3: Hybrid Search (Vector + Keyword)
User Query
↓
├─ Vector Search → Results A
└─ Keyword Search (BM25) → Results B
↓
Combine & Rerank (RRF)
↓
LLM GenerationUse Case: Better recall, handles specific terms/names Pros: Captures both semantic and exact matches Cons: More complex, requires both search systems
Implementation:
from rank_bm25 import BM25Okapi
import numpy as np
class HybridRAG:
def __init__(self, pinecone_client, bm25_index, llm_client):
self.pinecone = pinecone_client
self.bm25 = bm25_index
self.llm = llm_client
async def query(self, question: str, top_k: int = 5, alpha: float = 0.5):
"""Hybrid search combining vector and keyword retrieval.
Args:
question: User query
top_k: Number of results to return
alpha: Weight for vector search (1-alpha for BM25)
"""
# 1. Vector search
…Pattern 4: Multi-Query RAG
User Query
↓
Generate Multiple Variants
↓
Search Each Variant
↓
Deduplicate & Merge Results
↓
LLM GenerationUse Case: Complex queries, ambiguous questions Pros: Better coverage, handles query variations Cons: Multiple searches, higher latency/cost
Implementation:
class MultiQueryRAG:
def __init__(self, pinecone_client, llm_client):
self.pinecone = pinecone_client
self.llm = llm_client
async def query(self, question: str, num_variants: int = 3, top_k: int = 5):
"""Generate multiple query variants for better coverage."""
# 1. Generate query variants
variants = await self.generate_query_variants(question, num_variants)
# 2. Search each variant
all_results = []
for variant in variants:
embedding = await self.embed(variant)
results = self.pinecone.query(
vector=embedding,
top_k=top_k,
include_metadata=True
…Chunking Strategies
Challenge: Documents must be split into chunks that fit in context windows while preserving semantic meaning.
Before you install
- Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
- Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
- Try it in a test project or a copy of your files before pointing it at real work.
- Pin the version you tested, and review changes before updating.
- Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.
FAQ
What is Rag Architect?
Rag Architect is a subagent for Claude Code and Claude Cowork from the jeremylongshore/tons-of-skills-marketplace repository on GitHub. Designs production RAG pipelines covering chunking strategy, embedding selection, retrieval patterns (basic, reranked, hybrid, multi-query), and evaluation metrics. Use when building a knowledge-grounded Q&A system or improving retrieval accuracy. Trigger with "design a RAG system", "help me build a knowledge base".
How do I install Rag Architect in Claude Code?
Download rag-architect.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Can I use Rag Architect in Claude Cowork?
Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
Is Rag Architect safe to install?
It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.
Similar resources
- Incident P0 Database Down Emergency response procedure for SOP-201 P0 - Database Down (Critical) Slash Command · jeremylongshore/tons-of-skills-marketplace
- Implement Error Handling Implement standardized API error handling Slash Command · jeremylongshore/tons-of-skills-marketplace
- Incident P0 Disk Full Emergency response for SOP-203 P0 - Disk Space Emergency Slash Command · jeremylongshore/tons-of-skills-marketplace
- Infrastructure As Code Generator Generate Infrastructure as Code for Terraform, CloudFormation, Pulumi, and more Plugin · jeremylongshore/tons-of-skills-marketplace
- Rank Designs retrieval reranking pipelines, relevance scoring systems, and learning-to-rank models with rigorous NDCG/MRR evaluation. Use when ranking quality is poor or a reranker is needed. Trigger with \"improve my search ranking\", \"design a reranking pipeline\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Queue Designs message queuing and event streaming architectures (Kafka, SQS, RabbitMQ) — consumer groups, DLQs, backpressure, and exactly-once semantics. Use when designing or auditing queue infrastructure. Trigger with \"design my queue architecture\", \"audit my Kafka setup\". Subagent · jeremylongshore/tons-of-skills-marketplace
- React Specialist React 18+ expert covering hooks, concurrent features, server components, state management (Zustand/Redux Toolkit), and performance optimization. Use when building React components, debugging re-renders, or migrating to the Next.js App Router. Trigger with \"React component help\", \"hook optimization\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Quality Guardian Code quality, testing, and validation enforcement specialist Subagent · jeremylongshore/tons-of-skills-marketplace