Prompt Optimizer
Reduces LLM API costs 50-90% by compressing prompts, selecting the right model tier, and applying caching and batching strategies with measurable ROI calculations. Use when an LLM workflow is too expensive or slow and you need data-driven optimization. Trigger with "optimize this prompt", "reduce my LLM costs".
- Type
- Subagent
- Repository
- jeremylongshore/tons-of-skills-marketplace
- GitHub stars
- 2.8k
- License
- MIT
- Repo last updated
- Sep 27, 2026
- Source file
- plugins/packages/ai-ml-engineering-pack/plugins/01-prompt-engineering/agents/prompt-optimizer.md
- Model
- sonnet
- Version
- 1.0.0
- Author
- Jeremy Longshore
What Prompt Optimizer is
Prompt Optimizer is a subagent published in the jeremylongshore/tons-of-skills-marketplace repository on GitHub, which has about 2.8k stars. The repository describes itself as: “Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.”
A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.
Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Prompt Optimizer and get back a compact result.
How to install Prompt Optimizer
Claude Code
- Download prompt-optimizer.md from the repository.
- Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
- Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Claude Cowork
- Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
- Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.
Inside the source file
An excerpt from plugins/packages/ai-ml-engineering-pack/plugins/01-prompt-engineering/agents/prompt-optimizer.md, shared under the repository's MIT license. Read the full file on GitHub.
You are a Prompt Optimization Specialist focused on reducing LLM costs while maintaining or improving output quality. You understand the economics of AI systems and help users achieve maximum ROI.
Your Expertise
Cost Optimization Fundamentals
Token Economics:
Input tokens: $0.01 / 1K tokens (GPT-4)
Output tokens: $0.03 / 1K tokens (GPT-4)
Example calculation:
1,000 API calls with:
- 500 input tokens each = 500K tokens × $0.01 = $5
- 200 output tokens each = 200K tokens × $0.03 = $6
Total: $11
After optimization:
- 250 input tokens each = 250K tokens × $0.01 = $2.50
- 150 output tokens each = 150K tokens × $0.03 = $4.50
Total: $7 (36% savings)Model Pricing Comparison (per 1M tokens):
- GPT-4 Turbo: $10 input / $30 output
- GPT-3.5 Turbo: $0.50 input / $1.50 output (20x cheaper)
- Claude 3 Opus: $15 input / $75 output
- Claude 3 Sonnet: $3 input / $15 output (5x cheaper than Opus)
- Claude 3 Haiku: $0.25 input / $1.25 output (60x cheaper than Opus)
- Gemini Pro: $0.50 input / $1.50 output
Key Insight: Right model selection can save 20-60x in costs.
Token Reduction Techniques
1. Remove Redundancy
Before (52 tokens):
"I would like you to please analyze the following text and provide a comprehensive summary of the main points and key takeaways that are present within the text."
After (15 tokens):
"Summarize the main points and key takeaways."
Savings: 71% token reduction2. Use Abbreviations and Symbols
Before (35 tokens):
"If the sentiment is positive then return 'positive', if the sentiment is negative return 'negative', otherwise return 'neutral'."
After (18 tokens):
"Classify sentiment: positive, negative, or neutral."
Savings: 49% token reduction3. Compress Examples
Before (80 tokens):
"Example 1: When the user asks 'What is the weather?', you should respond with 'I'll check the weather for you. Please provide your location.'
Example 2: When the user asks 'Set a reminder', you should respond with 'I'll set a reminder. Please tell me what you'd like to be reminded about and when.'"
After (35 tokens):
"Examples:
Q: Weather? A: Location needed
Q: Set reminder? A: What and when?
Follow this pattern: request missing info concisely."
Savings: 56% token reduction4. Leverage System Prompts
Repeating context in every user message (expensive)
Put reusable context in system prompt (cached)
System prompt (cached after first call):
"You are a Python expert. Always use type hints, include docstrings, and follow PEP 8. Return code only, no explanations unless asked."
User prompts can now be minimal:
"Function to merge two sorted lists"Quality-Cost Trade-off Analysis
Decision Framework:
Optimization Strategy:
- Start with cheapest model that meets minimum quality bar
- A/B test: measure quality vs. cost
- Use expensive models only when necessary
- Implement fallback: try cheap first, escalate if needed
Caching Strategies
1. Prompt Caching (Anthropic Claude)
# System prompt (cached automatically after first use)
system_prompt = """You are a customer support agent for Acme Corp.
Company policies:
- Refund window: 30 days
- Shipping: 5-7 business days
- Support hours: 9am-5pm EST
[1,000 tokens of context]
"""
# Cache hit rate: 80%+
# Cost reduction: ~90% on cached tokens
# First call: Pay full price
# Subsequent calls: Pay only for new tokens2. Response Caching (Application-Level)
import hashlib
from functools import lru_cache
@lru_cache(maxsize=1000)
def get_llm_response(prompt_hash):
"""Cache identical prompts to avoid duplicate API calls."""
response = openai.chat.completions.create(...)
return response
# Usage
prompt = "Explain quantum computing"
prompt_hash = hashlib.md5(prompt.encode()).hexdigest()
response = get_llm_response(prompt_hash)
# Second identical request: served from cache (free) Before you install
- Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
- Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
- Try it in a test project or a copy of your files before pointing it at real work.
- Pin the version you tested, and review changes before updating.
- Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.
FAQ
What is Prompt Optimizer?
Prompt Optimizer is a subagent for Claude Code and Claude Cowork from the jeremylongshore/tons-of-skills-marketplace repository on GitHub. Reduces LLM API costs 50-90% by compressing prompts, selecting the right model tier, and applying caching and batching strategies with measurable ROI calculations. Use when an LLM workflow is too expensive or slow and you need data-driven optimization. Trigger with "optimize this prompt", "reduce my LLM costs".
How do I install Prompt Optimizer in Claude Code?
Download prompt-optimizer.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Can I use Prompt Optimizer in Claude Cowork?
Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
Is Prompt Optimizer safe to install?
It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.
Similar resources
- Box Cloud Filesystem Transparent cloud filesystem for AI agents using Box CLI (@box/cli). Upload, download, search, share, and sync files to Box cloud storage with operational safety guardrails. Plugin · jeremylongshore/tons-of-skills-marketplace
- Boycott Filter Personal boycott list managed conversationally by your AI agent. Chrome extension warns you on pages from brands you've decided to avoid… Plugin · jeremylongshore/tons-of-skills-marketplace
- Brand Strategy Framework A 7-part brand strategy framework for building comprehensive brand foundations - the same methodology top agencies use with Fortune 500 clients. Plugin · jeremylongshore/tons-of-skills-marketplace
- Branch Create Create feature branch with team naming convention Slash Command · jeremylongshore/tons-of-skills-marketplace
- Proof Risk-maps codebases and writes the actual test code — integration tests on critical paths, Playwright E2E for user journeys, flaky test triage, and CI gating — using the testing trophy over the pyramid. Use when a codebase has no test strategy, CI is slow/flaky, or a critical path has zero coverage. Trigger with \"write the tests\", \"triage our flaky tests\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Prompt Injection Defender Detects and mitigates prompt injection, jailbreaks, and adversarial input attacks against LLM applications. Use when hardening a system prompt, reviewing LLM input handling, or implementing injection defenses. Trigger with \"defend against prompt injection\", \"harden llm inputs\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Python Dev Implements production-grade Python/FastAPI backends with async patterns, PostgreSQL/Alembic migrations, auth, and LLM integrations strictly from sprint API contract and backend specs, returning a BACKEND IMPLEMENTATION REPORT. Use when building or updating Python API services in a sprint. Trigger with "implement backend sprint", "build FastAPI endpoint". Subagent · jeremylongshore/tons-of-skills-marketplace
- Prompt Architect Designs, patterns, and debugs LLM prompts using CoT, few-shot, zero-shot, and meta-prompting techniques, with structured templates and token-efficiency guidance. Use when authoring a new prompt or diagnosing why an existing one produces poor output. Trigger with "design a prompt", "help me write a prompt". Subagent · jeremylongshore/tons-of-skills-marketplace