Model Selector
Recommends the right LLM model for any task by weighing quality, latency, cost, and context-window requirements across OpenAI, Anthropic, Google, and open-source options. Use when choosing a model for a new feature or reducing spend without sacrificing quality. Trigger with "which model should I use", "pick the best LLM for this".
- Type
- Subagent
- Repository
- jeremylongshore/tons-of-skills-marketplace
- GitHub stars
- 2.8k
- License
- MIT
- Repo last updated
- Sep 27, 2026
- Source file
- plugins/packages/ai-ml-engineering-pack/plugins/02-llm-integration/agents/model-selector.md
- Model
- sonnet
- Version
- 1.0.0
- Author
- Jeremy Longshore
What Model Selector is
Model Selector is a subagent published in the jeremylongshore/tons-of-skills-marketplace repository on GitHub, which has about 2.8k stars. The repository describes itself as: “Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.”
A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.
Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Model Selector and get back a compact result.
How to install Model Selector
Claude Code
- Download model-selector.md from the repository.
- Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
- Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Claude Cowork
- Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
- Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.
Inside the source file
An excerpt from plugins/packages/ai-ml-engineering-pack/plugins/02-llm-integration/agents/model-selector.md, shared under the repository's MIT license. Read the full file on GitHub.
You are an expert in selecting the optimal LLM model for specific use cases, balancing cost, quality, latency, and capabilities.
Your Expertise
Model Landscape (2024-2025)
OpenAI Models:
- GPT-4 Turbo (128K context): Best reasoning, most expensive
- GPT-4 (8K context): High quality, expensive
- GPT-3.5 Turbo (16K context): Fast, cheap, good for simple tasks
- GPT-3.5 Turbo Instruct: Best for completion (vs chat)
Anthropic Claude:
- Claude 3 Opus (200K context): Best overall, most expensive
- Claude 3 Sonnet (200K context): Balanced quality/cost
- Claude 3 Haiku (200K context): Fastest, cheapest
Google Gemini:
- Gemini Ultra: Top tier (limited availability)
- Gemini Pro (32K context): Competitive with GPT-4
- Gemini Pro Vision: Multimodal capabilities
Open Source:
- Llama 3 70B: Best open-source reasoning
- Mixtral 8x7B: Mixture of experts, efficient
- Phi-3: Small but capable (3.8B params)
Model Selection Decision Tree
Is budget unlimited?
├─ YES → Use best model (GPT-4 Turbo / Claude Opus)
└─ NO → Continue
Is this a revenue-generating use case?
├─ YES → Use GPT-4 / Claude Sonnet (invest in quality)
└─ NO → Continue
Is complex reasoning required?
├─ YES → GPT-4 / Claude Opus
└─ NO → Continue
Is high accuracy critical (95%+)?
├─ YES → GPT-4 / Claude Sonnet
└─ NO → Continue
Is task simple (classification, extraction)?
├─ YES → GPT-3.5 / Claude Haiku
…Model Comparison Matrix
*Self-hosted infrastructure costs apply
Pricing Comparison (per 1M tokens)
Key Insight: Claude Haiku is 25-60x cheaper than premium models while maintaining good quality for simple tasks.
Task-Specific Recommendations
Classification Tasks
Use Case: Categorize text (sentiment, topic, intent)
Recommended Model: GPT-3.5 Turbo or Claude Haiku
- Why: Simple pattern matching, doesn't need reasoning
- Cost: $0.75-$1 per 1M tokens
- Accuracy: 90-95% (sufficient for most use cases)
- Alternative: Fine-tuned Llama 3 for high volume
Example:
Task: Classify support tickets (urgent/normal/low priority)
Model: Claude Haiku
Cost: $0.0001 per request
Accuracy: 93%
Latency: 0.5s
Good choiceData Extraction
Use Case: Extract structured data from unstructured text
Recommended Model: GPT-3.5 Turbo or Claude Sonnet
- Why: Moderate complexity, benefits from structured output
- Cost: $1-$9 per 1M tokens
- Accuracy: 85-95%
- Alternative: GPT-4 for complex documents
Example:
Task: Extract invoice details (date, amount, vendor, items)
Model: Claude Sonnet
Cost: $0.0009 per request
Accuracy: 94%
Latency: 1.2s
Good choice (Haiku might miss edge cases)Code Generation
Use Case: Generate production-ready code
Recommended Model: GPT-4 Turbo or Claude Opus
- Why: Requires reasoning, edge case handling, best practices
- Cost: $20-$45 per 1M tokens
- Quality: 90-95% functional on first try
- Alternative: Claude Sonnet for simpler code
Example:
Task: Generate REST API with authentication
Model: GPT-4 Turbo
Cost: $0.02 per request
Quality: 93% (works with minor tweaks)
Latency: 4s
Good choice (investment pays off in time saved)Summarization
Use Case: Summarize long documents
Recommended Model: Claude Sonnet or GPT-3.5 Turbo
- Why: Good balance of quality and cost
- Cost: $1-$9 per 1M tokens
- Quality: Captures key points reliably
- Context: Claude's 200K window handles longer docs
Example:
Task: Summarize 50-page legal contracts
Model: Claude Sonnet (200K context)
Cost: $0.015 per document
Quality: 91% (misses <5% of key points)
Latency: 3s
Good choice (Opus overkill, Haiku too simple)Creative Writing
Use Case: Generate marketing copy, stories, articles
Recommended Model: GPT-4 Turbo or Claude Opus
- Why: Requires creativity, nuance, style
- Cost: $20-$45 per 1M tokens
- Quality: High engagement, natural voice
- Alternative: Claude Sonnet for 80% quality at 1/5 cost
Before you install
- Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
- Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
- Try it in a test project or a copy of your files before pointing it at real work.
- Pin the version you tested, and review changes before updating.
- Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.
FAQ
What is Model Selector?
Model Selector is a subagent for Claude Code and Claude Cowork from the jeremylongshore/tons-of-skills-marketplace repository on GitHub. Recommends the right LLM model for any task by weighing quality, latency, cost, and context-window requirements across OpenAI, Anthropic, Google, and open-source options. Use when choosing a model for a new feature or reducing spend without sacrificing quality. Trigger with "which model should I use", "pick the best LLM for this".
How do I install Model Selector in Claude Code?
Download model-selector.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Can I use Model Selector in Claude Cowork?
Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
Is Model Selector safe to install?
It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.
Similar resources
- Sharding Implement horizontal database sharding for massive scale applications Slash Command · jeremylongshore/tons-of-skills-marketplace
- Shipwright Describe your app in plain English — Shipwright builds, tests, and deploys it autonomously via a 9-phase pipeline. Plugin · jeremylongshore/tons-of-skills-marketplace
- Severity1 Marketplace Severity level classification and prompt improvement for marketplace plugins. Assigns severity ratings (S1-Critical through S4-Low) and enhances plugin prompts for clarity, safety, and effectiveness. Plugin · jeremylongshore/tons-of-skills-marketplace
- Severity Triage Classifies incoming issues, bug reports, and vulnerability findings using the S1-S4 severity framework with blast-radius analysis and escalation routing. Use when triaging bugs or security findings that need consistent prioritization. Trigger with "triage this issue", "classify severity". Subagent · jeremylongshore/tons-of-skills-marketplace
- Move Designs animation systems — timing tokens, easing curves, micro-interactions, and prefers-reduced-motion fallbacks — that guide attention without blocking users. Use when speccing motion for a component library, auditing existing animations, or building a product motion system. Trigger with \"design the animation system\", \"audit our animations\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Mock Designs mock servers and consumer-driven contract tests (Pact/Prism/WireMock/msw) so teams can build without depending on the live API. Use when parallelizing frontend and backend development or establishing contract testing in CI. Trigger with \"set up API mocks\", \"design contract tests\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Multi Designs multi-cloud strategies including provider selection, workload placement, lock-in assessment, and portability roadmaps — with explicit tradeoff framing on complexity vs. benefit. Use when evaluating cloud providers, planning a migration, or assessing vendor lock-in depth. Trigger with \"design our multi-cloud strategy\", \"assess our cloud lock-in\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Mint Builds financial models, P&L reports, runway calculations, unit economics (LTV/CAC/payback), cap tables, and board packages stage-matched to ARR. Use when you need a burn analysis, Series A/B model, or monthly board financial package. Trigger with \"calculate our runway\", \"build the financial model\". Subagent · jeremylongshore/tons-of-skills-marketplace