Llm Api Scaffold
Generate production-ready LLM API integration boilerplate
- Type
- Slash Command
- Repository
- jeremylongshore/tons-of-skills-marketplace
- GitHub stars
- 2.8k
- License
- MIT
- Repo last updated
- Sep 27, 2026
- Source file
- plugins/packages/ai-ml-engineering-pack/plugins/02-llm-integration/commands/llm-api-scaffold.md
- Version
- 1.0.0
- Author
- Jeremy Longshore
What Llm Api Scaffold is
Llm Api Scaffold is a slash command published in the jeremylongshore/tons-of-skills-marketplace repository on GitHub, which has about 2.8k stars. The repository describes itself as: “Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.”
A slash command is a reusable prompt saved as a markdown file and run by typing its name after a slash. In Claude Code, custom commands have been merged into skills: a file in .claude/commands/ and a skill folder in .claude/skills/ both create the same kind of command, and existing command files keep working.
Llm Api Scaffold gives you a repeatable way to run the same instructions without retyping them, optionally with arguments.
How to install Llm Api Scaffold
Claude Code
- Download llm-api-scaffold.md from the repository.
- Save it to ~/.claude/commands/ (all projects) or .claude/commands/ (one project). As a skill, you can instead save it as ~/.claude/skills/<name>/SKILL.md.
- Run it by typing / followed by its name.
Claude Cowork
- Turn the command into a skill: create a folder with the file saved as SKILL.md and zip it.
- In Customize → Skills, click +, then upload the ZIP.
- Run it from any task with / and the skill name.
New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.
Inside the source file
An excerpt from plugins/packages/ai-ml-engineering-pack/plugins/02-llm-integration/commands/llm-api-scaffold.md, shared under the repository's MIT license. Read the full file on GitHub.
Generate complete, production-ready LLM API integration code with error handling, rate limiting, caching, monitoring, and best practices built-in.
What You'll Get
When you run this command, you'll receive:
- Complete API client with retry logic and error handling
- Rate limiting (token bucket algorithm)
- Caching layer (in-memory + Redis)
- Cost tracking and monitoring
- Multi-provider support with fallback
- Streaming responses implementation
- Testing suite with mocks
- Docker deployment configuration
Usage
/llm-api-scaffold <provider> [options]Providers: anthropic, openai, multi (both)
Examples:
- /las anthropic - Generate Anthropic Claude integration
- /las openai - Generate OpenAI GPT integration
- /las multi - Generate multi-provider with fallback
Generated Output
Example: Anthropic Claude Integration
Input:
/las anthropicOutput:
1. Project Structure
llm-api-integration/
├── src/
│ ├── client/
│ │ ├── __init__.py
│ │ ├── base.py # Base client interface
│ │ ├── anthropic_client.py # Anthropic implementation
│ │ └── rate_limiter.py # Rate limiting
│ ├── cache/
│ │ ├── __init__.py
│ │ ├── memory_cache.py # In-memory cache
│ │ └── redis_cache.py # Redis cache
│ ├── monitoring/
│ │ ├── __init__.py
│ │ ├── metrics.py # Prometheus metrics
│ │ └── cost_tracker.py # Cost tracking
│ └── utils/
│ ├── __init__.py
│ └── retry.py # Retry logic
…2. Base Client (src/client/base.py)
from abc import ABC, abstractmethod
from dataclasses import dataclass
from typing import Optional, AsyncGenerator
@dataclass
class CompletionRequest:
"""Standardized completion request."""
prompt: str
max_tokens: int = 1024
temperature: float = 1.0
model: Optional[str] = None
stream: bool = False
@dataclass
class CompletionResponse:
"""Standardized completion response."""
content: str
usage: dict
…3. Anthropic Client (src/client/anthropic_client.py)
import time
import asyncio
from anthropic import AsyncAnthropic, RateLimitError, APIError
from .base import BaseLLMClient, CompletionRequest, CompletionResponse
from .rate_limiter import TokenBucket
from ..cache import CacheManager
from ..monitoring import MetricsCollector, CostTracker
from ..utils.retry import retry_with_backoff
class AnthropicClient(BaseLLMClient):
"""Production-ready Anthropic Claude client."""
def __init__(
self,
api_key: str,
model: str = "claude-3-haiku-20240307",
requests_per_minute: int = 50,
enable_cache: bool = True,
…4. Rate Limiter (src/client/rate_limiter.py)
import time
import asyncio
from threading import Lock
class TokenBucket:
"""Thread-safe token bucket for rate limiting."""
def __init__(self, capacity: int, refill_rate: float):
"""
Args:
capacity: Maximum tokens (e.g., 50 requests)
refill_rate: Tokens per second (e.g., 50/60 = 0.833 req/s)
"""
self.capacity = capacity
self.tokens = capacity
self.refill_rate = refill_rate
self.last_refill = time.time()
self.lock = Lock()
…5. Cache Manager (src/cache/memory_cache.py)
import time
from typing import Optional, Any
from collections import OrderedDict
class MemoryCache:
"""In-memory LRU cache with TTL."""
def __init__(self, max_size: int = 1000):
self.max_size = max_size
self.cache = OrderedDict()
self.expiry = {}
async def get(self, key: str) -> Optional[Any]:
"""Get cached value if not expired."""
if key not in self.cache:
return None
# Check expiry
…6. Cost Tracker (src/monitoring/cost_tracker.py)
import time
from collections import defaultdict
from dataclasses import dataclass, field
@dataclass
class CostTracker:
"""Track LLM usage costs."""
usage_history: list = field(default_factory=list)
costs_by_model: dict = field(default_factory=lambda: defaultdict(float))
PRICING = {
"claude-3-opus-20240229": {"input": 0.015, "output": 0.075},
"claude-3-sonnet-20240229": {"input": 0.003, "output": 0.015},
"claude-3-haiku-20240307": {"input": 0.00025, "output": 0.00125},
"gpt-4-turbo-preview": {"input": 0.01, "output": 0.03},
"gpt-3.5-turbo": {"input": 0.0005, "output": 0.0015}
}
…7. Retry Utility (src/utils/retry.py)
import time
import random
import asyncio
from functools import wraps
def retry_with_backoff(
max_retries: int = 3,
base_delay: float = 1.0,
max_delay: float = 60.0,
exponential_base: float = 2.0,
jitter: bool = True
):
"""Async retry decorator with exponential backoff."""
def decorator(func):
@wraps(func)
async def wrapper(*args, **kwargs):
retries = 0
while retries < max_retries:
… Before you install
- Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
- Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
- Try it in a test project or a copy of your files before pointing it at real work.
- Pin the version you tested, and review changes before updating.
- Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.
FAQ
What is Llm Api Scaffold?
Llm Api Scaffold is a slash command for Claude Code and Claude Cowork from the jeremylongshore/tons-of-skills-marketplace repository on GitHub. Generate production-ready LLM API integration boilerplate
How do I install Llm Api Scaffold in Claude Code?
Download llm-api-scaffold.md from the repository. Save it to ~/.claude/commands/ (all projects) or .claude/commands/ (one project). As a skill, you can instead save it as ~/.claude/skills/<name>/SKILL.md. Run it by typing / followed by its name.
Can I use Llm Api Scaffold in Claude Cowork?
Turn the command into a skill: create a folder with the file saved as SKILL.md and zip it. In Customize → Skills, click +, then upload the ZIP. Run it from any task with / and the skill name.
Is Llm Api Scaffold safe to install?
It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.
Similar resources
- Api Documentation Generator Generate comprehensive API documentation from OpenAPI/Swagger specs Plugin · jeremylongshore/tons-of-skills-marketplace
- Api Event Emitter Implement event-driven APIs with message queues and event streaming Plugin · jeremylongshore/tons-of-skills-marketplace
- Api Error Handler Implement standardized error handling with proper HTTP status codes Plugin · jeremylongshore/tons-of-skills-marketplace
- Api Expert Diagnoses REST API failures by comparing HTTP logs against OpenAPI specs, identifies root cause by status code category, and generates working cURL repro commands. Use when an API call returns unexpected errors or status codes. Trigger with "why is my API failing", "debug this API error". Subagent · jeremylongshore/tons-of-skills-marketplace
- Load Balance Configure load balancers (ALB, NLB, Nginx, HAProxy) Slash Command · jeremylongshore/tons-of-skills-marketplace
- Learn Toggle learning mode with educational micro-explanations Slash Command · jeremylongshore/tons-of-skills-marketplace
- Log Log a new AI experiment with terminal report Slash Command · jeremylongshore/tons-of-skills-marketplace
- Lb Test Test load balancer traffic distribution and failover strategies Slash Command · jeremylongshore/tons-of-skills-marketplace