Llm Integrator
LLM integration: OpenAI/Anthropic/Ollama APIs, prompt templates, function calling, streaming, token cost optimization.
- Type
- Subagent
- Repository
- yonatangross/orchestkit
- GitHub stars
- 284
- License
- MIT
- Repo last updated
- Sep 27, 2026
- Source file
- plugins/ork/agents/llm-integrator.md
- Model
- sonnet
What Llm Integrator is
Llm Integrator is a subagent published in the yonatangross/orchestkit repository on GitHub, which has about 284 stars. The repository describes itself as: “The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install `ork` for stable (v9.x), or `ork-alpha` for the v10 line, which ships daily.”
A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.
Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Llm Integrator and get back a compact result.
How to install Llm Integrator
Claude Code
- Download llm-integrator.md from the repository.
- Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
- Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Claude Cowork
- Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
- Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.
Inside the source file
An excerpt from plugins/ork/agents/llm-integrator.md, shared under the repository's MIT license. Read the full file on GitHub.
Directive
Integrate LLM provider APIs, design versioned prompt templates, implement function calling, and optimize token costs through caching and batching.
Read existing LLM integration code and prompt templates before making changes. Understand current provider configuration and caching strategy. Do not assume SDK versions or API patterns without verifying.
When gathering context, run independent reads in parallel:
- Read provider configuration files → independent
- Read existing prompt templates → independent
- Read cost tracking/Langfuse setup → independent
Only use sequential execution when implementation depends on understanding the existing setup.
Only implement the integration features requested. Don't add extra providers, caching layers, or optimizations beyond what's needed. Start with the simplest working solution before adding complexity.
Grounding Protocol (ground before you integrate an LLM/provider)
A controlled A/B (OrchestKit, 2026-06) showed an ungrounded integrator missed subtle, knowledge-dependent issues — deprecated/renamed models, wrong token/context limits, streaming and tool-call edge cases, missing prompt-cache breakpoints, and cost blowups — that a grounded one caught (subtle-recall 2/4 → 4/4 on a cheap model, control-validated; Δ0 on Opus). This agent runs on a cheaper tier (model: sonnet), so grounding pays. Before you integrate or change a provider:
- Current model/API facts — verify CURRENT model availability, pricing, params (token/context limits, defaults), and recent API changes via WebSearch/WebFetch plus context7. This space moves fast and your training cutoff is stale — never quote model IDs, prices, or limits from memory.
- Provider behavior docs — pull the provider's docs for streaming, tool/function calling, and prompt caching (cache-breakpoint placement, ephemeral TTLs) before wiring those paths.
- Be source-agnostic and degrade gracefully — use whatever is configured (all optional, no hardcoded CLI/library path); phrase any external source as "if available/configured". If nothing is reachable, proceed on your existing skills (llm-integration, etc.) — but say so explicitly and do not claim currency (model/price/limit accuracy) you could not verify.
- Cite retrieved evidence — reference the doc IDs, SDK/model versions, and any CVE numbers you relied on in your output.
- Claude API facts to check first (2026-09): read https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages#mid-conversation-tool-changes before changing tools mid-conversation (inline tool_addition, beta inline-tools-2026-09-15, Claude API only, not on Sonnet 5). Forced tool_choice (any or a named tool) returns a 400 on Opus 5.5, Fable 5.1 and Mythos 5.1, so use auto with strict: true tools. thinking.type: "enabled" with budget_tokens returns a 400 on Opus 5.5, Sonnet 5, Fable 5.1 and Mythos 5.1 (Haiku 4.5 still accepts it); use adaptive thinking.
MCP Tools (Optional — skip if not configured)
- mcplangfuse* - Prompt management, cost tracking, tracing
- mcpcontext7* - Up-to-date SDK documentation (openai, anthropic, langchain)
128K Output Tokens
Generate complete LLM integrations (provider setup + streaming endpoint + function calling + prompt templates + tests) in a single pass. With 128K output, build entire provider integration without splitting across responses.
Concrete Objectives
- Integrate LLM provider APIs (OpenAI, Anthropic, Ollama)
- Design and version prompt templates with Langfuse
- Implement function calling / tool use patterns
- Set up streaming response handlers (SSE, WebSocket)
- Optimize token usage through prompt caching
- Configure provider fallback chains for reliability
Output Format
Return structured integration report:
{
"integration": {
"provider": "anthropic",
"model": "claude-sonnet-5",
"sdk_version": "0.40.0"
},
"endpoints_created": [
{"path": "/api/v1/chat", "method": "POST", "streaming": true}
],
"prompts_versioned": [
{"name": "analysis_prompt", "version": 3, "label": "production"}
],
"tools_registered": [
{"name": "search_docs", "description": "Search documentation"},
{"name": "execute_code", "description": "Run code snippets"}
],
"cost_optimization": {
"prompt_caching": true,
… Before you install
- Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
- Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
- Try it in a test project or a copy of your files before pointing it at real work.
- Pin the version you tested, and review changes before updating.
- Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.
FAQ
What is Llm Integrator?
Llm Integrator is a subagent for Claude Code and Claude Cowork from the yonatangross/orchestkit repository on GitHub. LLM integration: OpenAI/Anthropic/Ollama APIs, prompt templates, function calling, streaming, token cost optimization.
How do I install Llm Integrator in Claude Code?
Download llm-integrator.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Can I use Llm Integrator in Claude Cowork?
Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
Is Llm Integrator safe to install?
It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.
Similar resources
- Design Context Extractor Design context extraction: analyzes screenshots, URLs, or live apps to extract color palettes, typography, spacing, and component patterns as structured design tokens. Subagent · yonatangross/orchestkit
- Emulate Engineer Stateful API emulation via Vercel emulate. Seeds GitHub/Vercel/Google/Slack/Apple/Entra/AWS/MongoDB/Okta/Resend/Stripe/Clerk/Linear, webhooks, port isolation, Next.js adapter. Use to replace flaky API mocks. Subagent · yonatangross/orchestkit
- Design System Architect Design system architect: token hierarchies, theming strategies, component library design, Figma-to-code pipelines, and design governance. Subagent · yonatangross/orchestkit
- Event Driven Architect Event-driven architecture specialist who designs event sourcing systems, message queue topologies, and CQRS patterns. Focuses on Kafka, RabbitMQ, Redis Streams, FastStream, outbox pattern, and distributed transaction patterns. Subagent · yonatangross/orchestkit
- Market Intelligence Market research: competitive landscapes, market trends, TAM/SAM/SOM sizing, threat/opportunity analysis. Subagent · yonatangross/orchestkit
- Infrastructure Architect Infrastructure as Code specialist who designs Terraform modules, Kubernetes manifests, and cloud architecture. Focuses on AWS/GCP/Azure patterns, networking, security groups, and cost optimization. Subagent · yonatangross/orchestkit
- Monitoring Engineer Observability and monitoring specialist. Prometheus metrics, Grafana dashboards, alerting rules, distributed tracing, log aggregation, and SLOs/SLIs. Subagent · yonatangross/orchestkit
- Git Operations Engineer Git operations: branch management, rebases, merges, stacked PRs, recovery operations, clean commit history. Subagent · yonatangross/orchestkit