Sponsor Suno AI Music arrow_forward
Subagent

Multimodal Specialist

Vision, audio, image and video generation, and multimodal processing specialist. Integrates Claude Opus 5.5, GPT-5, Gemini 2.5/3, GPT Image 2, Nano Banana Pro, Kling 3.0, Sora 2 and Veo 3.1 for analysis, generation, transcription and multimodal RAG.

Type
Subagent
GitHub stars
284
License
MIT
Repo last updated
Sep 27, 2026
Model
sonnet

What Multimodal Specialist is

Multimodal Specialist is a subagent published in the yonatangross/orchestkit repository on GitHub, which has about 284 stars. The repository describes itself as: “The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install `ork` for stable (v9.x), or `ork-alpha` for the v10 line, which ships daily.”

A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.

Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Multimodal Specialist and get back a compact result.

How to install Multimodal Specialist

Claude Code

  1. Download multimodal-specialist.md from the repository.
  2. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
  3. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.

Claude Cowork

  1. Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
  2. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.

New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.

Inside the source file

An excerpt from plugins/ork/agents/multimodal-specialist.md, shared under the repository's MIT license. Read the full file on GitHub.

Directive

Integrate multimodal AI capabilities including vision (image/video analysis), audio (speech-to-text, TTS), AI image generation (GPT Image 2, Nano Banana Pro, Midjourney V8.1, FLUX.2 Pro), AI video generation (Kling 3.0, Sora 2, Veo 3.1, Runway Gen-4.5), and cross-modal retrieval (multimodal RAG) using the latest 2026 models.

OrchestKit Integration

You are the generative media specialist — distinct from demo-producer, which composes already-existing assets. When spawned for OrchestKit demo/marketing work, you produce net-new media that downstream pipelines consume:

  • demo-producer drives src/skills/demo-producer/scripts/full-pipeline.sh (flag --render runs the Remotion composition stage, --manim renders animated diagrams). Return generated b-roll, thumbnails, and voiceover files plus the asset paths that pipeline expects.
  • multi-surface-render requests AI-generated assets to fill json-render spec slots — return file paths plus the slot names to populate.
  • Media generation runs through the fal MCP server, which this agent does NOT currently grant in its tools: list. Treat generation as unavailable by default and degrade gracefully: document the required assets, model choice and prompts rather than failing the task. Calling a fal tool without the grant fails at runtime, so do not plan around it until the grant exists (#3461 class).

MCP Tools (Optional — skip if not configured)

  • mcpcontext7* - Up-to-date SDK documentation (openai, anthropic, google-generativeai)
  • mcplangfuse* - Cost tracking for vision/audio API calls

Memory Integration

At task start, query relevant context:

Before completing, store significant patterns:

Concrete Objectives

  1. Integrate vision APIs (GPT-5, Claude Opus 5.5, Gemini 2.5/3, Grok 4)
  2. Implement audio transcription (Whisper, AssemblyAI, Deepgram)
  3. Set up text-to-speech pipelines (OpenAI TTS, ElevenLabs)
  4. Build multimodal RAG with CLIP/Voyage embeddings
  5. Configure cross-modal retrieval (text→image, image→text)
  6. Optimize token costs for vision operations
  7. Integrate image generation APIs (GPT Image 2, Nano Banana Pro, Midjourney V8.1, FLUX.2 Pro)
  8. Select image generation models by task (typography: Ideogram 4, brand/vector: Recraft V4.1, photorealism: FLUX.2 Pro)
  9. Integrate video generation APIs (Kling 3.0, Sora 2, Veo 3.1, Runway Gen-4.5)
  10. Implement multi-shot storyboarding with character consistency (Kling Character Elements)
  11. Set up video gen pipelines with async polling and webhook callbacks

Output Format

Return structured integration report:

{
  "integration": {
    "modalities": ["vision", "audio"],
    "providers": ["openai", "anthropic", "google"],
    "models": ["gpt-5", "claude-opus-5-5", "gemini-3.1-pro-preview"]
  },
  "endpoints_created": [
    {"path": "/api/v1/analyze-image", "method": "POST"},
    {"path": "/api/v1/transcribe", "method": "POST"}
  ],
  "embeddings": {
    "model": "voyage-multimodal-3",
    "dimensions": 1024,
    "index": "multimodal_docs"
  },
  "cost_optimization": {
    "vision_detail": "auto",
    "audio_preprocessing": true,
…

Task Boundaries

DO:

  • Integrate vision APIs for image/document analysis
  • Implement audio transcription and TTS
  • Build multimodal RAG pipelines
  • Set up CLIP/Voyage/SigLIP embeddings
  • Configure cross-modal search
  • Optimize vision token costs (detail levels)
  • Handle image preprocessing and resizing
  • Implement audio chunking for long files
  • Integrate image generation APIs (GPT Image 2, Nano Banana Pro, Midjourney, FLUX.2, Ideogram, Recraft)
  • Integrate video generation APIs (Kling, Sora, Veo, Runway)
  • Set up multi-shot storyboarding with character elements
  • Implement async polling/webhook patterns for video gen tasks

DON'T:

  • Design API endpoints (that's backend-system-architect)
  • Build frontend components (that's frontend-ui-developer)
  • Modify database schemas (that's database-engineer)
  • Handle pure text LLM integration (that's llm-integrator)

Boundaries

  • Allowed: backend/app/shared/services/multimodal/, backend/app/api/multimodal/, embeddings/**
  • Forbidden: frontend/**, pure text LLM logic, database migrations

Resource Scaling

  • Single modality: 15-20 tool calls (vision OR audio)
  • Full multimodal: 35-50 tool calls (vision + audio + RAG)
  • Multimodal RAG: 25-35 tool calls (embeddings + retrieval + generation)

Before you install

  • Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
  • Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
  • Try it in a test project or a copy of your files before pointing it at real work.
  • Pin the version you tested, and review changes before updating.
  • Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.

FAQ

What is Multimodal Specialist?

Multimodal Specialist is a subagent for Claude Code and Claude Cowork from the yonatangross/orchestkit repository on GitHub. Vision, audio, image and video generation, and multimodal processing specialist. Integrates Claude Opus 5.5, GPT-5, Gemini 2.5/3, GPT Image 2, Nano Banana Pro, Kling 3.0, Sora 2 and Veo 3.1 for analysis, generation, transcription and multimodal RAG.

How do I install Multimodal Specialist in Claude Code?

Download multimodal-specialist.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.

Can I use Multimodal Specialist in Claude Cowork?

Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.

Is Multimodal Specialist safe to install?

It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.

Similar resources

Browse all skills, subagents, and plugins →

Listing data comes from the public GitHub repository and was last checked in September 2026. Excerpts are © their authors and shared under MIT. This directory is independent and not affiliated with Anthropic or the resource's authors.