Sponsor Suno AI Music arrow_forward
Plugin

Ollama Local Ai

Run AI models locally with Ollama - free alternative to OpenAI, Anthropic, and other paid LLM APIs. Zero-cost, privacy-first AI infrastructure.

Type
Plugin
GitHub stars
2.8k
License
MIT
Repo last updated
Sep 27, 2026
Version
1.18.0
Author
Jeremy Longshore

What Ollama Local Ai is

Ollama Local Ai is a plugin published in the jeremylongshore/tons-of-skills-marketplace repository on GitHub, which has about 2.8k stars. The repository describes itself as: “Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.”

A plugin is a package that bundles skills, slash commands, subagents, hooks, and MCP connectors so they install together. Plugins are plain files with a manifest at .claude-plugin/plugin.json, and they work in both Claude Code and Claude Cowork.

Installing Ollama Local Ai adds everything it ships in one step. Connectors inside a plugin still need to be connected separately, and hooks and subagents only run in Cowork and Claude Code, not in regular chat.

How to install Ollama Local Ai

Claude Code

  1. Add the repository as a plugin marketplace: claude plugin marketplace add jeremylongshore/tons-of-skills-marketplace
  2. Install the plugin: claude plugin install ollama-local-ai@<marketplace-name>, using the marketplace name from the repository's .claude-plugin/marketplace.json.
  3. Restart the session if the new skills or commands don't appear straight away.

Claude Cowork

  1. Open Customize → Plugins and choose Add marketplace.
  2. Enter jeremylongshore/tons-of-skills-marketplace (the owner/repo shorthand works for GitHub).
  3. Find Ollama Local Ai in the list, click Install, then connect any connectors it needs from its Connectors tab.

New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.

Inside the source file

An excerpt from plugins/ai-ml/ollama-local-ai/.claude-plugin/plugin.json, shared under the repository's MIT license. Read the full file on GitHub.

Free, self-hosted alternative to OpenAI, Anthropic, and paid LLM APIs

Run powerful AI models locally with zero API costs. Complete privacy, unlimited usage, no subscriptions.

Why Ollama?

  • 💰 Free Forever - No API keys, no subscriptions, no usage limits
  • 🔒 Privacy First - Your data never leaves your machine
  • ⚡ Fast - Local inference, no network latency
  • 🎯 Production Ready - Used by thousands of developers worldwide
  • 🔧 Easy Setup - One command installation

Quick Start

# Install Ollama
/setup-ollama

# Or manually:
curl -fsSL https://ollama.com/install.sh | sh

# Pull a model
ollama pull llama3.2

# Start using it!
ollama run llama3.2

Available Models

Code Generation

  • CodeLlama 34B - Best for code generation
  • Qwen2.5-Coder 32B - Excellent coding assistant
  • DeepSeek-Coder 33B - Strong code understanding

General Purpose

  • Llama 3.2 70B - Meta's flagship model
  • Mistral 7B - Fast and efficient
  • Mixtral 8x7B - High quality reasoning

Specialized

  • Phi-3 14B - Microsoft's efficient model
  • Gemma 27B - Google's open model
  • Command-R 35B - Cohere's command model

Replace Paid APIs

OpenAI GPT-4 → Llama 3.2 70B

# Before (Paid - $0.03/1K tokens)
from openai import OpenAI
client = OpenAI(api_key="sk-...")
response = client.chat.completions.create(
    model="gpt-4",
    messages=[{"role": "user", "content": "Hello"}]
)

# After (Free)
import ollama
response = ollama.chat(
    model="llama3.2",
    messages=[{"role": "user", "content": "Hello"}]
)

Anthropic Claude → Mistral

# Before (Paid - $0.015/1K tokens)
from anthropic import Anthropic
client = Anthropic(api_key="sk-ant-...")
message = client.messages.create(
    model="claude-3-5-sonnet-20241022",
    messages=[{"role": "user", "content": "Hello"}]
)

# After (Free)
import ollama
response = ollama.chat(
    model="mistral",
    messages=[{"role": "user", "content": "Hello"}]
)

System Requirements

Minimum:

  • 8GB RAM for 7B models
  • 16GB RAM for 13B models
  • 32GB RAM for 33B+ models

Recommended:

  • NVIDIA GPU with 8GB+ VRAM (10x faster)
  • Apple Silicon (M1/M2/M3) works great
  • AMD GPUs supported

⚠️ Rate Limits & Resource Constraints

No API Limits (Local Deployment)

Hardware-Based "Rate Limits"

Unlike API services, Ollama's constraints are hardware-based, not usage-based:

1. Memory Constraints

Multiple Agents on Same Machine:

# With 32GB RAM, you can run:
# - 3-4 agents using 7B models (8GB each)
# - 2 agents using 13B models (16GB each)
# - 1 agent using 33B model (32GB)

# Agent coordination example
from ollama import Client
import asyncio

async def agent_task(agent_name, model):
    client = Client()
    # Ollama automatically queues requests if busy
    response = await client.chat(
        model=model,
        messages=[{"role": "user", "content": f"Task for {agent_name}"}]
    )
    return response
…

2. Disk Space Requirements

Multi-Model Strategy:

# Storage planning for 3 agents
ollama pull llama3.2      # 4.7 GB (general purpose)
ollama pull codellama     # 13 GB (code generation)
ollama pull mistral       # 4.1 GB (fast responses)
# Total: ~22 GB disk space

3. Inference Speed "Limits"

Agent Best Practice - Queue Management:

# Smart request queuing for single GPU
from queue import Queue
import threading

class LocalLLMCoordinator:
    def __init__(self, max_concurrent=3):
        self.queue = Queue()
        self.max_concurrent = max_concurrent
        self.active_requests = 0

    def process_request(self, agent_id, prompt):
        # Wait if too many concurrent requests
        while self.active_requests >= self.max_concurrent:
            time.sleep(0.1)

        self.active_requests += 1
        try:
            response = ollama.chat(
…

Registration & Setup Requirements

Agent Strategies for Single Machine

Scenario: 10 Agents on One Machine (32GB RAM, NVIDIA RTX 3060)

# Strategy 1: Shared model pool (most efficient)
class AgentPool:
    def __init__(self):
        self.model = "llama3.2"  # Single model loaded once
        self.cache = {}  # Shared response cache

    async def agent_request(self, agent_id, prompt):
        # Check cache first (avoid redundant inference)
        cache_key = hash(prompt)
        if cache_key in self.cache:
            return self.cache[cache_key]

        # Process request
        response = await ollama.chat_async(
            model=self.model,
            messages=[{"role": "user", "content": prompt}]
        )
…

Before you install

  • Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
  • Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
  • Try it in a test project or a copy of your files before pointing it at real work.
  • Pin the version you tested, and review changes before updating.
  • Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.

FAQ

What is Ollama Local Ai?

Ollama Local Ai is a plugin for Claude Code and Claude Cowork from the jeremylongshore/tons-of-skills-marketplace repository on GitHub. Run AI models locally with Ollama - free alternative to OpenAI, Anthropic, and other paid LLM APIs. Zero-cost, privacy-first AI infrastructure.

How do I install Ollama Local Ai in Claude Code?

Add the repository as a plugin marketplace: claude plugin marketplace add jeremylongshore/tons-of-skills-marketplace Install the plugin: claude plugin install ollama-local-ai@<marketplace-name>, using the marketplace name from the repository's .claude-plugin/marketplace.json. Restart the session if the new skills or commands don't appear straight away.

Can I use Ollama Local Ai in Claude Cowork?

Open Customize → Plugins and choose Add marketplace. Enter jeremylongshore/tons-of-skills-marketplace (the owner/repo shorthand works for GitHub). Find Ollama Local Ai in the list, click Install, then connect any connectors it needs from its Connectors tab.

Is Ollama Local Ai safe to install?

It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.

Similar resources

Browse all skills, subagents, and plugins →

Listing data comes from the public GitHub repository and was last checked in September 2026. Excerpts are © their authors and shared under MIT. This directory is independent and not affiliated with Anthropic or the resource's authors.