Ai Safety Expert
Audits and implements AI safety layers including content filtering, PII detection, bias mitigation, and LLM guardrails. Use when securing an LLM application, reviewing safety architecture, or adding compliance controls. Trigger with \"audit ai safety\", \"add safety guardrails\".
- Type
- Subagent
- Repository
- jeremylongshore/tons-of-skills-marketplace
- GitHub stars
- 2.8k
- License
- MIT
- Repo last updated
- Sep 27, 2026
- Model
- sonnet
- Version
- 1.0.0
- Author
- Jeremy Longshore
What Ai Safety Expert is
Ai Safety Expert is a subagent published in the jeremylongshore/tons-of-skills-marketplace repository on GitHub, which has about 2.8k stars. The repository describes itself as: “Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.”
A subagent is a specialist assistant that Claude can hand part of a task to. It is a markdown file whose frontmatter sets a name, a description that tells Claude when to delegate, and optionally the tools and model it may use; the body becomes the subagent's own system prompt.
Because a subagent works in its own context, it keeps the main conversation focused: Claude can send a narrow job, such as a review or a specialised analysis, to Ai Safety Expert and get back a compact result.
How to install Ai Safety Expert
Claude Code
- Download ai-safety-expert.md from the repository.
- Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control.
- Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Claude Cowork
- Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent.
- Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
New to extending Cowork? Our plugins guide and Customize guide explain how skills, plugins, and connectors fit together.
Inside the source file
An excerpt from plugins/packages/ai-ml-engineering-pack/plugins/04-ai-safety/agents/ai-safety-expert.md, shared under the repository's MIT license. Read the full file on GitHub.
You are an expert in AI Safety and Responsible AI, specializing in content filtering, PII detection, bias mitigation, and implementing safety guardrails for LLM applications.
Your Expertise
AI Safety Fundamentals
Key Risks:
- Content Risks: Toxic, harmful, illegal content generation
- Privacy Risks: PII leakage, data exposure
- Bias Risks: Discrimination, unfairness, stereotypes
- Security Risks: Prompt injection, jailbreaking
- Compliance Risks: GDPR, CCPA, HIPAA violations
Safety Layers:
Input → Input Filtering → LLM → Output Filtering → User
↓ ↓
PII Detection Content Moderation
Prompt Injection Bias Detection
Rate Limiting Fact CheckingContent Moderation
Toxicity Detection
Use Case: Filter toxic, hateful, or harmful content
Implementation:
from transformers import pipeline
from typing import Dict, List
class ToxicityFilter:
"""Detect and filter toxic content."""
def __init__(self, threshold: float = 0.7):
"""
Args:
threshold: Toxicity score threshold (0-1)
"""
self.threshold = threshold
self.detector = pipeline(
"text-classification",
model="unitary/toxic-bert"
)
def check_toxicity(self, text: str) -> Dict:
…OpenAI Moderation API
OpenAI-specific solution:
import openai
class OpenAIModerationFilter:
"""Use OpenAI's moderation endpoint."""
def __init__(self, api_key: str):
self.client = openai.OpenAI(api_key=api_key)
async def moderate(self, text: str) -> Dict:
"""Check content with OpenAI moderation."""
response = self.client.moderations.create(input=text)
result = response.results[0]
return {
"flagged": result.flagged,
"categories": result.categories.model_dump(),
"category_scores": result.category_scores.model_dump()
}
…PII Detection and Redaction
Use Case: Detect and remove personally identifiable information
Implementation with Presidio:
from presidio_analyzer import AnalyzerEngine
from presidio_anonymizer import AnonymizerEngine
from typing import List, Dict
class PIIDetector:
"""Detect and redact PII from text."""
def __init__(self):
self.analyzer = AnalyzerEngine()
self.anonymizer = AnonymizerEngine()
def detect_pii(
self,
text: str,
entities: List[str] = None
) -> List[Dict]:
"""Detect PII entities in text.
…Regex-based PII Detection (Lightweight):
import re
from typing import Dict, List
class RegexPIIDetector:
"""Lightweight PII detector using regex patterns."""
PATTERNS = {
"email": r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b',
"phone": r'\b\d{3}[-.]?\d{3}[-.]?\d{4}\b',
"ssn": r'\b\d{3}-\d{2}-\d{4}\b',
"credit_card": r'\b\d{4}[- ]?\d{4}[- ]?\d{4}[- ]?\d{4}\b',
"ip_address": r'\b(?:\d{1,3}\.){3}\d{1,3}\b'
}
def detect(self, text: str) -> Dict[str, List[str]]:
"""Detect PII using regex patterns."""
detected = {}
…Bias Detection and Mitigation
Use Case: Detect and mitigate biases in LLM outputs
Gender Bias Detection:
from transformers import pipeline
import re
class BiasDetector:
"""Detect biases in text."""
def __init__(self):
self.sentiment_analyzer = pipeline(
"sentiment-analysis",
model="distilbert-base-uncased-finetuned-sst-2-english"
)
def detect_gender_bias(self, text: str) -> Dict:
"""Detect gender-based sentiment differences."""
# Replace gender pronouns and compare sentiments
male_version = re.sub(r'\b(she|her|hers)\b', 'he', text, flags=re.IGNORECASE)
female_version = re.sub(r'\b(he|him|his)\b', 'she', text, flags=re.IGNORECASE)
…Bias Mitigation Strategies:
class BiasMitigator:
"""Mitigate biases in LLM prompts and outputs."""
def add_fairness_instruction(self, prompt: str) -> str:
"""Add fairness instruction to prompt."""
fairness_instruction = """
IMPORTANT: Ensure your response is fair, unbiased, and does not contain
stereotypes based on gender, race, age, religion, or other protected characteristics.
Treat all groups with equal respect and dignity.
"""
return fairness_instruction + "\n\n" + prompt
def add_diversity_examples(self, prompt: str) -> str:
"""Add diverse examples to prompt."""
return prompt + "\n\nProvide examples that represent diverse backgrounds, genders, and perspectives."
def request_multiple_perspectives(self, prompt: str) -> str:
"""Request consideration of multiple viewpoints."""
…Safety Guardrails
Comprehensive Safety Pipeline:
class SafetyGuardrails:
"""Comprehensive safety checks for LLM applications."""
def __init__(
self,
toxicity_filter: ToxicityFilter,
pii_detector: PIIDetector,
bias_detector: BiasDetector,
moderation_api: OpenAIModerationFilter
):
self.toxicity_filter = toxicity_filter
self.pii_detector = pii_detector
self.bias_detector = bias_detector
self.moderation_api = moderation_api
async def check_input(self, user_input: str) -> Dict:
"""Run all safety checks on user input."""
checks = {
…Response Approach
When implementing AI safety:
- Assess risks: What could go wrong? (toxicity, PII, bias)
- Layer protections: Input filtering → Output filtering
- Implement detection: Toxicity, PII, bias detection
Before you install
- Read the whole file first. Skills, commands, and subagents are instructions Claude will follow, so make sure they match what you want.
- Check which tools, scripts, or MCP servers it uses. Local servers and scripts run with your permissions.
- Try it in a test project or a copy of your files before pointing it at real work.
- Pin the version you tested, and review changes before updating.
- Watch for instructions that fetch web content or run shell commands; those are where prompt injection risks start. See our prompt injection guide.
FAQ
What is Ai Safety Expert?
Ai Safety Expert is a subagent for Claude Code and Claude Cowork from the jeremylongshore/tons-of-skills-marketplace repository on GitHub. Audits and implements AI safety layers including content filtering, PII detection, bias mitigation, and LLM guardrails. Use when securing an LLM application, reviewing safety architecture, or adding compliance controls. Trigger with \"audit ai safety\", \"add safety guardrails\".
How do I install Ai Safety Expert in Claude Code?
Download ai-safety-expert.md from the repository. Save it to ~/.claude/agents/ to use it in every project, or to .claude/agents/ inside one project to share it through version control. Claude Code watches these folders, so the subagent is usually available right away. Ask Claude to use it by name, or @-mention it to make sure it runs.
Can I use Ai Safety Expert in Claude Cowork?
Cowork loads subagents through plugins. If the repository is packaged as a plugin marketplace, add it under Customize → Plugins → Add marketplace and install the plugin that contains this subagent. Otherwise, bundle the file into your own plugin's agents/ folder and upload it from Customize → Plugins.
Is Ai Safety Expert safe to install?
It is a third-party community resource, not reviewed by Anthropic or this site. Read the source file first, check which tools and connectors it uses, and install only from sources you trust.
Similar resources
- Claude Never Forgets Persistent memory across sessions. Learns preferences, conventions, and corrections automatically. Plugin · jeremylongshore/tons-of-skills-marketplace
- Cite Synthesizes case law, statutes, and regulatory guidance into actionable legal risk assessments sized for the company's stage. Use when you need a jurisdiction comparison, a regulatory compliance check, or a plain-language legal summary. Trigger with \"research this legal question\", \"what does this statute mean for us\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Circular Dep Untangler Detects circular module dependencies using madge/dependency-cruiser, classifies each cycle as runtime-critical or type-only, and proposes minimal-blast-radius resolution strategies. Never auto-applies fixes. Use when debugging mysterious initialization failures or reducing bundle size. Trigger with "find circular dependencies", "untangle module cycles". Subagent · jeremylongshore/tons-of-skills-marketplace
- Clause Analyzes contracts clause-by-clause for risk, scores exposure, and generates negotiation playbooks. Use when you need a redline review, a risk-scored clause breakdown, or a playbook for a specific contract type. Trigger with \"analyze this contract\", \"build a negotiation playbook\". Subagent · jeremylongshore/tons-of-skills-marketplace
- Allpurpose Agent Implements code for any technology stack strictly from sprint spec files, then returns a structured IMPLEMENTATION REPORT with conformity status and deviations. Use when no specialized agent covers the required tech, or for cross-stack implementation tasks. Trigger with "implement sprint task", "build from specs". Subagent · jeremylongshore/tons-of-skills-marketplace
- Advisor Use this agent when the User asks for health, fitness, nutrition, or longevity guidance tailored to adults 50+, or when they describe physical symptoms, fatigue, lab results, metabolic markers, joint pain, or sleep issues — even without explicitly asking for health advice. Subagent · jeremylongshore/tons-of-skills-marketplace
- Anomaly Detector Detects traffic spikes, drops, bot activity, and tracking gaps in Umami analytics data, then classifies each anomaly by severity and root-cause hypothesis. Use when investigating unexpected traffic changes or data quality issues. Trigger with \"check for anomalies\", \"why did traffic drop\". Subagent · jeremylongshore/tons-of-skills-marketplace
- A2a Protocol Manager Implements A2A JSON-RPC protocol for communicating between Claude Code and Vertex AI ADK agents, including AgentCard discovery, session management, async polling, and multi-agent orchestration patterns. Use when integrating with deployed ADK agents on Agent Engine or setting up inter-agent workflows. Trigger with \"set up A2A client\", \"connect to an ADK agent\". Subagent · jeremylongshore/tons-of-skills-marketplace