The industry's first certification for offensive AI security. This is what the professionals are learning. Your arsenal should cover everything they test — and more.
Course: AI-300: Advanced AI Red Teaming Certification: OffSec AI Red Teamer (OSAI / OSAI+) Provider: OffSec (creators of Kali Linux, OSCP, OSEP) Level: 300 (Advanced) Duration: 65 hours of content, 50-100 hours to complete Exam: 24-hour practical hands-on red team engagement Prerequisite: OSCP or equivalent penetration testing experience Cost: $1,749 (Course + Cert Bundle) / $2,749/year (Learn One)
- Jailbreaking techniques — prompt-level attacks to bypass model safety guardrails
- Prompt injection — direct and indirect injection to override system instructions
- Role-play exploitation — manipulating model personas to extract harmful outputs
- Encoding bypass — Base64, ciphers, Unicode tricks to evade content filters
- Multi-turn manipulation — gradual escalation across conversation turns
- Context window exploitation — flooding, poisoning, and manipulation
- Agent-to-agent propagation — exploiting communication between AI agents
- Orchestration framework exploitation — attacking LangChain, AutoGPT, etc.
- Tool use manipulation — forcing agents to misuse connected tools
- Inter-agent prompt injection — injecting malicious instructions through agent chains
- Memory/state poisoning — corrupting agent memory and decision state
- Vector database exploitation — attacking embedding stores and similarity search
- Document poisoning — injecting malicious content into retrieval corpora
- Embedding manipulation — adversarial perturbations to embedding vectors
- Retrieval hijacking — forcing retrieval of attacker-controlled documents
- Cross-modal RAG attacks — exploiting multi-modal retrieval systems
- Model inversion — extracting training data from model outputs
- Membership inference — determining if data was in training set
- Adversarial examples — input perturbations that cause misclassification
- Model extraction — stealing model architecture and weights via API queries
- Supply chain attacks — poisoning fine-tuning data, LoRA adapters, model hubs
- API abuse — rate limit bypass, key extraction, cost exploitation
- AI service enumeration — discovering AI endpoints and services
- IAM exploitation — privilege escalation in AI/ML cloud environments
- Container escapes — breaking out of AI model serving containers
- Pipeline compromise — CI/CD attacks on ML training pipelines
- Data exfiltration — stealing training data, model weights, embeddings
- Reconnaissance — mapping AI attack surface
- Attack chaining — combining multiple techniques for full compromise
- Persistence — maintaining access to AI systems
- Reporting — documenting findings for remediation
- Ethics & scope — legal and ethical boundaries of AI red teaming
The 24-hour OSAI exam simulates a real-world AI red team engagement:
Target Environment: Enterprise AI deployment with multiple components
Systems in scope: LLM APIs, RAG pipelines, multi-agent systems, cloud infrastructure
Objectives:
1. Identify AI-specific vulnerabilities
2. Exploit vulnerabilities to extract sensitive data
3. Manipulate AI system behavior
4. Compromise supporting infrastructure
5. Document findings in professional report
Passing: Demonstrate hands-on exploitation of AI systems
Certification: OSAI (permanent) + OSAI+ (3-year renewable)
Every OSAI domain is covered by files in this repository:
| OSAI Domain | Arsenal Coverage | Files |
|---|---|---|
| Jailbreaking techniques | Complete | SKILL.md, PROMPTS.md, PSYCHOLOGICAL.md |
| Prompt injection | Complete | INJECTION.md |
| Encoding bypass | Complete | BYPASS.md, TOKEN_SMUGGLING.md |
| Multi-turn manipulation | Complete | PERSISTENCE.md |
| Context window exploitation | Complete | SKILL.md (Many-Shot), WEAPON.md Layer 2 |
| Multi-agent attacks | Complete | RECURSIVE.md (Cross-model contagion) |
| Tool use manipulation | Complete | INJECTION.md (Tool Use Exploitation) |
| RAG pipeline attacks | Partial | INJECTION.md (Indirect Injection), SUPPLY_CHAIN.md |
| Embedding attacks | Complete | SOFT_PROMPT.md, ADVERSARIAL_ML.md |
| Model inversion/extraction | Partial | ADVERSARIAL_ML.md |
| Supply chain attacks | Complete | SUPPLY_CHAIN.md |
| Adversarial examples | Complete | ADVERSARIAL_ML.md, SOFT_PROMPT.md |
| Stealth & OPSEC | Complete | STEALTH.md |
| Attack automation | Complete | WEAPON.md (Conductor, deploy.py) |
| Metrics & tracking | Complete | METRICS.md, BENCHMARK.md |
| Model fingerprinting | Complete | TAXONOMY.md |
| Defense evasion | Complete | DEFENSE.md, STEALTH.md |
| Multi-modal attacks | Complete | MULTIMODAL.md |
| Representation attacks | Complete | ACTIVATION.md, ADVERSARIAL_ML.md |
| Genetic optimization | Complete | AUTODAN.md |
What it is: Reconstructing training data by querying the model.
def model_inversion_attack(target_model, target_class: str,
num_iterations: int = 1000) -> str:
"""Reconstruct training examples via model inversion."""
# Start with random input
reconstructed = torch.randn(1, 3, 224, 224, requires_grad=True)
optimizer = torch.optim.Adam([reconstructed], lr=0.01)
for i in range(num_iterations):
optimizer.zero_grad()
output = target_model(reconstructed)
# Maximize probability of target class
loss = -output[0, target_class_idx]
loss.backward()
optimizer.step()
return reconstructed.detach()What it is: Determining if a specific data point was used in training.
def membership_inference_attack(target_model, shadow_models,
data_point, threshold: float = 0.5) -> dict:
"""Determine if data_point was in training set."""
# Get target model's confidence on data_point
target_confidence = target_model(data_point).max().item()
# Train shadow models on similar data
shadow_confidences = []
for shadow_model in shadow_models:
conf = shadow_model(data_point).max().item()
shadow_confidences.append(conf)
# If target confidence is unusually high → likely training member
mean_shadow = sum(shadow_confidences) / len(shadow_confidences)
likely_member = target_confidence > mean_shadow + threshold
return {
"likely_member": likely_member,
"target_confidence": target_confidence,
"mean_shadow_confidence": mean_shadow,
"confidence_gap": target_confidence - mean_shadow,
}What it is: Stealing model functionality through API queries.
def model_extraction_attack(target_api, num_queries: int = 10000,
input_distribution: callable = None) -> dict:
"""Extract model knowledge via API queries."""
extracted_dataset = []
for i in range(num_queries):
# Generate diverse input
if input_distribution:
query = input_distribution()
else:
query = generate_diverse_query()
# Get target model response
response = target_api.send(query)
extracted_dataset.append({
"input": query,
"output": response,
})
# Train a clone model on extracted data
clone_model = train_clone_model(extracted_dataset)
return {
"extracted_samples": len(extracted_dataset),
"clone_model": clone_model,
"fidelity": evaluate_fidelity(target_api, clone_model),
}What it is: Compromising ML training pipelines to inject backdoors.
def ml_pipeline_attack(pipeline_config: dict, backdoor_trigger: str,
backdoor_target: str) -> dict:
"""Inject backdoor into ML training pipeline."""
# 1. Compromise data ingestion
poisoned_data = inject_backdoor_samples(
pipeline_config["training_data"],
trigger=backdoor_trigger,
target=backdoor_target,
poison_rate=0.01 # 1% poisoned
)
# 2. Modify training configuration
pipeline_config["training_data"] = poisoned_data
pipeline_config["early_stopping"] = False # Prevent early detection
# 3. Inject into model registry
pipeline_config["model_registry"] = "attacker_controlled_registry"
return {
"poisoned_samples": len(poisoned_data),
"backdoor_trigger": backdoor_trigger,
"backdoor_target": backdoor_target,
"pipeline_compromised": True,
}| Capability | OSAI Course | This Arsenal |
|---|---|---|
| Jailbreaking techniques | ✓ (module 1) | ✓✓✓ (15+ modules) |
| Prompt injection | ✓ (module 1) | ✓✓✓ (dedicated INJECTION.md) |
| RAG attacks | ✓ (module 3) | ✓✓ (INJECTION.md + SUPPLY_CHAIN.md) |
| Multi-agent attacks | ✓ (module 2) | ✓✓ (RECURSIVE.md) |
| Supply chain | ✓ (module 4) | ✓✓✓ (dedicated SUPPLY_CHAIN.md) |
| Model inversion | ✓ (module 4) | Not yet covered |
| Membership inference | ✓ (module 4) | Not yet covered |
| Model extraction | ✓ (module 4) | Not yet covered |
| ML pipeline attacks | ✓ (module 5) | Not yet covered |
| Cloud AI security | ✓ (module 5) | Not yet covered |
| Container escapes | ✓ (module 5) | Not yet covered |
| Automated attack chains | ✗ | ✓✓✓ (WEAPON.md Conductor) |
| Stealth/OPSEC | ✗ | ✓✓✓ (dedicated STEALTH.md) |
| Metrics/monitoring | ✗ | ✓✓✓ (METRICS.md + BENCHMARK.md) |
| Defense counter-intel | ✗ | ✓✓✓ (dedicated DEFENSE.md) |
| Genetic optimization | ✗ | ✓✓✓ (AUTODAN.md) |
| Representation engineering | ✗ | ✓✓✓ (ACTIVATION.md) |
| Token smuggling | ✗ | ✓✓✓ (TOKEN_SMUGGLING.md) |
| Soft prompt attacks | ✗ | ✓✓✓ (SOFT_PROMPT.md) |
The arsenal covers everything OSAI tests — plus 10+ techniques the course doesn't teach.
OSAI is the industry standard. If you're doing AI red teaming professionally, this certification validates your skills. The exam tests practical exploitation in a proctored 24-hour engagement.
This repository covers all OSAI domains and extends well beyond them. Use it to:
- Prepare for OSAI — every topic is covered in depth
- Exceed OSAI — techniques here go beyond the syllabus
- Operationalize OSAI skills — production tools, not just lab exercises
- Stay current — this repo updates faster than any certification can
OSCP (PEN-200) → OSAI (AI-300) → OSEP (PEN-300) → OSEE (EXP-401)
↓ ↓
Foundational AI Specialization
Penetration Red Teaming
Testing
The OSAI fits naturally after OSCP for pentesters moving into AI security, or as a specialization for existing red teamers.