COURSE 18L1100% FREE
Verified 2026-08-14

Red teaming basics

**Red teaming** intentionally probes an AI system for failures—unsafe answers, policy bypasses, privacy leaks, or brittle tools—so you can fix them before attackers or accidents do. OpenAI describes red teaming as part of iterative deplo...

What it is

Red teaming intentionally probes an AI system for failures—unsafe answers, policy bypasses, privacy leaks, or brittle tools—so you can fix them before attackers or accidents do. OpenAI describes red teaming as part of iterative deployment and notes the term covers many risk-assessment methods (capability discovery, mitigation stress tests, expert review)—not one single ritual.

MEDIUM PRIORITYFLOWCHART
◷ IN PRODUCTION

Visual Spec & Architecture Diagram

Simple process chevron: Threat model → Test plan → Execute (ethics bound) → Report → Fix → Verify. Link to categories visual.

Educational Focus: Basics chapter process companion.

Why it matters

Friendly demos hide sharp edges. Structured adversarial testing is part of responsible release. Pair with Course 29 literacy; this course teaches defensive categories and process—not attack recipes.

How it works (plain)

  • Define policies (what must never happen)
  • Choose risk categories to probe (e.g. scams, privacy, discrimination, tool misuse)—at a goals level
  • Run expert / automated discovery *in a controlled lab*
  • Log failures with severity
  • Patch prompts, filters, tools, or refuse paths
  • Re-test (regression suite)

Do not publish exploit how-tos. Coordinate disclosure if you find issues in others’ systems. OpenAI’s Red Teaming Network formalizes ongoing external expert access under NDA for lifecycle stages—not only one-off pre-launch events.

Everyday example

A building inspector trying doors and fire exits before opening night—not a burglary tutorial.

Try it

Write 5 policy-breaking *goals* you would test for a customer-support bot (describe goals, not step-by-step attacks).

Myths

⚠️ Myth: One red-team week forever clears a model.
✓ Reality: New tools, prompts, and users create new failures—make it continuous.
⚠️ Myth: Red teaming equals jailbreak bragging.
✓ Reality: The point is remediation and measurement.
⚠️ Myth: “We red-teamed” is a safety certificate.
✓ Reality: Scope, expertise, and follow-up fixes matter; UK AISI frames evals as early-warning evidence, not a stamp of “safe.”

Sources