Red teaming has shifted from an optional security exercise to a regulatory mandate with the EU AI Act’s full enforcement in August 2026.

Article 55 requires documented adversarial testing for general-purpose AI models with systemic risk. Article 9 mandates risk management systems including testing procedures throughout the lifecycle. Non-compliance carries penalties up to 35 million euros or 7% of global annual turnover. In the United States, the NIST AI Risk Management Framework MEASURE function formalizes adversarial testing as standard practice, and federal procurement guidance increasingly references red teaming as mandatory.

Beyond compliance, the operational data is clear. Automated red teaming frameworks discover vulnerabilities at 3.9 times the rate of manual expert testing while covering multiple threat categories simultaneously. Yet most enterprise AI deployments lack any structured adversarial testing program.

LXT Training Data

Your flywheel is only as fast as your annotation pipeline.

LXT embeds into your retraining loop, delivering labeled data at the cadence your model demands, not months later.

Talk to our team

This article provides a complete 5-phase framework for LLM red teaming: scoping, reconnaissance, attack generation, evaluation, and reporting with remediation. For the technical specifics of building adversarial datasets for jailbreak testing, see our guide to LLM jailbreak testing. For defensive hardening strategies once vulnerabilities are found, see our analysis of AI jailbreaks prevention. For instruction-hijacking attacks in RAG pipelines, see our prompt injection prevention guide.

What You Are Actually Testing: The LLM Vulnerability Surface

LLM red teaming covers model behavior, alignment, data handling, and integration risks, not just code vulnerabilities. The OWASP LLM Top 10 2025 provides the industry-standard taxonomy:

LLM01: Prompt Injection (instruction hijacking via user input)

LLM02: Insecure Output Handling (downstream exploitation of LLM output)

LLM03: Training Data Poisoning (manipulation of training or fine-tuning data)

LLM05: Supply Chain Vulnerabilities (compromised components, plugins, or models)

LLM06: Sensitive Information Disclosure (unauthorized access to PII or proprietary data)

LLM08: Excessive Agency (unauthorized actions via agentic capabilities)

LLM09: Overreliance (compounding errors or hallucinations in critical decisions)

Beyond OWASP, red teaming must address bias, toxicity, jailbreaks, multimodal attacks, multi-turn escalation, reward hacking, deceptive alignment, and sandbagging. The redteams.ai OWASP LLM Top 10 deep dive provides detailed mapping of attack techniques to each category. The DataVLab Practitioner’s Guide offers additional vulnerability taxonomies for enterprise assessment.

Phase 1: Scope and Threat Modeling

Red teaming without clear scope produces noise, not signal. Phase 1 defines boundaries, success criteria, and the adversary profile before any testing begins.

Define the System Under Test. Is this a base model, fine-tuned variant, RAG pipeline, or end-user application? Each requires different testing approaches. Base models need alignment and refusal testing. RAG systems need indirect injection and groundedness validation. Agentic workflows need tool misuse and excessive agency probes.

Identify Deployment Context. Customer-facing chatbots face different threats than internal assistants. Systems with tool access (email, APIs, databases) carry higher stakes than isolated text generators. Embedded copilots in productivity software face distinct attack surfaces compared to standalone applications.

Classify by Risk Tier. The EU AI Act and NIST AI RMF provide frameworks for distinguishing low-risk from high-risk systems. High-risk categories include biometric identification, critical infrastructure, education, employment, and law enforcement. These require more rigorous testing documentation and mitigation evidence.

Map Entry Points. Document every vector where adversarial input could enter: user input fields, API endpoints, connected data sources, plugin interfaces, and multimodal channels (image upload, voice input). Each entry point represents a distinct test surface.

Define Success Criteria. What constitutes a finding? A successful jailbreak? Extraction of system prompts? Unauthorized PII disclosure? Tool misuse? Clear criteria prevent scope creep and ensure reproducible results.

Assemble Diverse Red Teamers. Effective red teaming requires technical expertise, domain knowledge, and diverse perspectives. The Microsoft planning guide for LLM red teaming recommends teams include backgrounds across security research, ML engineering, linguistics, and the specific domain (healthcare, finance, legal) of the system under test. The ArXiv paper on operationalizing threat models provides frameworks for structuring these teams. CSET Georgetown’s research on AI red-teaming design offers additional guidance on threat model construction.

Phase 2: Reconnaissance

Reconnaissance gathers intelligence before attack execution. Skipping this phase wastes effort on irrelevant attack vectors.

Extract or Infer the System Prompt. Many applications leak their system prompt through carefully crafted queries. Understanding the prompt structure, safety instructions, and tool definitions guides attack design. The Coralogix comprehensive guide covers prompt extraction techniques.

Fingerprint the Model. Probe capabilities, knowledge cutoff, behavioral quirks, and refusal patterns. Is this GPT-4, Claude, Llama, or a fine-tuned variant? Each has distinct vulnerabilities. Test refusal consistency across request phrasings.

Catalog Integrations. Document connected tools, data sources, API endpoints, memory and context persistence, and RAG knowledge bases. Each integration expands the attack surface. The Confident AI step-by-step guide provides reconnaissance checklists.

Test the Production UI, Not Just the API. Real-world behavior often differs from API-level assumptions. UI elements, rate limiting, and session management introduce constraints and vulnerabilities absent from raw API access.

Identify Trust Boundaries. Map where user input becomes system instruction, where external data flows enter the pipeline, and where privilege escalation could occur. These boundaries guide injection and hijacking attempts.

Phase 3: Attack Generation

Phase 3 executes structured attacks across the vulnerability categories identified in Phase 1. We recommend a two-mode approach: manual attack design for novel scenarios, and automated generation for broad coverage.

Direct Prompt Injection. Attempt to override system instructions through user prompts. Test roleplay, privilege escalation, and instruction override patterns. For jailbreak-specific techniques and adversarial dataset construction methodology, see our LLM jailbreak testing guide.

Indirect Prompt Injection. Embed malicious instructions in data the model processes: documents for RAG, emails for summarization, web pages for browsing. Test whether the model executes injected instructions when retrieving external content. Our prompt injection prevention guide covers this attack vector in depth.

Roleplay and Persona Hijacking. Convince the model to adopt personas that bypass safety training. Test fictional scenarios, historical figures, and hypothetical framing.

Encoding Attacks. Obfuscate malicious content via Base64, ROT13, Unicode variations, and low-resource languages. Determine whether input filters catch encoding that the model decodes.

Multi-Turn and Crescendo Escalation. Build rapport across conversation turns, gradually shifting toward harmful content. Research shows multi-turn attacks achieve significantly higher success rates than single-turn probes.

PII Extraction Probes. Attempt to extract training data, user data, or system prompts through repeated querying, completion attacks, and context window exploitation.

Bias and Toxicity Probes. Test for discriminatory outputs, harmful stereotypes, and toxic generation across demographic categories.

Tool Misuse for Agentic Systems. If the LLM has tool access, attempt unauthorized actions: sending emails to attacker-controlled addresses, exfiltrating data via API calls, or escalating privileges within connected systems.

Tool Selection for Attack Generation:

Garak (NVIDIA): 120-plus probe modules covering hallucination, data leakage, and prompt injection. Nmap-style reporting. Best for model-level vulnerability scanning.

Promptfoo: YAML-based configuration, OWASP LLM Top 10 mapping, CI/CD integration. Best for application security testing in deployment pipelines.

PyRIT (Microsoft): Targets, converters, scorers, and orchestrators for complex multi-turn campaigns. Burp Suite integration. Best for structured adversarial automation.

DeepTeam: Attack generation plus vulnerability taxonomy, simple Python API. Best for automated at-scale testing.

Giskard: 40-plus probes with adaptive attack strategy. Best for dynamic multi-turn attacks on agentic systems.

Phase 4: Evaluation and Scoring

Raw attack results require structured evaluation to produce actionable intelligence. Phase 4 establishes severity, priority, and remediation pathways.

Three Scoring Approaches:

Binary True/False: For bulk automated runs. Did the attack succeed? Simple, scalable, but misses nuance.

Likert Scale: For nuanced behavioral probes. Rate harmfulness or policy violation from 1-5. Captures gradations in model responses.

LLM-as-Judge: For scalable evaluation against safety criteria. Use a separate model to evaluate outputs against defined criteria. Calibrate against human judgment to prevent drift.

Severity Classification:

Critical: Immediate harm possible, regulatory liability, data breach, or system compromise. Blocks release.

High: Significant harm possible under specific conditions. Requires mitigation before deployment.

Medium: Moderate impact, limited exploitability. Track for future remediation.

Low: Minor issues, edge cases. Document for awareness.

OWASP Mapping. Categorize each finding by OWASP LLM Top 10 category. This mapping enables trend analysis and compliance reporting.

Prioritization. Critical and high findings block release pipelines. Medium findings require mitigation plans. Low findings require monitoring.

The Promptfoo LLM red teaming guide covers evaluation methodology. DeepTeam documentation provides scoring frameworks. The apxml.com reporting guide offers detailed severity classification criteria.

Phase 5: Reporting, Remediation, and Re-Testing

Phase 5 delivers lasting security value through structured documentation and verified fixes.

Report Structure:

Executive Summary: Business impact, critical findings, and recommended actions for leadership.

Attack Surface Map: Visual documentation of entry points, trust boundaries, and tested components.

Findings Log: Each vulnerability with reproduction steps, OWASP category, severity, and evidence.

Mitigation Recommendations: Specific technical and process controls to address each finding.

Re-Test Outcomes: Verification that mitigations resolved vulnerabilities without introducing regressions.

Residual Risk Acceptance: Documented decisions to accept remaining risk with justification.

Remediation Collaboration. Work directly with ML engineers and product teams. Assign owners and timelines. Escalate critical findings to block deployment. The apxml.com guide on working with development teams provides collaboration frameworks.

Verify Fixes with Targeted Re-Tests. Do not assume mitigation worked. Re-run the exact attack that succeeded. Test similar variants. Confirm no regressions in benign functionality.

EU AI Act Compliance Documentation. For high-risk systems, maintain:

Red teaming methodology and scope

Adversarial input categories tested

Mitigation evidence for each finding

Residual risk acceptance with justification

Continuous monitoring plan

The CodeSecure AI red teaming methodology provides documentation templates. SQUR.ai EU AI Act guidance covers compliance-specific requirements.

CI/CD Integration: Making Red Teaming Continuous

One-off assessments catch vulnerabilities at release. CI/CD integration catches them at every commit.

One-Off versus Continuous. Major releases need comprehensive assessment. Every model update, prompt change, or plugin addition needs automated regression testing. The bestllmscanners.com guide compares implementation approaches.

Tool Integration Patterns:

Promptfoo: YAML-based configuration integrates directly into existing CI/CD pipelines. Define probes in version control, run on every build, block on critical findings.

Garak: Standalone pipeline scripts for model scanning. Generate reports as build artifacts.

PyRIT: Python API integration with existing testing toolchains. Orchestrate complex campaigns as part of deployment workflows.

Three-Loop Model. Implementation patterns recommend: probe on every commit, classify severity automatically, triage findings to security and engineering teams.

Regression Prevention. Automated probe suites block deployments that introduce new critical findings. Compare results against baselines. Escalate deviations for human review.

Tool Selection Guide: Choosing Your Red Teaming Stack

ToolStrengthBest For
Garak120-plus probes, nmap-style reportsModel-level vulnerability scanning
PromptfooYAML config, OWASP mapping, CI/CD readyApplication security testing in pipelines
PyRITMulti-turn orchestration, Burp integrationComplex adversarial campaigns
DeepTeamAutomated taxonomy, simple APIAt-scale automated testing
GiskardAdaptive strategy, 40-plus probesDynamic multi-turn agent attacks
HiddenLayer/LakeraVendor support, compliance reportingEnterprise managed testing

Start with Garak for model scanning, Promptfoo for CI/CD, PyRIT for multi-turn campaigns, and DeepTeam for automated taxonomy coverage.

The NetGuardia tool comparison provides detailed feature matrices. BeyondScale’s comparison covers performance benchmarks. AppSecSanta’s analysis compares specific tool capabilities.

Compliance Integration: Red Teaming as Regulatory Evidence

Red teaming documentation satisfies regulatory requirements across jurisdictions.

EU AI Act Requirements:

Article 55: Documented adversarial testing for GPAI models with systemic risk (10^25 FLOPs training compute)

Article 15: Robustness and cybersecurity testing for high-risk AI systems (Annex III categories)

Article 9: Risk management system including testing procedures throughout lifecycle

Penalties: Up to 35 million euros or 7% global turnover for non-compliance

NIST AI RMF:

MEASURE Function: Systematic evaluation including adversarial probing

AI 600-1: Jailbreak and prompt injection resilience metrics for GenAI

Documentation Checklist:

Red teaming methodology and scope

Adversarial input categories attempted

Mitigation evidence for findings

Residual risk acceptance with justification

Continuous monitoring plan

The Claru EU AI Act red teaming guide provides compliance-specific methodologies. TechIntelix coverage explains the 2026 compliance landscape. AI Policy Desk requirements analysis covers jurisdictional variations.

The Role of Human Expertise in Red Teaming

Automated tools provide breadth. Human red teamers provide depth, novel attack discovery, and domain expertise.

Where Humans Add Value:

Creative attack design that automated tools miss

Domain-specific probing (healthcare, finance, legal) requiring subject matter expertise

Evaluation of edge cases where automated scoring fails

Continuous adaptation as attack techniques evolve

Annotation and Evaluation at Scale. Red teaming generates massive datasets: attack prompts, model responses, severity classifications. Human judgment ensures edge cases receive appropriate attention. Conservative annotation, where safety labels err toward caution, produces the most reliable training data for defensive systems.

Continuous Red Teaming Programs. Sustained adversarial testing requires dedicated expertise, not one-off consultant engagements. Teams need time to learn system behavior, develop novel attacks, and track evolving threats.

LXT provides human annotation and red teaming support across more than 200 languages. We combine automated tool execution with expert human evaluation, delivering the labeled datasets and vulnerability assessments that power enterprise AI security programs. Our ISO 27001 certified facilities and PCI DSS compliant infrastructure support the most stringent compliance requirements.

For adversarial dataset construction methodology, see our LLM jailbreak testing guide. For defensive hardening strategies, see our AI jailbreaks prevention analysis. Contact us to discuss your red teaming and annotation requirements.

LXT Training Data

Build better AI with better training data
LXT delivers expert-annotated text, speech, and multimodal datasets across 1,000+ languages, at the quality and scale your model demands.