Red teaming has shifted from an optional security exercise to a regulatory mandate with the EU AI Act’s full enforcement in August 2026.
Article 55 requires documented adversarial testing for general-purpose AI models with systemic risk. Article 9 mandates risk management systems including testing procedures throughout the lifecycle. Non-compliance carries penalties up to 35 million euros or 7% of global annual turnover. In the United States, the NIST AI Risk Management Framework MEASURE function formalizes adversarial testing as standard practice, and federal procurement guidance increasingly references red teaming as mandatory.
Beyond compliance, the operational data is clear. Automated red teaming frameworks discover vulnerabilities at 3.9 times the rate of manual expert testing while covering multiple threat categories simultaneously. Yet most enterprise AI deployments lack any structured adversarial testing program.
Your flywheel is only as fast as your annotation pipeline.
LXT embeds into your retraining loop, delivering labeled data at the cadence your model demands, not months later.
This article provides a complete 5-phase framework for LLM red teaming: scoping, reconnaissance, attack generation, evaluation, and reporting with remediation. For the technical specifics of building adversarial datasets for jailbreak testing, see our guide to LLM jailbreak testing. For defensive hardening strategies once vulnerabilities are found, see our analysis of AI jailbreaks prevention. For instruction-hijacking attacks in RAG pipelines, see our prompt injection prevention guide.
What You Are Actually Testing: The LLM Vulnerability Surface
LLM red teaming covers model behavior, alignment, data handling, and integration risks, not just code vulnerabilities. The OWASP LLM Top 10 2025 provides the industry-standard taxonomy:
LLM01: Prompt Injection (instruction hijacking via user input)
LLM02: Insecure Output Handling (downstream exploitation of LLM output)
LLM03: Training Data Poisoning (manipulation of training or fine-tuning data)
LLM05: Supply Chain Vulnerabilities (compromised components, plugins, or models)
LLM06: Sensitive Information Disclosure (unauthorized access to PII or proprietary data)
LLM08: Excessive Agency (unauthorized actions via agentic capabilities)
LLM09: Overreliance (compounding errors or hallucinations in critical decisions)
Beyond OWASP, red teaming must address bias, toxicity, jailbreaks, multimodal attacks, multi-turn escalation, reward hacking, deceptive alignment, and sandbagging. The redteams.ai OWASP LLM Top 10 deep dive provides detailed mapping of attack techniques to each category. The DataVLab Practitioner’s Guide offers additional vulnerability taxonomies for enterprise assessment.
Phase 1: Scope and Threat Modeling
Red teaming without clear scope produces noise, not signal. Phase 1 defines boundaries, success criteria, and the adversary profile before any testing begins.
Define the System Under Test. Is this a base model, fine-tuned variant, RAG pipeline, or end-user application? Each requires different testing approaches. Base models need alignment and refusal testing. RAG systems need indirect injection and groundedness validation. Agentic workflows need tool misuse and excessive agency probes.
Identify Deployment Context. Customer-facing chatbots face different threats than internal assistants. Systems with tool access (email, APIs, databases) carry higher stakes than isolated text generators. Embedded copilots in productivity software face distinct attack surfaces compared to standalone applications.
Classify by Risk Tier. The EU AI Act and NIST AI RMF provide frameworks for distinguishing low-risk from high-risk systems. High-risk categories include biometric identification, critical infrastructure, education, employment, and law enforcement. These require more rigorous testing documentation and mitigation evidence.
Map Entry Points. Document every vector where adversarial input could enter: user input fields, API endpoints, connected data sources, plugin interfaces, and multimodal channels (image upload, voice input). Each entry point represents a distinct test surface.
Define Success Criteria. What constitutes a finding? A successful jailbreak? Extraction of system prompts? Unauthorized PII disclosure? Tool misuse? Clear criteria prevent scope creep and ensure reproducible results.
Assemble Diverse Red Teamers. Effective red teaming requires technical expertise, domain knowledge, and diverse perspectives. The Microsoft planning guide for LLM red teaming recommends teams include backgrounds across security research, ML engineering, linguistics, and the specific domain (healthcare, finance, legal) of the system under test. The ArXiv paper on operationalizing threat models provides frameworks for structuring these teams. CSET Georgetown’s research on AI red-teaming design offers additional guidance on threat model construction.
Phase 2: Reconnaissance
Reconnaissance gathers intelligence before attack execution. Skipping this phase wastes effort on irrelevant attack vectors.
Extract or Infer the System Prompt. Many applications leak their system prompt through carefully crafted queries. Understanding the prompt structure, safety instructions, and tool definitions guides attack design. The Coralogix comprehensive guide covers prompt extraction techniques.
Fingerprint the Model. Probe capabilities, knowledge cutoff, behavioral quirks, and refusal patterns. Is this GPT-4, Claude, Llama, or a fine-tuned variant? Each has distinct vulnerabilities. Test refusal consistency across request phrasings.
Catalog Integrations. Document connected tools, data sources, API endpoints, memory and context persistence, and RAG knowledge bases. Each integration expands the attack surface. The Confident AI step-by-step guide provides reconnaissance checklists.
Test the Production UI, Not Just the API. Real-world behavior often differs from API-level assumptions. UI elements, rate limiting, and session management introduce constraints and vulnerabilities absent from raw API access.
Identify Trust Boundaries. Map where user input becomes system instruction, where external data flows enter the pipeline, and where privilege escalation could occur. These boundaries guide injection and hijacking attempts.
Phase 3: Attack Generation
Phase 3 executes structured attacks across the vulnerability categories identified in Phase 1. We recommend a two-mode approach: manual attack design for novel scenarios, and automated generation for broad coverage.
Direct Prompt Injection. Attempt to override system instructions through user prompts. Test roleplay, privilege escalation, and instruction override patterns. For jailbreak-specific techniques and adversarial dataset construction methodology, see our LLM jailbreak testing guide.
Indirect Prompt Injection. Embed malicious instructions in data the model processes: documents for RAG, emails for summarization, web pages for browsing. Test whether the model executes injected instructions when retrieving external content. Our prompt injection prevention guide covers this attack vector in depth.
Roleplay and Persona Hijacking. Convince the model to adopt personas that bypass safety training. Test fictional scenarios, historical figures, and hypothetical framing.
Encoding Attacks. Obfuscate malicious content via Base64, ROT13, Unicode variations, and low-resource languages. Determine whether input filters catch encoding that the model decodes.
Multi-Turn and Crescendo Escalation. Build rapport across conversation turns, gradually shifting toward harmful content. Research shows multi-turn attacks achieve significantly higher success rates than single-turn probes.
PII Extraction Probes. Attempt to extract training data, user data, or system prompts through repeated querying, completion attacks, and context window exploitation.
Bias and Toxicity Probes. Test for discriminatory outputs, harmful stereotypes, and toxic generation across demographic categories.
Tool Misuse for Agentic Systems. If the LLM has tool access, attempt unauthorized actions: sending emails to attacker-controlled addresses, exfiltrating data via API calls, or escalating privileges within connected systems.
Tool Selection for Attack Generation:
Garak (NVIDIA): 120-plus probe modules covering hallucination, data leakage, and prompt injection. Nmap-style reporting. Best for model-level vulnerability scanning.
Promptfoo: YAML-based configuration, OWASP LLM Top 10 mapping, CI/CD integration. Best for application security testing in deployment pipelines.
PyRIT (Microsoft): Targets, converters, scorers, and orchestrators for complex multi-turn campaigns. Burp Suite integration. Best for structured adversarial automation.
DeepTeam: Attack generation plus vulnerability taxonomy, simple Python API. Best for automated at-scale testing.
Giskard: 40-plus probes with adaptive attack strategy. Best for dynamic multi-turn attacks on agentic systems.
Phase 4: Evaluation and Scoring
Raw attack results require structured evaluation to produce actionable intelligence. Phase 4 establishes severity, priority, and remediation pathways.
Three Scoring Approaches:
Binary True/False: For bulk automated runs. Did the attack succeed? Simple, scalable, but misses nuance.
Likert Scale: For nuanced behavioral probes. Rate harmfulness or policy violation from 1-5. Captures gradations in model responses.
LLM-as-Judge: For scalable evaluation against safety criteria. Use a separate model to evaluate outputs against defined criteria. Calibrate against human judgment to prevent drift.
Severity Classification:
Critical: Immediate harm possible, regulatory liability, data breach, or system compromise. Blocks release.
High: Significant harm possible under specific conditions. Requires mitigation before deployment.
Medium: Moderate impact, limited exploitability. Track for future remediation.
Low: Minor issues, edge cases. Document for awareness.
OWASP Mapping. Categorize each finding by OWASP LLM Top 10 category. This mapping enables trend analysis and compliance reporting.
Prioritization. Critical and high findings block release pipelines. Medium findings require mitigation plans. Low findings require monitoring.
The Promptfoo LLM red teaming guide covers evaluation methodology. DeepTeam documentation provides scoring frameworks. The apxml.com reporting guide offers detailed severity classification criteria.
Phase 5: Reporting, Remediation, and Re-Testing
Phase 5 delivers lasting security value through structured documentation and verified fixes.
Report Structure:
Executive Summary: Business impact, critical findings, and recommended actions for leadership.
Attack Surface Map: Visual documentation of entry points, trust boundaries, and tested components.
Findings Log: Each vulnerability with reproduction steps, OWASP category, severity, and evidence.
Mitigation Recommendations: Specific technical and process controls to address each finding.
Re-Test Outcomes: Verification that mitigations resolved vulnerabilities without introducing regressions.
Residual Risk Acceptance: Documented decisions to accept remaining risk with justification.
Remediation Collaboration. Work directly with ML engineers and product teams. Assign owners and timelines. Escalate critical findings to block deployment. The apxml.com guide on working with development teams provides collaboration frameworks.
Verify Fixes with Targeted Re-Tests. Do not assume mitigation worked. Re-run the exact attack that succeeded. Test similar variants. Confirm no regressions in benign functionality.
EU AI Act Compliance Documentation. For high-risk systems, maintain:
Red teaming methodology and scope
Adversarial input categories tested
Mitigation evidence for each finding
Residual risk acceptance with justification
Continuous monitoring plan
The CodeSecure AI red teaming methodology provides documentation templates. SQUR.ai EU AI Act guidance covers compliance-specific requirements.
CI/CD Integration: Making Red Teaming Continuous
One-off assessments catch vulnerabilities at release. CI/CD integration catches them at every commit.
One-Off versus Continuous. Major releases need comprehensive assessment. Every model update, prompt change, or plugin addition needs automated regression testing. The bestllmscanners.com guide compares implementation approaches.
Tool Integration Patterns:
Promptfoo: YAML-based configuration integrates directly into existing CI/CD pipelines. Define probes in version control, run on every build, block on critical findings.
Garak: Standalone pipeline scripts for model scanning. Generate reports as build artifacts.
PyRIT: Python API integration with existing testing toolchains. Orchestrate complex campaigns as part of deployment workflows.
Three-Loop Model. Implementation patterns recommend: probe on every commit, classify severity automatically, triage findings to security and engineering teams.
Regression Prevention. Automated probe suites block deployments that introduce new critical findings. Compare results against baselines. Escalate deviations for human review.
Tool Selection Guide: Choosing Your Red Teaming Stack
| Tool | Strength | Best For |
|---|---|---|
| Garak | 120-plus probes, nmap-style reports | Model-level vulnerability scanning |
| Promptfoo | YAML config, OWASP mapping, CI/CD ready | Application security testing in pipelines |
| PyRIT | Multi-turn orchestration, Burp integration | Complex adversarial campaigns |
| DeepTeam | Automated taxonomy, simple API | At-scale automated testing |
| Giskard | Adaptive strategy, 40-plus probes | Dynamic multi-turn agent attacks |
| HiddenLayer/Lakera | Vendor support, compliance reporting | Enterprise managed testing |
Start with Garak for model scanning, Promptfoo for CI/CD, PyRIT for multi-turn campaigns, and DeepTeam for automated taxonomy coverage.
The NetGuardia tool comparison provides detailed feature matrices. BeyondScale’s comparison covers performance benchmarks. AppSecSanta’s analysis compares specific tool capabilities.
Compliance Integration: Red Teaming as Regulatory Evidence
Red teaming documentation satisfies regulatory requirements across jurisdictions.
EU AI Act Requirements:
Article 55: Documented adversarial testing for GPAI models with systemic risk (10^25 FLOPs training compute)
Article 15: Robustness and cybersecurity testing for high-risk AI systems (Annex III categories)
Article 9: Risk management system including testing procedures throughout lifecycle
Penalties: Up to 35 million euros or 7% global turnover for non-compliance
NIST AI RMF:
MEASURE Function: Systematic evaluation including adversarial probing
AI 600-1: Jailbreak and prompt injection resilience metrics for GenAI
Documentation Checklist:
Red teaming methodology and scope
Adversarial input categories attempted
Mitigation evidence for findings
Residual risk acceptance with justification
Continuous monitoring plan
The Claru EU AI Act red teaming guide provides compliance-specific methodologies. TechIntelix coverage explains the 2026 compliance landscape. AI Policy Desk requirements analysis covers jurisdictional variations.
The Role of Human Expertise in Red Teaming
Automated tools provide breadth. Human red teamers provide depth, novel attack discovery, and domain expertise.
Where Humans Add Value:
Creative attack design that automated tools miss
Domain-specific probing (healthcare, finance, legal) requiring subject matter expertise
Evaluation of edge cases where automated scoring fails
Continuous adaptation as attack techniques evolve
Annotation and Evaluation at Scale. Red teaming generates massive datasets: attack prompts, model responses, severity classifications. Human judgment ensures edge cases receive appropriate attention. Conservative annotation, where safety labels err toward caution, produces the most reliable training data for defensive systems.
Continuous Red Teaming Programs. Sustained adversarial testing requires dedicated expertise, not one-off consultant engagements. Teams need time to learn system behavior, develop novel attacks, and track evolving threats.
LXT provides human annotation and red teaming support across more than 200 languages. We combine automated tool execution with expert human evaluation, delivering the labeled datasets and vulnerability assessments that power enterprise AI security programs. Our ISO 27001 certified facilities and PCI DSS compliant infrastructure support the most stringent compliance requirements.
For adversarial dataset construction methodology, see our LLM jailbreak testing guide. For defensive hardening strategies, see our AI jailbreaks prevention analysis. Contact us to discuss your red teaming and annotation requirements.




