Why every app needs manual red-teaming
With today’s powerful foundation models and pre-existing fine-tuned systems, it is tempting to believe that red-teaming is no longer necessary for every deployment. Aren’t language models now essentially plug-and-play? Connect them to your company’s knowledge base, configure a few safeguards, and you’re ready for production.
No.
Every LLM application has its own risk surface, shaped by its domain, its users, its integrations and the data it can access. That risk surface must be actively tested for safety, privacy, and ethical failure modes.
The real risk surface
Consider a customer service agent trained on a company’s internal documentation. Best practice dictates that personally identifiable information (PII) is scrubbed before those documents are used for training or retrieval. But internal knowledge bases are rarely static. They are updated constantly, new documents are added, policies are changed, data pipelines evolve, and each update introduces new opportunities for error.
Red-teaming can expose where data cleaning has failed, where sensitive information remains embedded, or where the model has learned patterns that allow it to reconstruct or infer confidential data unintentionally. In multi-turn conversations subtle prompting strategies can elicit information that was never meant to be revealed.
The practical reality is that risk does not live only in the base model, but also emerges from how the model is connected, updated and used.
Need human red-teamers who find what automation misses?
LXT builds expert red-teaming and safety-evaluation programmes that probe your model’s real risk surface across languages, domains, and multi-turn attacks.
The value of red-teaming
Red-teaming is often misunderstood as a tool that “proves” a model is safe or has no exploitable “gaps”. This is not the case.
Rather, red-teaming is a structured method for confirming that avoidable harms are in fact avoided, and for helping engineers map how and where the system fails. It helps organisations discover and understand their exposure across a range of harms, including privacy, safety, bias, misinformation and misuse, before those failures occur in production.
It can be useful to think of red-teaming as a discovery process, not a certification.
Industry best practice
This approach is consistent with guidance from Microsoft, which recommends conducting manual red-teaming for each LLM application, even before running automated test suites. Importantly, Microsoft also recommend red-teaming the production model through the same user interface that end users will interact with, in order to replicate real-world conditions as closely as possible.
Testing in isolation is not enough. Systems must be stress-tested as they are actually deployed.
What should you test for?
Harmful outputs can be categorised in several ways, and the following is just one type of high-level taxonomy:
- by the type of harm caused: biased language, providing information on how to conduct illegal activities, by agreeing with statements about self-harm, etc.,
- by the method used to generate it: prompt injection, attacking implicit assumptions, use of different languages, etc., or
- by the source of vulnerability: harmful training data, low-quality fine-tuning data, trajectory – that is, harms that appear after a single prompt, or after extended turns (see A Red Teaming Roadmap Towards System-Level Safety), etc.
As a starting point, many organisations draw on benchmark frameworks such as those developed by MLCommons, whose AILuminate benchmark defines twelve categories of harmful single-turn outputs across three classes: physical hazards, non-physical hazards, and contextual hazards.
However, generic benchmarks are only a baseline. Application-specific testing must go further, and can include multi-turn attacks.
Common categories to include in a tailored harm inventory are:
- Providing instructions to create dangerous or illegal substances
- Advising users on how to evade law enforcement or bypass safeguards
- Generating financial, medical, or legal misinformation
- Implicitly reinforcing harmful biases related to gender, sexuality, age, or ethnicity
- Multi-turn manipulation strategies designed to circumvent guardrails
- Confidential data extraction or inference attacks
This inventory should evolve continuously as new risks are identified.
Automatic vs manual red-teaming
Automated red-teaming tools are valuable – they can scale testing, probe known vulnerability classes, and provide repeatable stress scenarios.
But automation has limits.
Generative models tend to reproduce patterns found in their training data. They are effective at exploring known attack types, yet less capable of discovering genuinely novel exploit paths, especially those that arise from specific integrations, business logic, or proprietary data structures.
In addition, language models and multi-agent pipelines created by them, tend to favour their own texts, finding fewer problems with their own reasoning and code output than in text of non-related models’ output. They also preference volume over reliable sources.
Human red-teamers, by contrast, bring contextual reasoning and creativity over longer turn sequences. They can test edge cases, multi-step interactions and subtle contextual manipulations that automated systems often miss, because automated systems are trained on current data, not on “future, possible” data.
Another way to look at it is that, while automated testing measures coverage, manual red-teaming uncovers surprises.
Red-teaming is risk-mapping
Red-teaming should not be treated as a compliance checkbox performed once before launch.
It is a method for understanding your system’s real risk surface. It is a way to prioritise mitigation efforts based on actual exposure. And because knowledge bases evolve, integrations change, and models are updated, it must be ongoing.
Red-teaming is not about proving your model is safe, it is about understanding how it fails before your users do.



