Security

AI red teaming

Before launch, test your AI application the way an attacker would: prompt injection, data leakage, actions beyond its permissions, checked item by item against the OWASP LLM Top 10.

01

Who it’s for

  • AI customer service, chatbots or internal AI assistants that are live or about to launch
  • AI applications connected to internal data, or able to take actions such as sending email or placing orders
  • You need to explain your AI security measures to customers, auditors or regulators
  • After a model or feature update, you’re not sure your existing safeguards still work

02

What we do

Scoping

We map the application’s architecture, the data and tools the AI can reach, and each type of user’s permissions.

Prompt injection testing

Direct injection and indirect injection (instructions hidden in documents, web pages or email), to see whether the AI’s behavior can be changed.

Data leakage and privilege abuse

Whether the system prompt and sensitive data can be extracted, and whether the AI can be led into actions beyond its permissions.

Output handling

Whether the AI’s output, when used by other systems, can cause injection or unintended execution.

Report and retest

Each finding comes with a severity, reproduction steps and a recommended fix; we retest after fixes.

03

How it works

1. Scope and rules

Agree on the test scope, environment, time window and stop conditions.

2. Testing

Test item by item against the OWASP LLM Top 10, recording every attempt.

3. Report

Explain findings, impact and recommended fixes, and go through them one by one with your developers.

4. Retest

Test again after fixes to confirm the issues are resolved.

04

Deliverables

  • Test scope and methodology
  • Findings list (severity, reproduction steps, recommended fixes)
  • Retest report

05

Engagement cycle

Each cycle runs three months, six months or a year, depending on scope. At the end of each cycle we sit down with you and compare the results against the goals set at the start, then decide what the next cycle should cover, or whether to stop there.

Every cycle: audit → design → implement → check against the goals, then decide what's next

06

Pricing

Each engagement is estimated on its own: how many systems and how much data are involved, the people and time needed, and how long the cycle runs. Talk to us first and we’ll give you a number based on the actual scope, rather than quoting a price and then fitting the scope to it.

07

Common questions

What is AI red teaming?

Simulating an attacker to try to make an AI application do what it shouldn’t: leak data, ignore its original instructions, or take actions beyond its permissions. The goal is to find the problems before a real attacker does.

How is AI red teaming different from a regular penetration test?

A regular penetration test targets weaknesses in networks, servers and web applications. AI red teaming targets risks specific to language models, such as prompt injection, where text content changes the model’s behavior. The two complement each other.

Will testing affect our production environment?

We recommend testing in a test environment. If production must be tested, we agree on the scope, time window and stop conditions in advance.

Are we safe once the fixes are in?

There is currently no one-time cure for prompt injection. The aim is to reduce risk and limit the permissions and data the AI can use. After a model or feature update, we recommend testing again.

08

Further reading

Tell us where you are

Email us about where you’re stuck, and we’ll reply with what could work and the next step.