Field Notes / Compliance
EU AI Act red teaming: what you must do before 2 August 2026
The EU AI Act makes adversarial testing a legal requirement for high-risk AI systems. Here is what the regulation asks for, whether your systems are caught by it, and the work you can still finish before the deadline.
01 The deadline that matters
The EU AI Act is the first comprehensive AI law anywhere, and most teams have read the headlines without reading the clock. Enforcement does not arrive in one moment. It lands in waves, and one of those waves carries a testing obligation that a firewall and a passing audit will not satisfy.
The Act entered into force on 1 August 2024 EU 2024/1689 . The ban on prohibited AI practices took effect on 2 February 2025. Rules for general-purpose AI models, the systems behind tools like GPT-4, Claude and Gemini, applied from 2 August 2025. The wave that catches most businesses is the next one. On 2 August 2026, the full Chapter III obligations for high-risk AI systems become enforceable, and mandatory testing sits inside them. Remaining provisions, including certain Annex I systems, follow on 2 August 2027.
The reason the 2026 date deserves your attention is the penalty attached to it. Breaching the high-risk obligations, testing included, carries a fine of up to €15 million or 3% of global annual turnover, whichever is higher. Prohibited practices are punished more harshly still, at €35 million or 7%. Supplying incorrect information to authorities costs up to €7.5 million or 1%. For a small business the lower figure applies. For an enterprise turning over more than half a billion euros, the percentage is what hurts.
02 Whether your AI counts as high-risk
The Act sorts AI into four risk tiers. High-risk systems, defined in Article 6 and Annex III, carry the strictest duties, and mandatory adversarial testing is one of them. Your system is likely high-risk if it drives decisions in any of these areas: employment and recruitment, such as CV screening or performance monitoring; credit and financial assessment, such as scoring, insurance pricing or loan approval; critical infrastructure, such as energy, water, transport or telecommunications; education, such as admissions or student assessment; law enforcement, such as risk assessment or border control; access to essential services, such as healthcare or social benefits; and biometric identification, such as facial recognition or behavioural categorisation.
There is a second route into the high-risk tier that catches teams off guard. An AI system that acts as a safety component of a product already covered by EU harmonised legislation, think medical devices, vehicles, machinery or aviation, qualifies automatically. You do not choose the label. The use case chooses it for you.
03 What Article 9 actually requires
Article 9 sets up a risk-management system that must run across the whole life of the AI system, not a one-off form filled in before launch. Four parts of it map directly onto red teaming. Article 9(2)(a) requires you to identify known and foreseeable risks when the system is used as intended and under conditions of reasonably foreseeable misuse. That phrase is the hinge. Adversarial attacks against AI are now well documented, so they are foreseeable by definition, and a risk register that ignores them is incomplete on its face.
What does foreseeable misuse look like in practice? System-prompt extraction succeeded in 89% of the AI assessments our operators ran. RAG pipeline weaknesses showed up in 82% of the retrieval deployments we tested. Prompt injection affects roughly 73% of production AI applications OWASP LLM 2025 . None of these are exotic. They are the baseline threat model, and Article 9(2)(a) expects them on your register.
Article 9(2)(b) then asks you to estimate and evaluate those risks with both quantitative and qualitative methods. A red team assessment produces exactly that: specific findings, severity scores, exploitability ratings and impact analysis. Article 9(6) requires appropriate testing procedures at stages of development and before the system reaches the market, measured against clearly defined metrics, run under real-world conditions, and covering reasonably foreseeable misuse. Article 9(7) closes the loop by insisting that whatever risk remains after mitigation is documented and communicated to deployers. A red teaming report is not merely advisable at that point. It is the legal artefact that shows the loop was closed.
Regulatory guidance from the European AI Office is clear on what counts as adequate. Testing has to be conducted or validated independently, because internal testing alone will not satisfy a high-risk obligation. It has to cover the full attack surface, meaning the APIs, data pipelines, access controls and integration points, not the model in isolation. It has to use current attack methods drawn from sources like the OWASP Top 10 for LLMs and the NIST AI Risk Management Framework NIST AI RMF , rather than a static checklist. And the results have to feed back into the risk-management system as concrete remediation actions.
04 Closing the gap before August
The uncomfortable arithmetic is in the timeline. A red team assessment for a standard enterprise AI deployment runs two to four weeks. Remediation adds another two to six weeks depending on severity. Re-testing takes one to two weeks more. That is a minimum of five to twelve weeks from first engagement to demonstrable compliance, which means any organisation that has not started is already inside the risk window for the deadline.
Supply makes the arithmetic worse. Demand for AI red teaming is climbing fast. The market was valued at around $1.3 billion in 2025 and is growing at roughly 30.5% a year Market.us 2025 . As August nears, qualified assessors get harder to book, not easier. The teams that certify cleanly will be the ones that scheduled the work early.
The good news is that this is a finite, mappable programme, and our AI security assessment was built to produce the evidence Article 9 names. Threat modelling answers Article 9(2)(a). Input-validation and output testing answer Article 9(6). Access-control and data-pipeline review answer Article 9(2)(b) and foreseeable misuse. Integration testing answers the real-world-conditions clause. Compliance mapping answers the residual-risk duty in Article 9(7). Each step leaves a documented artefact, and together they form the audit trail a regulator will ask to see. Once the findings are closed, RTP Robin lets your team re-run the same attacks as the system changes, so the picture stays current between formal assessments rather than ageing the moment the report is signed.
- Catalogue every AI system in use and classify each against Annex III
- Commission independent adversarial testing for each high-risk system
- Remediate critical and high-severity findings, then re-test to confirm closure
- Document residual risk under Article 9(7) and complete the technical file
- Set a recurring re-test cycle so compliance holds after the system changes
References
Sources
- European Parliament and Council. Regulation (EU) 2024/1689 (Artificial Intelligence Act). Official Journal of the European Union, 2024. artificialintelligenceact.eu
- European Commission. European AI Office. Directorate-General for Communications Networks, Content and Technology, 2024. digital-strategy.ec.europa.eu
- OWASP. Top 10 for Large Language Model Applications, 2025 Edition. OWASP Foundation, 2025. owasp.org
- NIST. Artificial Intelligence Risk Management Framework (AI RMF 1.0). National Institute of Standards and Technology, 2023. nist.gov
- Market.us. AI Red Teaming Services Market: Global Forecast. Market.us, 2025. market.us