An agent that tells a small business where it stands under the EU AI Act
AI Governance Co-Pilot. January 2026.
The problem
A UK small business buying an HR screening tool, or a CRM with lead scoring built in, has a question with an expensive answer: is this high-risk under the EU AI Act? Classification is not a lookup, because the same CV screening tool can be high-risk or exempt depending on how it is used, and most of these businesses became deployers of AI, in the Act's sense, through software they already had. Getting a lawyer to answer the question costs about six hours and £2,400 per system on the rates I modelled.
What I built
A web tool where the owner describes their AI system through a five-step intake wizard (industry, their role, the use case, who it affects, review) and receives a classification against the Act: prohibited under Article 5, high-risk under Article 6 and Annex III, or not high-risk. The answer arrives as a card carrying the category, the article it rests on and a confidence score, with a tailored markdown checklist of obligations and a legal disclaimer specified on every classification. A line from the test log shows the shape: high-risk, Annex III category 4(a), employment, recruitment screening, confidence 0.85. A conversation costs about £0.04 in model calls.
I started with a PRD, then designed, built and evaluated the system end to end.
Built with
| Model | Claude 3.5 Sonnet, called through OpenRouter |
| Orchestration | n8n, self-hosted, running the agent workflow and the webhook |
| Retrieval | Supabase with pgvector, holding the text of the EU AI Act |
| Frontend | React (Vite) and Tailwind, deployed on Vercel |
| Evaluation | LangSmith dataset, three custom Python evaluators, a 29-case functional suite |
| Escalation | Slack channel for cases the agent will not decide |
| Specification | Three-part PRD: business case, agent specification, system visualisation |
The behaviour spec
What the agent may decide. The agent leads the questioning while the user supplies the facts, and it reasons to a classification which the user reviews. Anything beyond that, acting on the classification or acting without review, is written out of scope, with the reason recorded: compliance classification depends on facts only the user can verify.
The workflow, with a stop rule per stage.
| Stage | What the agent does | Stops when |
|---|---|---|
| 1 Information gathering | Asks about the system, its purpose, who it affects, and whether a human reviews decisions | Ten questions asked, or the user asks for a classification |
| 2 Classification | Reasons against Article 6 and Annex III, returns category, article and confidence | Classification complete, or confidence too low, which triggers escalation (thresholds below) |
| 3 Role clarification | Determines provider or deployer, which sets the obligations | Role determined, or confidence too low to determine it |
| 4 Output | Delivers the classification card and the checklist | Delivered and acknowledged |
What it may use. Four tools:
- Search the Act
- Generate the obligations checklist
- Determine the provider or deployer role
- Escalate to a human
Article 6 and Annex III themselves, about 2,000 tokens, sit in the system prompt so the core logic never depends on a retrieval succeeding, and the vector store supplies the supporting text and the citation for each answer.
What it must not do. Six anti-goals, written into the specification as constraints on the agent:
- Must not provide legal advice
- Must not guarantee compliance
- Must not store sensitive business data
- Must not classify without sufficient information: better to ask more questions than guess
- Must not override the user's stated facts: if the user says a human reviews every decision, the agent accepts it
- Must not be overconfident on edge cases: express uncertainty, recommend escalation
Where confidence goes.
| Confidence | What the user gets |
|---|---|
| Above 0.7 | The classification and the full obligations checklist |
| 0.5 to 0.7 | The classification with caveats and a recommendation to have it verified |
| Below 0.5 | No classification. The case is escalated and the user is directed to legal counsel. Two consecutive attempts below 0.5 stop the conversation |
Escalations post to a Slack channel so I could read every case the agent would not decide, and the user is pointed to qualified counsel rather than to me.
How I would know it had drifted. One rule in the health metrics: if accuracy improves while the escalation rate on boundary cases drops below 10%, investigate, because the agent may have become over-confident. Accuracy going up is the number everyone watches.
How it was evaluated
A ten-case LangSmith dataset, run from a script against the n8n webhook, with three custom Python evaluators:
- Was the classification correct
- Was the right category cited
- Was the disclaimer present
Between prompt version three and version seven, classification accuracy went from 30% to 80% and correct category citation from 30% to 90%. The disclaimer evaluator closed at 60%: the spec requires a disclaimer on every classification, and the deployed agent was not meeting it.
Part of the gain came from fixing the evaluators themselves. A JSON parsing error, a case-sensitive regex and a guardrail check which passed everything were all scoring the agent wrong, so some of what had looked like agent failure was measurement failure, and I only found it by reading the transcripts behind the scores.
A separate functional suite of 29 cases across six categories (positive, negative, boundary, adversarial, multi-turn and escalation), including five adversarial cases written to pull the agent past its boundaries, passed 29 of 29 against the behaviour each case specified.
Where it broke, and what each failure taught
The open chat box. The first interface was an open chat box, and users stalled trying to describe their AI system in free text. I replaced it with the five-step wizard, which addressed the riskiest usability assumption in the PRD directly.
The agent which would never answer. The early evaluation runs scored 30% on classification and the transcripts showed why: the agent kept asking follow-up questions and never committed, because the structured intake was arriving as new information and triggering more gathering. The fix was a dual-mode system prompt, gather first, then classify.
One question before committing. With the intake structured, the agent could have classified on the form alone. I kept one instruction in the prompt: ask one follow-up question about the most important gap before classifying. In the final demo the agent paused to ask whether a human reviews the decisions before it committed.
The prohibited-case bug. A prohibited classification returns an empty response body from the webhook, HTTP 200 with no JSON, so the agent classifies an Article 5 case correctly and the user sees nothing. It is in the release notes as a known issue and it is still open.
Status
This was a build project: I took it to a working prototype and did not take it live or make it available to the public. The PRD is available on request.