EVALUATE · You build AI into your products
Due diligence of your AI models
An internal review of the AI models you have developed: what risks they carry, whether those risks are properly identified and whether they are properly evaluated.
AI Act Art. 9 · prEN 18228 · ISO/IEC 23894
It is a starting point: the recipe is adapted to your company.

Evaluate
Verify and validate its safety and impact: what can go wrong, who it affects and whether the controls work.
You get
For each model, in writing: the missing risks, the evaluations that do not hold up and what is worth strengthening.
What sets it apart
We do not stop at the documentation: what is claimed about the model is checked with technical tests.
How we work
The models and their purpose
GOVERN · first step
Which models you have developed, what they are used for and whom their outputs affect. That tells us where to start.
Are the risks properly identified?
EVALUATE · with your team
We compare the risks you have recorded with those produced by the use and by the model itself. Missing ones are identified in a session with the canvas.
The AI risk canvas →Are they properly evaluated?
EVALUATE · with their evidence
We review the criteria, the estimated probability and severity and the evaluation of each risk: whether they hold up with evidence and follow the standard's method.
Test what is claimed
PROVE · with your data
Where the evaluation says something about the model, we check it with your data and metrics: performance, robustness and differences between groups.
The method, step by step
- You tell us who is affected; we bring the method, the measurement and the tests.
- You approve the test thresholds.
- Someone in your organisation accepts the result.
- It protects your team’s work before whoever examines you. It is not an audit of the team.
Report and decision
GOVERN · with whoever decides in your organisation
We hand over the report for each model and go through it with whoever decides in your organisation. The review ends there; what comes next is your call.
What we leave running
The reviewed and expanded risk register
With the missing risks, identified with your team in a canvas session.
Model tests
Performance, robustness and differences between groups, measured with your data and repeatable.
Criteria your team can apply
How to identify and evaluate risks with the method of the standard, for the next version.
What changes for your team
Stays the same
- Your models and your data
- Your team's way of working
Changes
- Risks start from the model's real use
- Every evaluation is backed by evidence
- What is claimed about the model is tested
Other starting points
You build AI into your products