EU SME AI Data Assess

Before your AI feature ships to EU customers, try deleting one customer from everything it reads.
That test turns up copies nobody put on the compliance slide. The support tickets sit in the warehouse, chunked copies sit in a vector store, a sample went into a fine-tuning file, and some prompts live in the vendor’s logs.
Here’s the check I run with teams before launch, yes or no:
- Can you list every place the feature’s input data gets copied to?
- Can you remove one person from all of those within a month?
- Is there a named owner for each source the model reads?
- Can you say which version of the data the feature saw on a given date?
- Have you written down whether the use case lands in an AI Act high-risk category? Most SME features won’t, and the reasoning is worth keeping.
If either 1 or 2 is a no, the data work comes before the launch date. In the teams where I’ve done this, that’s meant a few weeks of unglamorous plumbing, and much cheaper than doing it after a complaint.
I’m not a lawyer, and your DPO still signs off. The copies and the owners are architecture, though, and that part you can fix yourself.
The data-layer groundwork an AI feature needs is here: https://thomasnys.com/ai-data-architecture/
Fractional Data Architect helping startups and scaleups build data platforms that scale.
More about Thomas Nys →