Fractional Data Architect
Book a discovery call →

Semantic For Unstructured First

Semantic For Unstructured First
Semantic For Unstructured First

Most AI data problems I see start in a column called payload, long before any embedding.

The AI plans I’ve reviewed this year start with embeddings. The data they want to feed in usually sits somewhere duller: a payload column in Postgres, filled by three versions of the same webhook.

Field names changed in 2023. Half the rows store customer_id as a string, half as an integer. Nobody wrote down which event types exist.

Put that in a vector store and you get confident search over things nobody can define. The model finds the chunk. It still can’t tell you whether “status: 3” means paid or refunded.

What I’d do first is boring. I’d pull and count the distinct keys, name the event types, then pick one shape per type and write it down as a contract, even if it’s a markdown file for now.

I’m still unsure how much of this an LLM can infer for you. For the obvious keys, quite a lot. For the one field finance cares about, I wouldn’t trust it yet.

How many distinct keys are hiding in your biggest JSON column?

Written by Thomas Nys

Fractional Data Architect helping startups and scaleups build data platforms that scale.

More about Thomas Nys →

Recognise the problem? Let's talk about it.