Notes on dataarchitecture.
Short essays on architecture, cost, hiring and AI: one cartoon, one idea, most days of the week. No tutorials, no listicles.
Orchestration Consolidation Matrix
Count the places a job can be scheduled in your data stack. Include the cron on that one VM.
Orchestration Consolidation Matrix
Count the places a job can be scheduled in your data stack. Include the cron on that one VM.
Read →BigQuery Scan Cost Surprises
LIMIT 10 doesn't reduce what BigQuery scans on an unclustered table. A lot of teams learn that from the invoice.
Read →dbt Coalesce Metadata Shift
The description fields in your dbt project used to be read by nobody. Now an agent reads them.
Read →Governance Before AI Act Fines
The provenance record behind an AI feature fits in a YAML file next to the pipeline code.
Read →Process Before People Fix
Before you replace the data engineer who seems slow, count how many people can hand them work.
Read →Stack Impact Analysis
Before any stack change I ask how many consumers it touches and whether we can undo it by Friday.
Read →Rust Is Quietly Becoming the Language of Data Infra
The fastest tools in your data stack are written in a language nobody on your team writes.
Read →Reg Dual Compliance Map
GDPR and the AI Act ask their questions at the same moment: when a dataset gets reused for a model.
Read →Pipeline Novelty Implementation
Novel data projects tend to die between the team that built the prototype and the team stuck running it.
Read →EU SME AI Data Assess
Before your AI feature ships to EU customers, try deleting one customer from everything it reads.
Read →Want expert eyes on your data architecture?
No pitch. An honest conversation about whether I can help, and what shape it would take if I can.