Cost-Aware Data Engineering - FinOps for Data

Your data team doesn’t know what their pipelines cost. That’s the first problem.
I asked a data engineering lead last month: “What does your most expensive pipeline cost per run?” Blank stare. Nobody had ever looked.
Turns out, one pipeline was doing a full table refresh every night while only a small slice of the data changed. Switching to incremental was a small change with a visible effect on the bill.
This is the FinOps for Data gap. For years, the data world operated on “spend what you need, we’ll sort it out later.” Later never came. Now budgets are tightening and CFOs want to know what drives the Databricks bill.
Three places I always look first:
- Full-refresh jobs that could be incremental
- Dev/test clusters left running 24/7 because nobody set auto-suspend
- Inappropriate storage tiers that treat cold data like hot data
Most teams find avoidable spend without removing features. The first gains usually come from measuring what is running, who owns it, and whether anyone still needs it.
Nobody enjoys the first audit. But the bill after makes it worth it.
When’s the last time someone audited your data platform’s cloud costs?
Fractional Data Architect helping startups and scaleups build data platforms that scale.
More about Thomas Nys →