Fractional Data Architect
Book a discovery call →

Sizing Your Orchestrator To Your Team, Not Your Ambition

Sizing Your Orchestrator To Your Team, Not Your Ambition
Sizing Your Orchestrator To Your Team, Not Your Ambition

Three engineers maintaining Kubernetes Airflow for 20 DAGs is a sizing mistake, and a common one.

Teams I talk to tend to land in one of two places. That one, or the opposite: cron and a prayer until something breaks in front of the board.

Both usually happen before anyone has counted the jobs.

So count what runs today, look at how that number grew over the last year, and project it 18 months out. Then use rough bands:

  • Under ~20 jobs, one team: your ingestion tool’s scheduler plus dbt’s own runner is usually enough.
  • 20 to ~150, with dependencies across teams: a managed orchestrator. Dagster, Prefect or managed Airflow. Pick the one your team can read.
  • Beyond that, or with hard SLAs: self-hosting might earn its keep, if someone owns it.

These bands are my own rule of thumb. The question under all of them is who maintains the orchestrator itself. If the answer is “the same person who writes the pipelines”, go managed.

I’ve over-built this before, for scale that never came. We spent more time upgrading the orchestrator than using it.

How many jobs does your team run today, and who keeps the orchestrator alive?

Written by Thomas Nys

Fractional Data Architect helping startups and scaleups build data platforms that scale.

More about Thomas Nys →

Recognise the problem? Let's talk about it.