FractionalDataArchitect
Book a discovery call

Data Architecture Due Diligence Checklist

A practical data architecture due diligence checklist for funding, acquisition, and technical review: evidence, risks, ownership, security, cost, and delivery.

Playbooks· Last updated 3 September 2026· 5 min read

The short version

Data architecture due diligence should establish whether the platform can support the business plan without hiding material security, reliability, cost, ownership, or delivery risk.

The useful output is a decision record:

  • What has been verified?
  • Which claims remain unverified?
  • Which risks can change the deal, valuation, integration plan, or next funding stage?
  • Who owns each follow-up action?

Screenshots and architecture diagrams help. Repositories, logs, invoices, access policies, incident records, and named owners are stronger evidence.

1. Connect the platform to the business case

  • Which products, operations, reports, contracts, or regulatory duties depend on the platform?
  • Which growth assumption creates the largest change in data volume, latency, access, or team workload?
  • Which capabilities are already required by signed customers?
  • Which roadmap items depend on architecture work that hasn’t started?
  • What happens to revenue or operations when critical data is late or wrong?

Record the business consequence beside every material technical risk. A long list of technical debt without consequence is hard to prioritize.

2. Verify the current architecture

  • Ask the team to walk one important data flow from source to decision.
  • Compare the diagram with deployed infrastructure and repositories.
  • Identify undocumented integrations and manual steps.
  • List major platforms, contracts, versions, and end-of-support dates.
  • Check whether the target architecture is being presented as current state.
  • Find components with no replacement path or export strategy.

Evidence: architecture diagrams, cloud inventory, repositories, deployment configuration, data-flow samples, contracts, and decision records.

3. Test reliability and recoverability

  • Identify critical pipelines, datasets, and service objectives.
  • Inspect recent failures and recurring causes.
  • Verify alert ownership and escalation.
  • Ask for the most recent restore or recovery test.
  • Check whether jobs can restart safely.
  • Identify single points of failure in systems and people.
  • Compare stated reliability with logs and incident records.

A backup policy proves intent. A restore test proves recovery.

4. Check data trust and ownership

  • Identify the owner of each critical domain, metric, and data product.
  • Compare definitions of customer, revenue, order, active user, and other key terms.
  • Inspect reconciliation and quality checks.
  • Check how breaking schema or semantic changes are approved.
  • Identify manual corrections that happen outside the platform.
  • Verify whether known limitations are visible to consumers.
  • Find decisions that depend on one person’s memory.

Ownership requires decision rights. A name in a spreadsheet isn’t enough if that person can’t approve definitions, access, quality thresholds, or changes.

5. Review security, privacy, and compliance

  • Map sensitive data and where copies exist.
  • Review human and service access.
  • Check least privilege, access reviews, audit logging, and secret management.
  • Verify retention, deletion, residency, and processor obligations.
  • Inspect how production data reaches development and test environments.
  • Connect data incidents to the organization’s response and reporting process.
  • Record open audit findings and accepted exceptions.

Use the relevant legal and regulatory context. The checklist doesn’t replace legal advice.

6. Rebuild the cost story from evidence

  • Reconcile cloud invoices with products, workloads, teams, or clients.
  • Identify commitments, minimum spend, egress exposure, and exit costs.
  • Inspect idle compute, storage retention, data copies, and expensive recurring queries.
  • Separate current run cost from migration and remediation cost.
  • Challenge savings estimates that have no baseline.
  • Identify vendor concentration and skills constraints.
  • Model the cost of leaving the platform unchanged.

Cost estimates should state assumptions, owner, source period, and sensitivity. A precise number built on a guessed workload is still a guess.

7. Assess delivery capacity

  • Compare the roadmap with the team’s actual skills and available capacity.
  • Inspect deployment frequency, rollback, review, testing, and environment controls.
  • Identify recurring interrupt work.
  • Find repositories or systems only one person can change.
  • Review open roles and whether their definitions match the work.
  • Check external supplier responsibilities and handover provisions.
  • Identify the first capability that must be internal after the transaction or funding event.

Architecture risk often appears as a hiring plan. Verify that the proposed roles solve the observed work.

Classify the findings

Use 4 classes:

ClassMeaningDecision
Verified strengthEvidence supports the claim and the capability fits the planPreserve it
Remediation itemA bounded weakness has an owner and feasible correctionPlan and cost it
Material riskThe issue can affect the deal, integration, compliance, or growth planEscalate it
UnknownEvidence is missing or contradictoryResolve before relying on the claim

Avoid turning every weakness into a material risk. The assessment should show which constraints matter to this transaction or funding plan.

Minimum evidence request

  • Current and target architecture diagrams
  • System, integration, and data inventories
  • Cloud and vendor contracts
  • Recent invoices and usage reports
  • Repositories and deployment process
  • Incident, support, and recovery records
  • Data classification and access model
  • Quality checks and reconciliation evidence
  • Ownership and metric definitions
  • Roadmap, team structure, open roles, and supplier responsibilities
  • Relevant security, privacy, and compliance assessments

Send the request early. Missing evidence discovered at the end creates delay and weakens confidence in otherwise sound work.

Primary references

Frequently asked

What does data architecture due diligence cover?
It covers the connection to the business plan, current architecture, reliability, data trust and ownership, security and compliance, platform cost, and the team’s ability to deliver and operate the system.
What evidence should a reviewer request?
Request deployed architecture and repositories, logs, incident and recovery records, access configuration, quality checks, invoices, contracts, ownership records, roadmap, and team responsibilities. Use interviews to explain evidence rather than replace it.
How is due diligence different from a normal architecture review?
Due diligence evaluates technical facts against a funding, acquisition, or investment decision. It separates verified claims, bounded remediation, material risks, and unresolved unknowns.

Last updated: 3 September 2026

Written by Thomas Nys

Fractional Data Architect helping startups and scaleups build data platforms that scale.

More about Thomas Nys →

Need the evidence before the review starts?