Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The seven steps in the Network World Big Data delivery model are Collect, Process, Manage, Measure, Consume, Store, and Data Governance. They describe connected responsibilities across a data lifecycle, not a one-way checklist: governance applies throughout, while storage and quality decisions must be made before users rely on the results.

What are the seven steps in Big Data delivery?

The model names seven stages. In practice, work may move between them as sources change, quality issues are found, or users need different outputs.

1. Collect

Gather data from source systems and make it available to processing systems. Collection may involve distributing data across nodes so different portions can be processed in parallel. Decide how to extract it without putting undue load on production systems, and record where it came from.

2. Process

Run computations on collected data, often in parallel, then combine the results into datasets that people or software can use. Processing can include cleaning, joining, aggregating, or applying analytical logic. Keep the transformation steps understandable so results can be checked and, when necessary, recreated.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

3. Manage

Make diverse data understandable and safe to use. Define fields and business terms, add context, cleanse inconsistencies, and audit access and security. Management concerns the condition and usability of data; governance sets the policies and decision rights that direct this work.

4. Measure

Choose measures that follow from the business requirement, then track whether the data delivery is meeting them. Useful measures can include integration or correction rates, quality checks, and the outcome the analysis was intended to support. A metric without a connection to the original decision can report activity without showing whether delivery is useful.

5. Consume

Specify how people and systems access, use, or update the delivered data. The access method and format should fit the intended task—such as analysis by a person or input to an automated process—and remain consistent with the original requirement.

6. Store

Choose storage deliberately for both active processing and longer-term retention. A system optimized for analysis may not be the best place to preserve raw inputs for replay, machine-learning work, or later reprocessing. Retention and access requirements should inform the choice rather than being left until after a pipeline is built.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

7. Data Governance

Set business-driven definitions, policies, decision rights, escalation paths, privacy enforcement, and oversight. Although governance appears as the seventh named stage, it spans collection, processing, management, measurement, consumption, and storage. Establish it at the outset and maintain it as the pipeline changes.

How do the seven stages translate into a working pipeline?

A March 7, 2026 implementation guide gives a practical sequence that starts with the decision to support and ends with operational checks. These are implementation actions, not replacement names for the seven-stage model.

1. Define the question

State what decision, question, or KPI the data must support. Work backward from the intended use: this helps determine which sources, transformations, freshness, and access are actually needed.

2. Inventory sources and plan extraction

List the relevant source systems, the data each can provide, and the extraction method. Account for source-system load so collection does not damage production performance. Clarify what happens when a source is unavailable or its structure changes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

3. Choose the destination

Match storage to the job. A cloud data warehouse is suited to SQL analytics; object storage can hold raw or semi-structured data and support machine-learning use cases. Some designs use both: object storage for retained inputs and a warehouse for query-ready datasets. The choice affects retention, reprocessing, and how consumers access data.

4. Choose batch or streaming

Let the required decision freshness determine the delivery pattern. Batch processing is generally simpler to build and backfill, and is appropriate when updates can arrive on a schedule. Streaming is justified when a person or system must act within seconds. Do not choose streaming solely because the data is large; it adds operational concerns such as late events and state management.

5. Load raw data, then transform it

The guide recommends loading raw data first and transforming it in the warehouse, an ELT approach. Retaining the raw inputs preserves a basis for reprocessing when transformation logic changes or an earlier result needs to be investigated.

6. Orchestrate dependencies

As workflows grow, use a scheduler or orchestrator to coordinate task order and dependencies. Plan for retries and backfills, and make failures visible rather than allowing downstream work to run on incomplete inputs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

7. Test quality and monitor freshness

Add data-quality checks and freshness monitoring before consumers depend on the output. Tests should reflect the expected shape and meaning of the data; freshness checks should identify when an update has not arrived as expected. Define who investigates failures and what consumers should do while data is stale or invalid.

Should raw data go to object storage or a warehouse?

The answer depends on how the data will be queried and whether preserving source-form inputs matters. The implementation guide gives these roles as a starting point:

Destination Best-fit role in the guide Trade-off to consider
Cloud data warehouse SQL analytics and query-ready datasets Useful for analytical querying; do not assume it replaces a retained raw-data layer when replay or ML use is needed.
Object storage Raw or semi-structured data and machine-learning data Can preserve inputs for later use; plan separately how users or downstream jobs will query and transform them.
Both Retained raw inputs plus warehouse datasets for analytics Supports distinct retention and query roles, while adding another component to manage.

These are architectural roles, not a requirement to use a particular vendor. The guide names Snowflake, BigQuery, Redshift, and Databricks SQL as warehouse examples, and S3, GCS, and ADLS as object-storage examples.

When should you use batch instead of streaming?

Choose based on the time window in which someone or something must act on the data. If the decision can wait for a scheduled update, batch is typically simpler and easier to backfill. If action is required within seconds, streaming may be warranted. Before committing, consider not only latency but also operational burden: retries, late-arriving events, state management, and changes to source schemas.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How do you monitor data quality and freshness?

Monitoring should make delivery problems visible before they quietly affect decisions. Establish checks before consumers rely on the output, and connect each check to an expected condition or delivery promise.

  • Quality: Test that fields and records meet defined expectations, and track correction or quality rates where relevant.
  • Freshness: Monitor when data last arrived or was updated against the cadence the use case requires.
  • Failure response: Assign ownership for investigating failed checks, source delays, and schema changes, and define how downstream consumers are notified.
  • Outcome: Measure whether the delivered data supports the business question or KPI it was built for.

The guide names Great Expectations and Soda as data-validation examples. They are options, not mandatory tools.

How do you compare delivery designs?

Use the same decision criteria for competing approaches rather than comparing vendors in isolation:

  • Business outcome: Is the pipeline tied to a specific decision or KPI?
  • Freshness and latency: Does the use case need data in seconds, minutes, hours, or daily?
  • Source impact: Can extraction avoid overloading production systems?
  • Storage role: Is the priority replay, machine-learning use, analytics queries, or a combination?
  • Quality and governance: Are definitions, tests, privacy rules, and decision rights explicit?
  • Operational burden: How will the design handle retries, backfills, late events, state, and schema changes?
  • Cost and reversibility: Which choices will be expensive or difficult to change after adoption?

The March 7, 2026 guide lists Airbyte and Fivetran as connector examples, dbt for transformation, and Airflow, Dagster, and Prefect for orchestration. These names illustrate tool categories; the appropriate choice depends on the design and its operating requirements.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is there another seven-part Big Data framework?

Yes. ABS Group’s 2016 framework describes an analytics operating cycle rather than the Network World delivery stages:

  1. Set the vision and foundations.
  2. Prioritize strategic business opportunities.
  3. Select the best analytics approach.
  4. Secure the data.
  5. Extract the insights.
  6. Define and execute the response.
  7. Continuously improve.

Its first two factors establish enterprise direction, the next four guide analytics work, and the final factor keeps the cycle improving. Use it as a complementary view of how an organization chooses and acts on analytics, not as a substitute for the seven delivery-stage names.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.