Nobody decided which store wins

Most organisations already have somewhere they put data that spans systems. A warehouse, a lakehouse, a reporting database that predates both. Putting a customer data platform inside the CRM means one of two things has been decided. Either you are accepting a second analytical store on purpose, or you have decided this one supersedes the other and you intend to retire what it replaces.

That decision is almost never made out loud. What happens instead is that marketing builds segments in the new store, finance keeps reporting from the old one, and eighteen months later two teams hold different counts of active customers with no agreed way to say which is right. Both stores get funded, neither gets reconciled, and the reconciliation turns into a standing agenda item that never closes.

The first question in front of an architecture board concerns scope rather than features. Which of your existing stores loses scope when this one arrives, and who signs off on that. If nothing loses scope, you have chosen to run two, and the cost of running two belongs in the business case from the start.

Proximity to the action is the real argument

The strongest case for Data Cloud is how close it sits to the thing you want to happen. Data that will drive segmentation, personalisation, routing or the grounding of an agent benefits enormously from living where that action runs, because the alternative involves a round trip. Export to the warehouse, compute something there, push the result back on a schedule.

That schedule is always slightly stale. A propensity score calculated overnight and written back to a field holds up fine until someone needs to act inside the window. The rep opens the account at eleven in the morning and sees yesterday's view of a customer who has called support twice since. Nobody notices the staleness in a demo. Everyone notices it the first time a churn signal lands after the renewal call.

That is not a marketing claim dressed up in architecture language. Proximity to the action is the real reason the product exists, and if the workload you have in mind genuinely drives something inside Salesforce, the argument is strong enough to carry most of the decision on its own.

Deciding two records are the same person is hard

The second real argument is identity resolution. Deciding that a web visitor, a support contact and a billing record describe one person is genuinely difficult at any scale worth having, because the sources disagree about names, addresses and consent state, and the exceptions need somewhere to go.

Doing that work once, in a place the operational system can see, is worth a lot. The alternative resolves identity in the warehouse and then tries to reconcile the answer back into the CRM, which becomes a permanent project because the CRM keeps creating new records while the reconciliation runs. We have written before about what a shared customer record really needs, and the same point holds here. Resolution has to live where the writes happen.

Those two arguments, proximity and resolution, are the honest ones. Most of the rest of the pitch is a variation on them, or a claim about volume that nobody has tied to an outcome yet.

Reporting-only data does not need to be here

Data that exists to feed a dashboard has no reason to be ingested. If the output is a chart a human reads once a week, the warehouse you already run answers that question at a cost you already understand. Moving it into the CRM's data platform adds an ingestion bill and a second definition of the same measure without changing what anyone does on Monday morning.

The same applies to data that has to be joined with systems Salesforce does not hold. Revenue recognition, inventory positions, service margin, anything that needs the ledger or the supply chain sitting alongside the customer record. Those joins belong where the other side of the join lives, which is the argument we made about the boundary between a lake and a ledger and the one SAP customers are having about their own warehouse question.

A workable test sits behind both cases. If the data never triggers an action inside Salesforce and never takes part in identity resolution, you are holding a reporting asset, and reporting assets should stay where reporting already happens.

The bill follows the ingestion pattern

Consumption pricing changes who controls the cost. A seat-based line item moves when HR hires people. A consumption line moves when an engineer changes a batch frequency, widens a field list or points a streaming source at the platform. The people making those calls are usually nowhere near the room where the budget gets defended.

An ingestion pattern chosen once bills every day afterwards. Somebody sets a source to full refresh because incremental was fiddly to configure during the build, and that choice quietly re-ingests the same rows for the next three years. Nobody revisits it, because from the outside it looks settled.

Rates, included volumes and what counts as a billable operation depend on your agreement and on what you bought, so read your own contract rather than anyone's blog post, including ours. What stays stable across agreements is the shape of the exposure. Ingestion, storage and processing all meter, so the design decisions and the cost decisions are the same decisions. Put an estimated monthly consumption figure next to each source before the build, then revisit it at renewal with actuals in hand.

Volume was never where the value was

The failure people hit hardest is bringing everything in because bringing everything in is possible. The platform ingests, so teams ingest. Eighteen months later there are forty sources, a bill that surprised finance, and a handful of segments that could have been built from four of them.

The discipline that prevents it starts from the action. Name the thing you want to happen. A campaign triggered by a behaviour, a case routed by entitlement, an agent that can answer a billing question without a handoff. Then work backwards to the minimum data that supports it, and ingest that much. The rest waits until something needs it.

This sounds obvious and is almost never what happens, because the build sequence usually runs the other way. Someone lands the sources first so the data will be there when the use cases arrive, the use cases arrive slowly, and by then the bill has its own history and its own defenders. Ingest against a named action or do not ingest.

One store of customer data, one larger obligation

A platform that unifies customer data across sources concentrates personal data that used to sit in separate places under separate rules. Consent captured on a web form, a preference recorded in the call centre, a retention rule that applies to billing records and not to marketing events. Each of those stayed manageable inside its own system. Together they become one obligation that somebody has to be able to describe to a regulator.

Deletion is where this shows up first. A request arrives, and the answer has to cover the unified profile, the source objects it was built from, the segments it belongs to and anything already activated downstream. If nobody has traced that path before the first request lands, the first request takes weeks and the answer you give is a best effort.

Give it a named owner. Not a committee, and not the platform team by default, but a person who can say which consent rule applies to a given attribute and who signs off when a new source is added. That person should own the quality standard too, because a unified profile inherits every problem its sources have, which makes it the kind of data quality work a migration demands rather than a cleanup task someone picks up between releases.

Before the next source goes in, answer one question in writing. Which specific action inside Salesforce will this data change, and what breaks if it arrives a day late. A source that cannot answer the first half belongs in the warehouse, and a source that shrugs at the second half belongs there too.