The meaning lives in the logic, not the tables

Decide whether an SAP number survives being copied and most of the warehouse argument resolves itself. A revenue figure in S/4HANA is correct because of everything wrapped around it. Document types, posting rules, currency translation settings, reversal handling, validity dates on master data, and a chart of accounts somebody spent a year agreeing. Pull the underlying tables and you get the digits without any of that.

I have watched a team rebuild six months of that logic in SQL and still miss the case where a credit memo posts back into a prior period. Their report matched the ledger for eleven months of the year and diverged every December, which is the month people look hardest. Nobody had been careless. They had reimplemented a system that a dozen people had tuned over fifteen years, and they had budgeted three weeks for it.

That is the honest case for keeping analytical work close to SAP. A layer that already understands SAP objects hands you the meaning along with the rows, and the reimplementation it saves is larger than almost every estimate I have seen. Which layers your organisation can actually use depends on your landscape, your release levels and your licence agreement, so confirm your own position with your account team before it goes on a slide.

The questions the business asks cross SAP's edges

The counter-argument is just as real. Questions asked in the operations review rarely stay inside SAP. Margin by customer segment, where segment lives in the CRM. On-time delivery by carrier, where the scans sit in a warehouse system nobody in finance has heard of. Cost to serve by support tier, which needs ticket volumes from the service desk.

No SAP-native layer answers a question about SAP data joined to data it never sees. You can push the other systems into SAP's orbit, and some organisations do, but then you are paying to land customer and operational records in a place designed for a different job. The integration problem moves rather than disappears.

Both positions are right, and they are right about different questions. The design mistake is picking one of them and declaring the other case closed, which is how organisations end up with a vendor-managed layer nobody outside finance queries and a warehouse the controllers refuse to quote.

Copy the number and you own a second definition

Every copy creates a second definition of revenue, and somebody has to keep the two matched. That job never appears on the project plan. It appears eighteen months later, when a regional director's dashboard and the statutory pack disagree by an amount large enough to need an explanation and small enough that nobody can find it before the meeting.

Semantic drift starts small. SAP configuration moves. A new company code arrives, a document type gets added for an intercompany flow, a revenue account is split in two. The extraction was written against last year's configuration and nobody told the data team, because the data team does not sit on the change advisory board. The fight over who gets to define revenue is the same fight in a different meeting.

If you copy, fund the reconciliation as continuing work with a named owner, not as a go-live task. One person, a monthly tie-out against the source, and the authority to stop a dashboard being published when it fails. The discipline that keeps ERP master data governance honest applies to the definitions sitting on top of it.

The extraction path you pick becomes permanent

Whatever pattern you choose for getting data out will still be there in five years, carrying more volume and more dependencies than anyone planned. It deserves more thought than it usually gets, because the cheapest option at the start is almost always the expensive one by year three.

The nightly full-table dump is quick to build and quietly awful. It grows with the source, it lands in the same window as the close batch, and it carries no record of what changed, so every downstream question about history gets answered by guesswork. Change-based extraction costs more up front and ages far better. Extraction through released, supported interfaces costs the most to stand up and is the only one that tends to survive an upgrade without a rewrite.

Which of those you may use depends on your release, your deployment model and what your agreement permits, and those terms have moved more than once. Get your specific position confirmed in writing rather than working from a conference slide, whether you are building on SAP BTP or somewhere else entirely.

Authorisation is the part that becomes a disclosure problem

This is the one I raise first in a design review. SAP enforces access at a granularity that is genuinely hard to rebuild anywhere else. Company code, profit centre, cost centre, plant, sales organisation, and for people data a set of personnel structures that decide who may see a salary. Those checks run on every read, against objects the security team has maintained for years.

A warehouse that copies the rows without those checks hands a broad reporting population figures they were never cleared to see. Payroll cost by named cost centre. Margin on a deal an account team should not know about. Restructuring provisions visible to the division being restructured. The consequence is a disclosure incident rather than a reporting inconvenience, and it usually surfaces because somebody browsed a dashboard out of curiosity.

Rebuilding those rules in the reporting layer is possible and it is real work. Somebody has to map SAP's authorisation objects onto row-level rules in the target, keep the mapping current as the organisation reorganises, and show an auditor that the two agree. Budget that properly, or scope the copy narrowly enough that the sensitive dimensions never leave SAP at all.

A copy that ignores period close reports on moving numbers

Financial data is period-bound and copies forget it. The ledger knows whether a period is open, closed, or open only for specific activities. A pipeline that reads at two in the morning knows the time it ran and nothing else.

So the margin report goes out on the third working day, the controllers post another forty adjusting entries on the fourth, and the number everyone quoted on Monday was never the number. Nothing in the pipeline was broken. It reported a period that was still moving, and no part of the page said so.

Carry period status through with the data and show it on anything financial. Reporting that mixes a controlled source with a copy owes its reader the extract time and the close state, an argument made at more length in the difference between a lake and a ledger.

Write the question register before anyone builds a pipeline

The frame holds up in most rooms. Questions that stay inside SAP objects and depend on SAP's own semantics belong close to the source. Questions that cross into customer, operational or people data SAP never sees belong in the shared platform. Most organisations of any size need both, and the architecture document should say which is which instead of pretending one layer wins everything.

The expensive mistake is copying everything. Teams extract the full schema because choosing feels risky, then carry the cost of every field forever: authorisation exposure on columns nobody asked for, reconciliation surface on definitions nobody uses, upgrade risk on tables nobody reads. Copy the fields the questions require and stop there. That only works if you have done the work of making analytics questions a written requirement.

Before any extraction gets designed, write the register. Every analytical question the business actually asks, the SAP objects that answer it, the fields required, the authorisation dimension that has to survive the copy, the period rule, and the person who owns the definition. Take it to finance and to the security team and get the financial rows signed. Most of the pipeline arguments disappear once that page exists, and the ones that remain are worth having.