Guide

CRM and Pipeline Data Diligence: What a Buyer Should Test in a Target's CRM

CRM and pipeline data diligence is the pre-close work of testing whether a target's customer and opportunity records can support the revenue forecast a buyer is being asked to underwrite. It sits inside commercial and digital due diligence but answers a narrower question: not whether the growth story is plausible, but whether the data behind it is complete, consistent and reproducible enough to be recomputed independently. The buyer does not own the system, cannot change how anyone uses it, and usually has one export and a few weeks to decide.

Most processes never do this. The financial workstream ties revenue to the audited accounts. The commercial workstream interviews customers and sizes the market. The technology workstream inventories systems and counts licenses. The pipeline, which is the single input doing most of the work in the forward model, arrives as a management summary and is taken at face value. S&P Global's 2026 research found 71 percent of GPs and 53 percent of LPs now put operational value creation ahead of multiple expansion, which makes the operating data the thing being bought. It is still the least tested data in the room.

Testing a system you do not own yet

This is not CRM consolidation. Consolidating CRMs after an acquisition is a design job you run once you own both systems, with the authority to rule on stage definitions, account ownership and segment boundaries, and to rebuild whatever does not survive contact with the new group. Diligence is the opposite position. You cannot interview the sales team freely, cannot instrument anything, cannot ask for a field to be backfilled, and every question costs credibility with a seller running a process on a clock. You are reading a system through a keyhole.

That constraint sets the method. Every test below has to run on data that already exists, be reproducible by someone else from the same export, and produce a number rather than an impression. "CRM hygiene looked weak" does not survive an investment committee. "Thirty-one percent of open pipeline value has had no recorded activity in ninety days" does. Marketing due diligence tests the acquisition engine that produces demand; this tests the record of what that engine produced and what happened to it afterward.

Ask for the export, not the report

The first test happens before any analysis. Ask for a row-level export: every open opportunity as a row, every closed opportunity for the trailing 24 to 36 months, and the account and contact tables sitting behind them. Each opportunity row should carry a unique identifier, account key, owner, created date, current stage, stage-entry date, expected close date, amount, source and last activity date. Ask for the field definitions and picklist values alongside it, because a stage list of eleven values with no written exit criteria tells you something before you read a single row.

What comes back is itself a finding. A clean export in two days means the system holds the shape of the business. A report that management has already summarized, aggregated or filtered means either the raw data does not support the cut you asked for, or somebody would rather you did not see it. A refusal on grounds of technical difficulty is rarely true of any modern CRM and usually means the pipeline lives partly outside it. In one industrial engagement we found double-digit shares of annual pipeline sitting in spreadsheets, invisible to every system of record.

Six tests that produce numbers

Completeness

Compute the fill rate of each load-bearing field across open opportunities, weighted by value rather than by count. Owner, amount, stage, expected close date and source are the five that matter. A 95 percent fill rate by count alongside a 70 percent fill rate by value means the large deals are the ones nobody is documenting, which is the worse of the two results and invisible if you count rows.

Duplication

Duplication at the account level, not the contact level, is what distorts commercial conclusions. Match on normalized name, then domain, then billing address, and report the rate as a share of accounts and as a share of trailing revenue. Where the same buyer exists three times, every repeat rate, cross-sell estimate and concentration calculation in the model is wrong in the same direction: understating how dependent the business is on a small number of customers.

Ownership

Count open opportunities whose owner has left or no longer sells, again by value. Orphaned pipeline is not a hygiene issue, it is a forecast issue: nobody is working those deals and nobody has told the forecast. It is also the cleanest available proxy for sales turnover the seller has not volunteered, because the records outlast the exit interviews.

Stage discipline

Take the written stage exit criteria, if any exist, and test whether behavior matches them. Do late-stage deals carry the artifacts the criteria require. Does every deal pass through every stage, or do half of them jump from stage two to closed won. Do different owners or sites show materially different stage distributions on similar work. Wide variance between owners means the stages are personal habits rather than a process, and stage-weighted forecasting on top of them is arithmetic applied to noise. Consolidation treats inconsistent stages as something to fix; in diligence the job is to measure how far apart they are.

Staleness and slippage

Age every open opportunity from its created date and from its last activity date, and value-weight both. Then measure slippage: how many times the expected close date has moved, and by how much in total. Two patterns show up repeatedly and both are worth pricing. Close dates clustered on the last day of each quarter are a reporting convention, not a prediction. A block of pipeline created in the weeks immediately before the process opened is a pipeline built for a buyer.

Reconciliation to the ledger

Sum closed-won value by month from the export and tie it to booked revenue in the same months from the finance system. The two will not match, and the size and direction of the gap is the finding. A CRM reporting consistently more than the ledger is recognizing deals that never convert to invoices. A CRM reporting less has revenue arriving through routes it never sees, which is common in services businesses and not always a problem: in one PE-backed services company we audited, 38 percent of revenue came through word of mouth that no system recorded. That business was healthy. Its reporting was fiction, and every forward number built on it inherited the error. State the variance as a percentage with a cause attached rather than reconciling it away.

Recomputing the two numbers you were given

Win rate and sales cycle length appear in every management presentation and are almost never recomputed. Do both from the raw export, and do them the honest way. Win rate is closed won divided by all closed outcomes, including losses and deals marked dead or disqualified, by segment and by site rather than blended. Cycle length is the median rather than the mean, measured from created date to closed date, again by segment.

The divergence is the finding, and its direction tells you which problem you have. A management win rate well above the recomputed one usually means losses are being excluded or quietly deleted rather than recorded. A management cycle length well below the recomputed one usually means the clock starts at a later stage than the one deals actually enter at. Either way the forward model is running on a conversion assumption the historical data does not support, and the pipeline coverage required to hit the plan is higher than the plan assumes.

What each failure does to the number you are underwriting

The reason to run these as separate tests rather than roll them into one hygiene score is that each failure corrupts a different part of the model, and the remedies are not interchangeable.

  • Low completeness by value. The pipeline cannot be segmented, so every cut in the model is a guess. Discount confidence in the mix, not in the total.
  • Account duplication. Understates customer concentration and overstates new-customer acquisition. Recheck the concentration threshold and the growth-versus-retention split.
  • Orphaned pipeline. The near-term forecast is overstated by roughly the orphaned value multiplied by the recomputed win rate, and there is a hiring cost and a ramp delay attached to recovering it.
  • Weak stage discipline. Stage-weighted forecasting is unusable. Underwrite on cohort conversion from created date instead, which only works if the history exists.
  • Staleness and slippage. The timing of revenue is wrong even where the total is right, which is a cash and covenant problem before it is a valuation one.
  • Ledger variance. The growth attribution story is unsupported, and the channels being credited with growth may not be the ones producing it.

Only orphaned pipeline and demonstrable pipeline inflation give a defensible basis to move the price. The other four change what the first year of ownership has to contain, and cost time rather than money. A buyer who argues all six as a haircut loses credibility with the seller and the committee at once. A buyer who converts four of them into a scoped work plan arrives at close knowing what the commercial team has to build and roughly what it costs.

When there is no single pipeline to test

In multi-site and services businesses the assumption behind all of the above often fails: there is no single CRM. There is a job management system at three sites, a quoting tool at two more, a shared inbox at the newest acquisition, and a head office spreadsheet that adds them up. The opportunity object may not exist at all, because the commercial unit is a quote, a survey booking or a job.

Do not force those systems into one schema during diligence. The translation takes longer than the deal allows and buries the most useful finding, which is the variance itself. Run the tests per site, on whatever the local equivalent of an opportunity is, and report the spread rather than the average. Two sites converting quotes at twice the rate of the others is a real and bankable finding. A blended conversion rate across six incompatible systems is a number nobody should act on. The revenue stream map is the right place to settle the segmentation these tests then run inside.

The absence of a common system is a cost in its own right, and it belongs in the plan rather than in the price. Somebody has to build what the group does not have, and that work has a sequence and a duration.

What cannot be reconstructed after close

One argument for running these tests before signing rather than in the first month of ownership is that some of the data is perishable. Last activity dates, stage-entry timestamps and source fields are written at the moment something happens, and cannot be backfilled honestly afterward. If the target migrated systems eighteen months ago and kept only open records, the trailing history needed to compute cohort conversion does not exist, will not exist for another two years, and no amount of post-close investment shortens that. Everything else on this list is fixable after close, and most of it gets fixed as part of consolidation anyway.

That asymmetry is the practical reason to put the export request in early rather than saving it for confirmatory diligence. A missing field can be built. A missing year cannot.

Have a revenue problem the board is asking about? Start a conversation.

Frequently asked questions

What is CRM due diligence?

CRM due diligence is the pre-close work of testing whether a target's customer and opportunity records can support the revenue forecast a buyer is being asked to underwrite. It is narrower than commercial due diligence. The question is not whether the growth story is plausible but whether the data behind it is complete, consistent and reproducible enough that an outsider can recompute the headline numbers from the raw records.

What should a buyer ask a target for from its CRM?

A row-level export rather than a report. Every open opportunity as a row, every closed opportunity for the trailing 24 to 36 months, and the account and contact tables behind them. Each opportunity row should carry a unique identifier, account key, owner, created date, current stage, stage-entry date, expected close date, amount, source and last activity date, with the field definitions and picklist values alongside. What comes back, and how quickly, is itself a finding.

How much CRM history does a buyer need?

Twenty-four months of closed history is the working minimum, because twelve months cannot show a full cycle of seasonality or repeat purchase. Beyond thirty-six months the data usually crosses a system migration, at which point field definitions change underneath the analysis. History is also the one item on this list that cannot be created after close, so its absence is a finding to establish during the process rather than after it.

What are the biggest red flags in a target's pipeline data?

Close dates clustered on the last day of each quarter, which is a reporting convention rather than a prediction. A block of pipeline created in the weeks immediately before the process opened. Open opportunities owned by people who have left. Wide variance in stage distribution between owners or sites doing similar work. And a reported win rate that cannot be reproduced from the raw records once losses and disqualified deals are included in the denominator.

Is CRM data diligence the same as CRM consolidation?

No, and the distinction matters. Consolidation is a design job run after close, with the authority to rule on definitions and rebuild what does not work. Diligence is the opposite position: the buyer does not own the system, cannot instrument anything, cannot ask for fields to be backfilled, and has to judge quality from outside on data that already exists. Most of what diligence finds gets fixed during consolidation. The point of finding it early is to price it and plan it.