How to Run a Benchmarking Analysis (PDF)
Enter your email to unlock this guide. 9 pages, print-ready, with all the key takeaways.
Overview
Transfer pricing benchmarking analysis is how multinational groups prove that their intercompany transactions are priced the way independent companies would price them. For in-house tax teams and transfer pricing advisors, a defensible benchmarking study is both a compliance requirement and the first line of defense in an audit. This section sets out the framework: the arm’s length principle behind it, how to choose a method and tested party, how comparability is judged, the profit level indicators, and the arm’s length range the study produces.
What a Benchmarking Study Is
When two companies in the same multinational group transact with each other, the price they use directly shifts profit, and therefore taxable income, from one country to another. This happens whenever one subsidiary sells goods to another, or provides services, lends money, or licenses technology. Tax authorities care because an artificially high or low intercompany price can move profit out of a high-tax jurisdiction and erode its tax base.
To prevent this, the arm’s length principle requires that related parties transact at the same price independent companies would have agreed to under comparable circumstances. It is the cornerstone of the OECD Transfer Pricing Guidelines and most local rules, including US section 482. Being part of the same group should confer no pricing advantage a company could not have negotiated at arm’s length with a third party.
That raises the operative question: what price would independent parties actually charge? A benchmarking study answers it. By examining the financial results of comparable independent companies, the study derives an arm’s length range, the band of outcomes the market itself produces, against which the controlled transaction is tested. It earns its central place for three reasons:
- Defensibility. A study that withstands scrutiny is the taxpayer’s first line of defense in an audit, and is required documentation in most jurisdictions.
- An objective answer. It produces a market-derived range rather than a single “correct” price.
- Two audiences. The taxpayer uses it to set and defend policy; the tax authority uses it to test that policy.
The Five Transfer Pricing Methods
Transfer pricing rules recognize five methods for testing whether a transaction is arm’s length. The right choice depends on the transaction and the data available. The five fall into two groups: transactional methods, which test prices or gross margins on individual transactions, and profit-based methods, which test net profitability.
| Method | Type | Typical use |
|---|---|---|
| CUP (Comparable Uncontrolled Price) | Transactional | A near-identical independent price exists (commodities, loans, royalties). |
| Resale Price | Transactional, gross | Distributors reselling without adding much value; tests gross margin. |
| Cost Plus | Transactional, gross | Manufacturers and service providers; tests gross markup on costs. |
| TNMM / CPM | Profit-based | Most widely used. Tests a net profit indicator of the routine party. |
| Profit Split (PSM) | Profit-based | Two-sided, intangible-rich arrangements; splits combined profit. |
Benchmarking studies rely almost entirely on the two profit-based methods. The Transactional Net Margin Method (TNMM) under the OECD Guidelines, and its US equivalent the Comparable Profits Method (CPM), test the net profit earned by one party against the net profits of independent comparable companies. They tolerate moderate product and functional differences, which is why they are by far the most commonly applied methods in practice, and why net-margin comparables can be assembled from public financial filings even where reliable prices or gross margins cannot.
Choosing the Tested Party
A one-sided method like TNMM works only if a single routine party can be isolated and tested. That party is the tested party: the least complex entity in the transaction, which performs routine functions, owns no valuable intangibles, and bears limited risk. It is chosen because reliable comparables for a simple profile are far easier to find, and its results are the most defensible. The more complex counterparty keeps the residual (system) profit.
- Distribution. A local limited-risk distributor buying from a foreign principal is usually the tested party; you benchmark its operating margin.
- Services. A captive IT or back-office center serving group affiliates is the tested party; you benchmark its cost-plus markup.
- Manufacturing. A contract manufacturer with no market or IP risk is the tested party, not the principal that owns the technology and brand.
When to Use the Profit Split Method
When no single routine party can be isolated, a one-sided benchmark breaks down. The Profit Split Method (PSM) applies where both parties contribute uniquely valuable intangibles or share significantly in value creation. Common examples are a joint development of technology, or two principals that each own key IP. Rather than benchmark one side to external companies, PSM divides the combined profit between the parties in proportion to their relative contributions.
Rule of thumb. If you can identify one “simple” party whose return external comparables can stand in for, a benchmarking (TNMM / CPM) method works. If value creation is genuinely shared and intangible-rich on both sides, reach for a profit split.
Comparability and the Five Factors
A benchmark is only meaningful if the independent companies it relies on are genuinely comparable to the related party being tested. Comparability analysis is the disciplined process of confirming that, or adjusting for differences where they exist. Under both the OECD Guidelines and US section 1.482, comparability is assessed across five factors. For each, the practical question is whether a difference between the tested party and a candidate comparable is large enough to distort the profit indicator being tested.
- Characteristics of the property or services transferred. What is actually being sold, licensed, or provided, including any embedded intangibles. Close product similarity matters most under transactional methods; under TNMM it matters less, but a fundamentally different product can still distort margins.
- Functions performed (taking into account assets used and risks assumed). The activities each party carries out, the assets it deploys, and the risks it bears. This functional profile is the most decisive factor, because functions, assets, and risks are what drive how much profit an entity should earn.
- Contractual terms. How rights, obligations, and risks are allocated between the parties, including volume, warranties, credit terms, and ancillary services. Terms that shift risk change the expected return.
- Economic circumstances of the parties. Market-level conditions such as the size and level of the market, geography, and the intensity of competition. Comparables from a very different economic setting may earn returns the tested party’s market would never produce.
- Business strategies pursued by the parties. Strategies such as market penetration or a temporary margin sacrifice for share. A company mid-strategy can post results that do not reflect its steady-state return.
As comparability rises, the number and size of differences that could distort the analysis falls. Where differences remain, adjustments can improve comparability, but each adjustment introduces its own estimation error. A function-first approach screens candidates on their core activity before applying numeric filters, so the final set reflects economic substance rather than a shared industry code.
Profit Level Indicators
Under TNMM and CPM, the profit level indicator (PLI) is the specific financial ratio used to express the tested party’s profitability and compare it to the comparables. The right PLI must align with the tested party’s functions and be measured consistently on both sides, pairing profit with the base that best reflects what drives the tested party’s value-add: costs, sales, or assets.
| PLI | Formula (concept) | Best suited to |
|---|---|---|
| Operating Margin (OM) | Operating profit / Sales | Distributors and resellers. Value tied to sales volume. |
| Net Cost Plus Markup (NCPM) | Operating profit / Total costs | Service providers and contract manufacturers. Value tied to cost base. |
| Berry Ratio | Gross profit / Operating exp. | Intermediaries whose value-add is opex-driven, not inventory. |
| Return on Assets (ROA) | Operating profit / Op. assets | Asset-intensive activity such as full-fledged manufacturers, where assets drive returns. |
| Return on Capital Employed | Op. profit / Capital employed | Capital-heavy operations and capital booking locations; a fuller asset-based view. |
| Full Cost Markup (gross) | Gross profit / COGS | Cost-plus manufacturing under the traditional Cost Plus method. |
- Match the base to the driver. Use a cost-based PLI (NCPM) when reward should track activity or effort (services); a sales-based PLI (OM) when reward tracks turnover (distribution); an asset-based PLI (ROA) when capital is the key input.
- Berry-ratio caution. Useful for pure intermediaries, but unreliable where the company carries significant inventory or its operating costs do not proxy for its functions.
- Consistency. The PLI must be computed the same way for the tested party and every comparable, using comparable accounting definitions.
In practice, the two workhorses are NCPM for service providers and operating margin for distributors. Most studies default to one of these unless the functional profile clearly points elsewhere.
The Arm’s Length Range
A set of comparables never yields one number. Independent companies with similar profiles still post a spread of results. Together they form the arm’s length range: the band of profitability the market actually produces for that kind of activity. This range is the defensible target a taxpayer should design its transfer pricing policy to land within for the tested party.
- Why a range, not a point. Comparability is never perfect, so a band better reflects genuine arm’s length behavior than a single figure.
- The interquartile range (IQR). The most common convention spans the 25th percentile (lower quartile) to the 75th percentile (upper quartile) of the comparable results. Results below the 25th and above the 75th are set aside, so the range is anchored on the central, more reliable observations.
- Multi-year data. Comparable results are usually pooled over several years (most often a three-year period) rather than a single year, smoothing out business-cycle swings. The tested party is measured on the same basis.
| Statistic | Value |
|---|---|
| Maximum | 9.8% |
| 75th percentile (upper quartile) | 6.6% |
| Median (50th) | 4.2% (typical adjustment point) |
| 25th percentile (lower quartile) | 2.0% |
| Minimum | 0.4% |
If the tested party’s profitability falls within the range, it is generally treated as arm’s length and no adjustment is required. A result outside the range invites an adjustment, and because transfer pricing is a two-sided story, the direction of the miss points to which country bears the exposure.
- Below the range. The tested party earned less than arm’s length and was undercompensated. The exposure sits in the tested party’s own country, whose tax base has been understated; that authority would adjust the result upward, typically to the median.
- Above the range. The tested party earned more than arm’s length. From its own country this may look like a low-risk position, but the profit had to come from somewhere: the exposure sits with the counterparty’s country, which gave up profit to an entity that should have earned only a routine return.
The Step-by-Step Guide
The framework above becomes a defensible study only through a disciplined, documented search. The workflow moves from defining the universe of candidate companies, through independence and functional-ratio screens and a qualitative review, to the computed range. Every choice should be documented with its rationale so the search can be reproduced.
Set the Starting Point
Define the universe of candidates using industry classification and keyword logic that capture the tested party’s activity. Comparable financials are drawn from a financial-information database. Practitioners typically rely on a commercial provider supplemented by public regulatory filings, and where a maintained standard set exists for the function and region, it is normally used as the base population.
- SIC codes / industry categories. Selected on the database’s own classification criteria, with the rationale documented.
- Keyword search. Business-description terms that capture the activity and exclude adjacent ones.
Apply Independence and Data-Sufficiency Screens
Two threshold tests confirm each candidate is both an eligible comparable and a usable one. The independence screen ensures the company is a genuinely uncontrolled party; the data-sufficiency screens ensure enough financial data exists to compute a reliable result.
| Screen | Rationale | Removes |
|---|---|---|
| Independence | Comparables must be independent of any group. Controlled companies reflect intra-group dealings, not open-market pricing. | Subsidiaries and majority-owned companies |
| Sales availability | Sales must be reported for at least two of the three years in the measurement window. | Companies with insufficient sales data |
| Operating income availability | Operating income must be reported for at least two of the three years so the PLI can be computed. | Companies with insufficient profit data |
| Three consecutive years of operating losses | Persistent losses may signal going-concern issues or a different functional profile. Flagged companies are reviewed rather than cut automatically, since mechanically dropping loss-makers introduces an upward bias. | Flag for review (not an automatic cut) |
Run Quantitative (Functional Ratio) Screens
Quantitative screens use financial ratios to detect companies whose functional profile differs from the tested party’s. The logic is indirect. A ratio that is out of line, whether heavy inventory, significant manufacturing assets, large marketing spend, or material R&D, signals functions, assets, or risks the tested party does not have. Only the ratios relevant to the tested party’s function are applied.
| Screen | What it signals (and excludes) |
|---|---|
| Raw materials & work-in-progress ratio | Manufacturing or assembly activity; a distributor would not carry significant raw materials. |
| Finished-goods inventory to sales | Distribution or stock-holding activity in a contract-manufacturer search. |
| Inventory to sales (service / management) | Non-service functions, where a service provider holds significant inventory. |
| Inventory to sales (distribution) | Elevated inventory risk from holding stock for extended periods. |
| Inventory to average operating assets | A significantly different inventory-risk profile than the tested party. |
| PP&E to sales | Capital-intensive functions different from a service profile. |
| PP&E to average operating assets | Manufacturing activity, signalled by high fixed assets. |
| Advertising to sales | Ownership of marketing intangibles from significant brand investment. |
| Average total assets to sales | Asset-heavy operations with significant equipment or resources. |
| Operating expenses to sales | Additional functions and costs beyond a typical manufacturer or distributor. |
| R&D to sales | Ownership of valuable intangibles that compromise a routine-return profile. Report the tested party’s own value where this screen is used. |
Conduct the Qualitative Review
Numeric screens narrow the population but cannot confirm comparability on their own. The qualitative review reads each surviving company against the comparability factors, focusing on functions, assets, and risks. A key principle: functional relevance usually matters more than product similarity. A company in a different product line but with the same routine functions, risks, and asset profile is often a better comparable than one selling identical products while performing very different functions. This is the heart of the function-first approach. Each accept or reject decision must be recorded with a clear reason.
Make Comparability Adjustments
Where consistent differences remain, adjustments can improve comparability before the range is computed. Each adjustment’s rationale, inputs, and formula must be specified, since every adjustment introduces estimation error.
- Working-capital adjustments. Applied to inventory, AR, and AP using an appropriate interest rate, when differences reflect genuine business expectations. Many non-US authorities do not expect, and may reject, these adjustments.
- LIFO / FIFO adjustment. Restates comparables onto a consistent inventory-costing basis where they differ from the tested party, so margins are measured like-for-like.
- Adjusting to cash. Reconciles accrual versus cash timing differences that would otherwise distort comparability.
Compute the Range and Document the Search
Once the final set is fixed, compute the chosen PLI for each accepted comparable and aggregate into the arm’s length range. Settle the conventions before viewing results, so the outcome is not reverse-engineered: the range basis (interquartile vs. full), the averaging method (simple vs. pooled or weighted average), and the tested-party measurement basis.
The OECD percentile method. Quartiles rarely fall exactly on an observed data point, so they are interpolated between the two nearest observations:
- Rank the data. Order all comparable PLIs from lowest to highest.
- Locate the percentile position. Compute the rank position for the target percentile; the OECD approach commonly uses a position such as
P x (n - 1) + 1within the ranked list. - Interpolate. If the position falls between two ranks, take the lower observation plus the fractional distance multiplied by the gap to the next observation.
- Read the range. The 25th and 75th percentiles bound the interquartile range; the 50th is the median and the usual point of adjustment.
The exact positioning formula varies by convention (US regulations and some local rules differ from the OECD method), so the chosen method should be stated.
Building a Defensible Study
A benchmarking study is only as strong as its documentation. What makes it defensible is not the range it produces but the trail that leads to it: a reviewer following the recorded steps should arrive at the same comparable set and the same conclusion. A study that cannot be reproduced can always be re-argued by an examiner, and usually is.
That trail runs through every stage of the search. The starting population is defined with documented classification and keyword criteria; the independence and data-sufficiency screens are stated with their rationale; the functional-ratio screens applied are only those relevant to the tested party’s profile, each with its reason; every accept-or-reject call in the qualitative review is recorded; and each comparability adjustment carries its inputs and formula, because each one adds estimation error. The conventions that shape the result, the range basis, the averaging method, the tested-party measurement basis, and the percentile method, are all specified.
Two habits keep the study honest. First, fix the conventions before looking at the numbers, so the outcome is never reverse-engineered from a desired answer. Second, record the reason for every screen and every accept-or-reject decision, so the study reproduces rather than has to be defended from memory. A study built this way holds up in an audit because anyone can retrace it, which is the whole point of benchmarking: not a single correct price, but a market-derived range a taxpayer can stand behind.
Related reading
- →Selecting the Tested Partythe structured framework for choosing the least complex entity to benchmark
- →Comparability Analysis: The Five Factorsthe five factors that decide whether a candidate company is a reliable comparable
- →Selecting the Profit Level Indicatorhow to match the profit level indicator to the tested party's functional profile