Skip to main content

      Generative AI has established itself in transfer pricing (TP) with astonishing speed. Many teams now use Copilot, ChatGPT, Gemini, Claude or similar systems as a matter of course, for example to process information, prepare for tax audits or structure local files – that is, the country-specific transfer pricing documentation of multinational companies. 

      However, in discussions with companies, we are frequently told of a recurring observation: the AI delivers surprisingly good results for one task, only to fail shortly afterwards when faced with a seemingly similar problem. This unpredictability makes the use of AI very challenging for companies. 

      Why it is difficult to assess the quality of AI results

      The problem does not lie solely in the quality of the AI models. Nor is it because companies have not yet gained enough experience with AI. It lies in the fact that the performance limits of generative AI do not follow a linear path: Good results for one task are no guarantee that the technology will perform just as reliably when faced with a similar problem.

      Researchers at Harvard Business School and the Digital Data Design Institute at Harvard have coined a term for this phenomenon: the ‘jagged frontier’. 

      Figure: The Jagged Frontier: AI can perform well on one task but fail on another that appears similar. Source: Own illustration based on Dell'Acqua et al. (2026), ‘Navigating the Jagged Technological Frontier’.

       

      In transfer pricing, standardisable documentation tasks are often closely intertwined with tasks involving complex requirements that call for specialist experience and judgement. It is crucial for TP teams to be able to reliably assess which steps can be effectively supported by AI and which still require human expert judgement. This assessment cannot be delegated to the technology; it is a management responsibility.

      auto_stories

      Analyse und Einordnung: Wir zeigen anhand von Umfrageergebnissen die KI-Wahrnehmung in der Bevölkerung auf und erklären, wie in Zeiten von Deepfakes und viral gehender Falschinformationen der sichere, vertrauenswürdige, transparente und ethisch vertretbare KI-Einsatz ermöglicht werden kann.

      Better prompts do not solve the problem

      To obtain reliable results, a better prompt is often not enough due to the ‘jagged frontier’. Well-formulated instructions can improve the quality of AI results, but they do not push the model’s performance limits. 

      The quality of a result often depends on whether the AI has access to the relevant information in the first place. In transfer pricing, for example, this includes local documentation requirements, country-specific peculiarities, guidelines from the organisation’s own transfer pricing policy, or insights from past tax audits. If this context is missing, no prompt – however precise – can replace it.

      Furthermore, how well a model is generally rated says little about how reliably it performs the very tasks that the team encounters most frequently. The management task therefore begins with an assessment of the team’s own task portfolio, rather than with the selection of the supposedly best model. 

      Context is more important than the prompt

      AI requires not only a task description, but also a clearly defined working framework. It must know what the objective is, what rules apply and how a good result can be recognised. The term ‘context engineering’ has become established to describe this approach. It means that the AI is not merely given a task, but also the necessary subject-matter context.

      In transfer pricing, for example, this includes:

      • internal TP policies
      • local file templates
      • country-specific documentation requirements
      • approved wording from previous projects
      • guidelines on the choice of methodology and documentation standards

      The more carefully this context is curated and provided, the more consistent and comprehensible the AI results become.

      Context engineering alone is not enough

      In practice, another challenge becomes apparent: a significant proportion of transfer pricing knowledge is not documented in many companies. In their day-to-day work, experienced experts draw on knowledge gained from years of experience, which is rarely set out in guidelines or templates. This includes, for example, why a particular method is routinely accepted in a given country, which phrasing has proved effective in tax audits, or where arguments should be documented in a particularly robust manner. As long as this knowledge remains largely in the minds of individual staff members, it can neither be utilised systematically nor supported by AI.

      How knowledge moves from people’s minds into systems

      In addition to context engineering, it is therefore important to document implicit knowledge and make it available as usable context building blocks. This creates a better foundation for AI applications whilst also strengthening knowledge management. Lessons learnt from projects, audits and methodological decisions are recorded in a transparent manner and are thus also available to new staff, stand-ins and future generations. 

      What successful transfer pricing teams do differently

      Another important point: in practice, we see that companies which successfully deploy generative AI do not usually regard the technology as the sole solution. Instead, they analyse their processes. The key issue here is deciding which individual work steps can be usefully supported by AI and which should remain entirely in human hands. This decision is not a technical one, but a management task: tasks, context, risks and quality standards are actively shaped, rather than simply entering queries into an AI tool. Dealing with AI is therefore similar to delegating a task to a team member:

      • Focus on individual steps rather than overall processes

        AI will rarely be able to take over an entire transfer pricing process. Those responsible should therefore break the overall process down into its individual steps and assess where the use of AI is both sensible and responsible, and where a specialist tax assessment is essential.

      • Clearly define roles

        Not every use of AI has the same objective. Sometimes it serves as a research aid, sometimes as a sparring partner, and sometimes as the basis for an initial draft. It is therefore important to be clear, before using it, what contribution the AI is intended to make at each stage of the process. The requirements for an AI that structures information differ from those for an AI that develops lines of argument or challenges existing positions. The more clearly this role is defined, the easier it is to categorise, evaluate and verify the results. 

      • Make reviews a binding requirement

        The persuasive language of modern models is one of their greatest strengths. Yet this is precisely where the risk lies. Well-formulated answers often appear more plausible than they actually are. For this reason, AI outputs should always be regarded as working drafts that require verification. An AI-generated draft for a benchmarking narrative should not merely be linguistically convincing. It should also highlight assumptions, data points, uncertainties and open questions requiring further verification.

      Conclusion: start small, learn in a structured way

      You don’t need a comprehensive vision to get started. It makes more sense to carry out a clearly defined practical test: select a process – such as the local file preparation cycle or preparations for the next tax audit – evaluate the individual steps, provide the necessary context and establish review rules.

      The real management task is not to automate as much as possible, but to decide clearly which tasks generative AI can take on, where expert oversight is required, and how the two interact. Because the models are improving by leaps and bounds, such an assessment is never ‘complete’ and should be reviewed regularly.

      More KPMG Insights

      Your contact

      Magdalena Bonna

      Partner, Tax, Global Transfer Pricing Services – Head of Center of Excellence for Transfer Pricing Innovation & Technology

      KPMG AG Wirtschaftsprüfungsgesellschaft