We keep hearing how AI only works as well as the data behind it. But for many organisations, years of underinvestment in data means they are now trying to unlock the value of AI while also dealing with questions around quality, governance, ownership and trust.
Chris Middlehurst, Partner in Data and AI at KPMG UK, joins the podcast to explore what it really means to get your data AI-ready. He explains why good quality data matters, how organisations can identify the data AI should and should not have access to, and why human oversight remains essential as AI starts to support data stewardship itself.
We also discuss kite-marked data, data lineage, medallion architecture, governance, and why the best place to start is with the value case you are trying to solve.
What you need to know
- AI works best when it is built on trusted, good quality data
- Organisations should start with the value case, then work backwards to identify the critical data that feeds it
- Not all data should be made available to AI. Scope, access and governance all matter
- Kite-marked data gives organisations a clearer view of which data has been checked, validated and can be trusted
- Data stewards play a key role in setting standards and checking quality
- AI itself can support data management, including first drafts of catalogues, lineage maps and quality checks
- Bias needs to be considered use case by use case, based on the outcome the AI is being asked to support
You may also be interested in a previous episode, 'Why am I not seeing value in AI?'. We discuss why many organisations are struggling to translate AI experimentation into measurable impact and what needs to change to unlock value at scale.
Providing the insights on this episode:
Chris Middlehurst
John Robertson
All in just 15 minutes.
The Insight in 15 is KPMG UK's flagship podcast for business leaders and decision makers.
Join us every fortnight for a fresh perspective on the issues shaping the future for your business, people and communities.
No filler. We cut to the chase, setting out the risks and opportunities, and providing insights you can put into action straight away.
Episode transcript
John Robertson: Hi, I'm John Robertson, and this is The Insight in 15. Now, if you're an avid listener to the podcast, you'll know that we love a bit of AI. We've talked about agentic AI, we've talked about the value of AI. And on our previous episode we talked about the cost of AI. And we're going to talk about AI again today.
But today we're going to be talking about data and AI. And I'm joined by Chris Middlehurst, who's a Partner in Data and AI here at KPMG in the UK. Hi, Chris. Thanks for joining us.
Chris Middlehurst: Good morning. No problem.
John: So Chris, we only have 15 minutes on the podcast, so we'd better get cracking. I guess the first question is why are we talking about data and AI today? What's the issue organisations are having with that?
Chris: Well, it's a good question actually. And I talk to a lot of our clients about this. Look, everybody's on the hunt for value that's coming out of AI. And everyone is also racing to the top for AI as well. So they want to outcompete their competitors. And AI, they see, is the way of doing that. Now, fundamentally AI works on great data. So the fact that they've underinvested in data for many, many years is now causing them a slight challenge because they need to access it, it needs to be good quality, in order to trust the outcomes that AI is bringing.
John: So what does good quality data look like for AI?
Chris: Because I have this picture that AI can just go and look at any data and kind of work with it. And it can, and therein lies a slight challenge, right? So AI really doesn't distinguish between good quality and bad quality data. But we as humans need to understand what we believe is good quality data. And we feed AI that data.
So we like to use a term called kite-marked, which is when data has gone through a process that is quality checked. So you know that it's complete as far as it can be. You know that it's accurate and you can verify its accuracy in the context of your own business. Okay. And so if you can do that and kite-mark that data and then offer it up to an AI model, then at least you know you trust the input.
What AI does with it after that is a different matter. And the data scientists spend a lot of time trying to tweak those models to get the best outcome. But that's what we describe as good quality data.
John: Okay. If I'm looking for AI to do a particular task for me or take on a process if it's kind of an agentic workflow, how do I determine what data I need to give AI access to? How do I go about doing that?
Chris: Yeah. So there's a couple of things. So let's take a process in a business for example. Most processes in business are well mapped out. And we understand the data that services that process. Because data isn't just there for insights and AI, it's there to help you execute a process. Okay. So we can look at each of the steps of a process and we understand what we call the critical data elements that feed that process, that make it work if you like. And that's your first point of call. It's work backwards and go and find the source of those critical data elements. Number one.
And then secondly, do you want to append any other data to enrich what the AI model is doing? And a good example might be weather data which comes from outside of the company. But you bring it in and then you pull all of that together into a context window effectively, and your AI points at that. But it is an important point that you don't let the AI model just look at everything.
You want to try and put a rope of scope around the AI, if you like, to make sure that it's looking at the right stuff and that it's kite-marked and you trust it.
John: I think of AI doing things really quickly, right, quicker than maybe a person could do. Does that mean I need to have real-time data as well? Am I kind of running to keep up to feed good quality data?
Chris: No. So, I mean, in some cases real-time data is very useful. Okay. A good example of that might be in preventive maintenance. Think of an oil rig. Think of spinning equipment on an oil rig. And it's spitting out information all of the time.
Now capturing that real time is very useful because then AI is getting the data happening right at the minute, and it can understand how to prevent a failure. But 99% of the time, real-time data is not required and it's quite expensive to get to it.
John: I mean, you talked about kite-marking, I think we've talked about cleaning data. What goes into that? What is kite-marked clean data?
Chris: Well, if you think of data, and I'll give you an example of one particular type of data, let's say customer data. Okay. So you're a commercial director in the business. You're looking at your customers all of the time. That customer data is made up of many attributes about you as a customer. And they're trying to understand more and more about you. So the attributes grow and grow as you go. Right. So they've got your name and they've got your address and your age and the products you've bought and products you're likely to buy.
And you're building this picture of a person all the time. So first of all, we need to verify that you're a real person. And we can use various techniques to do that. We can check your email address. We can check against different data sources. So if we can validate you're a real person, we know that it's an accurate record.
And then we need to understand how many of the fields that we're collecting about you are empty. And if there's lots of them, that's probably not a great quality data set. So we try and enrich that data with as much data as we possibly can to learn more about you. Right. So that's enrichment and completeness of data.
And then the next one might be timeliness. So when did you last update that record and check it? Right. And it's basically a checklist, John. So we go through that checklist and we say at this point in time we're happy that we've done everything we can to make sure that that customer record is accurate. So let's kite-mark that and let's put that in the pool that AI can look at.
John: You've been talking about "we" do this. Who is the "we" who's going and kite-marking? Because this sounds like quite a mammoth operation.
Chris: It is, it is, it is. And there's a role within many organisations called a data steward. And they have the unfortunate task, but big opportunity, to get this right for organisations.
Right. So these people get up every single day and they love the data that they're working with. Right. And they typically are functional experts. So they could be finance folks or they could be procurement folks. And they're looking after the data that sits within their particular domain: finance data, procurement data, etc. And those are the people that have designed the standard for that data.
And they're also checking against that standard. And they'll use multiple tools like data quality tools, etc. But they're tasked with making sure that data is as accurate as possible. And in some cases, that means talking back with the business and educating people who are entering the data and saying, here's why it's so important you enter it accurately.
John: What sort of tools or processes should organisations be applying to make sure they've got the right data, they've got clean data?
Chris: Yeah. There are many. Some basic ones would be a data catalogue, and it does what it says on the tin. It's a catalogue of all your data. So you understand where it is, which systems it exists in and who owns it, etc.
Data quality tooling, which will take your standards and will measure your data at points in time through its lifecycle to make sure that inaccurate data hasn't crept in, and the data steward will be looking at that all of the time.
And then wrapped around all of that, you typically have some data governance tools. So if somebody wants to make a change to a standard or add a new attribute to data, the governance tool is where they request that and there's a workflow created. It comes to the data owner and steward. Can we do this? Yep, that's a good idea. Bang, let’s do that. And so they’re all stitched together, those tools.
And then the last one, which is useful for what you asked earlier on, John, in terms of how do you know which data to point at, is data lineage. So data lineage is a way to look at, let's call it, "from report to source". So where does your data appear in the report and how does it flow through your system all the way back to where it was created?
And that tooling allows you to draw a visual map of that. So you can tell if you affect some data here, everything that it affects on the right-hand side upstream in the organisation.
John: We're talking about preparing data for AI. But I'm just wondering if AI is one of the tools that I’m using to help me with my data stewarding.
Chris: Yes, indeed it is. It is indeed. So what AI can do pretty effectively is it can take inputs like data standards, process maps, architectural diagrams, etc. and it can determine, or try to determine, what data exists in your organisation, what standards you have to measure it by, where it exists. So it can actually create for you the first draft of your catalogue.
It can create for you a first draft of a lineage map. It can pull all of the standards in, and it can start to measure data against those standards. So it's quite new in the market, but it's definitely something that's taking off. And the second point of this is AI can be useful to clean data as well.
So if we can feed AI what we believe to be good quality data and train it on that, and you see something that's different, it will flag that.
John: At what point does the human come in there then if you've got the AI creating the lineage, the AI checking the data that another AI has sourced to say it wants to use?
Chris: Yeah. So, well, we still believe very strongly in human in the loop. And that data steward is that human in the loop.
Right. So what we don't want AI doing at the moment is making any changes to source system data. It needs to stop at that point with recommendations that the data steward can then approve, and only after approval does the change actually take place.
John: One of the things that I think hits the news about AI and data is around bias. So how do you guard against that? How do you govern that?
Chris: Well, of course, it's not easy to do because typically you're pointing AI at data that exists already today within your organisation. So you've not built a data set all the time that is fresh and ready to go. Right. So you're pulling in existing data.
But I think you very much have to look at the use case. And it's use case by use case. What is it you're trying to drive as an output from that particular AI? And if, for example, you were using AI, and I know that many people don't do this, but if you were using AI to do some recruitment, okay, you know the outcome might be you want to sift down to the best five resumes that you can possibly find.
You need to make sure that the data set that you're using therefore doesn't bias that to just white middle-aged men, right, for example. And so you look at the data set based on the use case and govern it that way. Because if you try and look at the whole elephant, you'll never be able to eat that one, right, it's too big.
John: We've talked about identifying the data you need. We've talked about governance. We've talked about the role of the data steward. I'm just wondering, where does this data sit and how do I control what the AI has access to?
Chris: So common practice, and good practice, is to extract that data that you're trying to use out of source systems, which could be everywhere in the organisation, and place it into a data platform that's controlled and governed.
And you typically take that data through multiple stages of cleaning and kite-marking, right. So medallion-type architecture is common where you bring it into a bronze layer. It's raw data. You just pull it in and then you start to clean it up and harmonise it and create a data set that you trust, ready to be consumed at the gold level.
So it sits there and you point AI at particular places. You could, for example, point AI at a SharePoint page with some unstructured data in and it will only look at that. You could just point it at the data platform area. But that, I would say, is good practice. Not just let it loose on everything in the organisation.
Now, interestingly, things like Copilot typically have access to everything that you can see because it's based on role-based access, right. So you are largely looking at an ungoverned data set at that point, which is why it's very important to put yourself back in the loop as a human and double check and validate everything that Copilot spits out.
John: So you just mentioned medallion in there. Yeah. For anyone like me who's unfamiliar with what that is, is that a provider or a particular approach?
Chris: It's not. It's an approach. Think of the Olympics, where they're giving the medals out and you've got people on a podium: gold, silver, bronze. And that's exactly what it is. It's stages in a data platform where you pull data into it.
And as I said, the lowest stage, bronze, you typically pull raw data in. So that's where you're pulling data from source systems and you put it in one place. And then you start to clean that data and make sure that it meets your standards, etc. And then you pull it into the silver layer, and then you normalise that data in a way that can be consumed by AI models, and then it goes into the gold layer.
So it's merely a staging layer in a data platform.
John: I think I'm getting the nod to wrap up here. We always end with the same question, which is basically, what's your one top tip or bit of advice to our listeners?
Chris: My top tip is break the problem down, right. Data is too big of an elephant to eat within any organisation, even small ones. Break the problem down into value cases. That's where you start. You don't do anything unless it drives value in the business. So go and find your value case.
Work your way back to the data that feeds that. Fix that first.
John: Well thanks very much for joining us today, Chris.
Chris: Pleasure.
John: Hope you've enjoyed the episode. If so, please do like and subscribe on Apple Podcasts, YouTube or Spotify, and hope to see you again on The Insight in 15.
Get in touch
Discover why organisations across the UK trust KPMG to make the difference and how we can help you to do the same.