📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

The Importance of Verified Data to AI Systems

IBISWorld25:11

Transcription

[music] [music] Hi, I'm Matt Murphy. I'm the senior vice president of client advocacy at Ibisworld. And today I'm joined by Sam Johnson. Sam is the manager of Ibisworld's data team. And today we're going to be talking about the fact that companies these days are trying to take all the different data sources that they have and put them in more centralized locations or AI tools that make that data more easily accessible for their employees.

Uh Sam, can you start by maybe just telling us a little bit about yourself and I guess also thanks for joining me today?

>> No worries. Pleasure to be chatting with you. I'm excited to of what we're going to cover. Um as you said, my name is Sam Johnson. I'm managing the data team at Ibisworld. We've been here for a little over 9 years. Um the last few in the data team and uh it's a definitely a very exciting time to be in this space. There's a lot of change both for us but more generally as well. So uh yeah.

>> Yeah. Yeah, like you said, you generally speaking, you know, uh, companies large and small have access to dozens, if not hundreds and thousands of data sets and data points these days, and they're really trying to make sense of them. And and that often um includes trying to centralize the data um into a third-party system. Um, that could mean including the data in a CRM system, an ERP system, business intelligence tools, uh, maybe proprietary dashboards that the company has built, or like I mentioned earlier, an AI tool. And of course, all of those systems are only as good as the data that they're being fed. So, making sure that the data that your company is ingesting is structured and verified is essential in this process. So before we go any further, Sam, maybe that's a good first question for us to answer is, you know, what is structured data and why is it so important to these AI or integration workflows?

>> Yeah, that's a good starting point. I mean, like like you said, you listed a number of systems and making sure that the data is structured in a certain way so that it can communicate with all of those systems is such a crucial first step. So I guess to define upfront a little bit, structured data is essentially data or information that's organized in a predefined and consistent format. Typically it's just rows and columns in a relational database. And that makes it easy to search through, easy to analyze, and easy to otherwise process. Um, having data in that format makes it easy for other systems to read because it's predictable as I mentioned. Um, so there's a number of ways a lot of approaches to take when you're structuring data. So I'll cover, you know, a couple of the highle ones that we focus on where we add value. And the the first one is using classification systems and taxonomies. So we pull in a ton of different data. The main way that we group it together is by industry. Obviously, that's kind of our core offering and and focus, but we'll also group things by product or by geography or by statistic type and all these different things so that we're using a consistent set of everything so that it's predictable and and tidy. Another one which is maybe a little less exciting but as important if not more important is just accurate and consistent labeling. you [snorts] know, bringing in uh sources that are all disparate and not matching it. You know, you can't overstate the value of having dates in the exact same format so that you know, AI knows exactly how to read it or other systems know exactly how to read it. Other things like percentages and rounding and the units that a dollar figure are expressed in, they all need to be uniform and consistent. And so those are kind of some of the ways that we might approach structuring. Um, in terms of other other characteristics of what structured data is or what it means, um, another good word to remember is that it's machine readable. And what that means is a lot of data, there's obviously a ton of available data out there, but a lot of it is designed for human consumption. Whether that's reading online or it's listening or it's downloading, the data is set up in a way so that a person like us can kind of read it or absorb it in whatever way. Um, but setting it up that way means it's a bit harder for AI or systems to read because it's not quite as um structured as the name implies. [snorts] Um, when you have it designed in that way, less structured and less predictable, it means it let's say let's use AI as an example, it means that there's more room for AI to have to interpret and make assumptions. And when you have AI, you know, generally getting creative and trying to deduce things, that's when you leave more room for introducing errors or hallucinations or AI making incorrect assumptions. So that's one of the another one of the benefits of structure. And I'll just I'll just cover one more point which is is pretty helpful especially again for AI which is that um structured data usually comes with or it should come with a data dictionary which is essentially an accompanying data data set of metadata which is descriptive of what all of the structured data is specifically. So to use an IBIS world example, industry revenue is probably one of, if not our most popular data point. Um, in our structured data, there's a dictionary that comes with it that says this is exactly what industry revenue is. This is how it's measured. And so AI doesn't need to try and interpret uh, you know, a human written query to figure out what that revenue means. It knows exactly what it means. So it knows where to go based on what um, someone's querying.

>> Yeah, I think you used a a good term there and that's desperate. You know, the reason that a lot of companies are trying to integrate this data or deploy AI tools is because right now they might have staff that are accessing um, you know, six or seven sometimes more than 10 different data sources to complete a research project and and centralizing that information is really important um when you're trying to achieve productivity. Um, I mean the desperate means that you know it could be internal data that the company has on hand. It could be an external data source like Ibisworld. Um, really when it comes to the the data sources you know, we like to bucket them into like three core types of data that we see Ibisworld clients using most often and that's customer data that the business has. You know, think if you're a B2B manufacturer, you know, you probably have a lot of data regarding your clients. Um, what industries they fall into, where they're located geographically, um, what their historical pricing has been. Um, that's customer data. Or maybe you're a bank and you've got a lot of loan data and you know, you know, what industries those borrowers are in and you know how many loans you have in a particular industry, you know, what dollar value that equals, you know, risk levels associated with those borrowers. This is data that companies are are storing in data sources already. You also have operational data which is the second kind and that could mean everything from inventory data to supply chain logistics to employee performance. Uh then you have the third type of data where IBISWS world comes in that's market data. So yes, it could be industry intelligence like Ibisworld provides, but it could also mean, you know, customer demographics or it could mean, you know, looking at more economic type of data like uh inflation or interest rate data. So, you're really trying to pull together all these different data sources and and put it in a place that's easy for um whoever's doing that research to digest. And and what does that do for you? Well, first of all, it breaks down research silos. Um, historically different departments are probably using different data sets. Um, if you can centralize that or use AI to break down those barriers, that means that all the departments in your organization are using the same data sets and the same data points that are hopefully up-to-date and verified. Um, that saves you from data overload and having to look at more data than you need to. So if you're using a dashboard and maybe integrating IBIS world with API, you're only delivering to your staff the data that you know that they need and really cutting out the noise. Um, that saves a ton of time which opens up cost efficiencies. So yes, you know, those team members um aren't doing as much resource or logging into as many resources as they were previously because it's all in one place. Um, but it also saves on cost because those employees at the end of the day um they can hopefully spend their time focused on more strategic or creative work for your organization. Um, so you know, these are things that are really important um because at the end of the day, the the process that you need to go through to your point um, you know, it's one where you have to make sure the data is clean, right? You might have a lot of data sources, those data sources need to find their way into a data lake where they might be structured like IVIS world's data, some of that data might be unstructured, so you need to go through an ETL process, which basically stands for extract, transform, and load, so you they can then get that data into to a data warehouse like for instance Snowflake, um, and then deploy those at the front end. And the front end ultimately means those systems that your staff are accessing. So it could be Salesforce as a CRM system, could be a, you know, a business intelligence tool like PowerBI or or an AI tool like Copilot at the end of the day. So there is is a process involved here, but the end results are what give you that time savings and those cost efficiencies at the end of the day. Um, does that sound kind of like how ISIS World, you know, has its own data and and arrives at our own front end? Like what's IVISWorld's process for collecting data and how can our customers ensure that that data is accurate?

>> Yeah, absolutely. Well, the the workflow you've described definitely has a lot of crossover with how we do things as well. Um, you know, step one, for example, we have a lot of a really broad range of sources globally that we bring in and we the first step is we put them all in our data lake. So same as same start as what you described. Um, the core place where we really add value is our structured data warehouse, which is we use snowflake for that. For us, that contains our raw data vault. So we go from our data lake to our raw data vault and then we have a transformations focused data vault as well, which is where we sort of model the data, which I'll get into in a moment. Um, the sources that we bring in, they're like I said, there's several, I think I said there's several hundred that we bring in and that we store. They're largely government data sets, which I for anyone who's worked with government data sets in the past, they can probably attest, there can be pretty frequent changes to the format or the frequency or what's included and what's not included. And so there's a lot of inconsistency and it's often only semistructured or quite unstructured at best. So we need to do a lot of work to clean that to get it into our structured data warehouse in some of the ways that I described before with labeling and taxonomies and stuff like that. So that's the first step that we take. Um, we set up automated flows to pull that data in. But of course, when governments go and change data sources, which is quite frequent, we need to then go revisit those automated flows and make sure that they're still working as we need them to. Um, we use tool we use snowflake as I mentioned, but the tools that we use at the various stages of getting that data in um include S3 for a data lake and glue and terraform and lambdas and things like that. Um, once we have it in our world in our snowflake world, our structured data warehouse, we um, we need to structure it for its end use. So that's the transformations focused data vault that we have, which is our sort of second main stage of the the data warehouse. Um, I guess to to go back to the industry revenue example, if you're getting that figure from a range of different sources, it might be called revenue, it might be called income, it might be called turnover, it might be called sales revenue. You know, it can have a lot of different names. Sometimes that can be the same thing with a bunch of different names. So, we just need to align the name itself and sometimes those can be slightly defined slightly differently. So in those cases, we would need to, you know, transform the figure a little bit to make sure it aligns with how we define industry revenue. So that's some of that structure. And then from there is our IVIS world specific value add that we do, which is filling gaps in the data, running bulk calculations to create new figures that clients can use and forecasting. So those are probably the three main categories of the the value add transformations that we run, which are also in the structured data warehouse. And from there it can plug into wherever it needs to. Sometimes it goes over to our research teams, sometimes to our clients, you know, then it's available for other systems to use essentially once we've finished those transformations. Um, and so a lot of it's very based around organization and tidiness as a first step and then value add as a second step. Um, but it those two together are what make the data what make our data reliably um consistent and accurate.

>> Yeah, we we've used the the term accurate or accuracy a few times and and verified as well. Um, but I think it's really important to note that Ibisworld's uh data is human verified at the end of the day. Can you talk a little bit about the human element that goes into what what might sound to some people as like a very like automated or like you know technology ccentric uh approach to this research?

>> Yes, absolutely. I will actually start because there there are very much two elements to our verification and like you said the human one is the uh the critical part and the part that we absolutely um need and it adds a lot of value at the end. But the first step is also includes some sort of technological verification as well before it goes through to that human verification. So I'll touch on both just to give the full picture of what that looks like. Um, and part of the part of the aim of both, uh, verified information is, uh, I don't know if you've heard the term like garbage in, garbage out in reference to to AI. Um, but it's, you know, it's exactly what it sounds like. If you're feeding unverified information or unreliable information into an AI model, the outputs that you're going to get are going to be low quality as well or less reliable. [snorts] So, um, this is obviously a concern for businesses that are wanting to rely on output from AI for their decision-m or for whatever they need it for. And it's a big concern. Um, actually, there's I have a stat from it was a US Census Bureau survey um, recently that found that 78% of businesses reported using AI in some way in their organization, but only 10% reported using it in a meaningful way. So there's a big gap in using it versus effectively operationalizing it. And so to to overcome some of those issues in our own data is where that verifi these kind of two elements of the verification steps come in. So the first one which is on our side the data team side is the more automated one. There's two main approaches to how we verify. The first is automated tests across everything with hard rules. So if the data doesn't pass those hard rules, then it can't be published. We need to go and fix what's happening. And then there's soft rules. So those can point us towards where there could be areas that we can investigate or potentially improvements to our model to bring things into a more reasonable range. That's the the automated side of things. And then once we have our data finalized in in that way, we will pass it over to research teams. And that those are the ones who can apply the expertise and contextual um critical thinking to the data as well. And so even if we've sort of modeled everything at scale, they know the specific considerations that need to be taken into account for the specific industry that they're looking at and they can sort of either verify or tweak the data or send it back to us and say you need to revise this. And that's that layer that gives the um the real confidence in the in the final data to to where it then goes on to next. So it's not like IBIS World is using AI to write its industry research there. There are real people still behind the numbers and behind the analysis that we're applying to these industries.

>> Yeah, absolutely. All the analysis is is human written. We don't have any uh any AI writing the analysis on the data.

>> Yeah. Because someone you know might say devil's advocate, well, I could just go do my research through chat GPT now. Um, but you know, without having that verified information um, and that up-to-date information, which I think is especially important when it comes to doing industry research, is knowing that you have the most up-to-date data from those government sources or otherwise means that you're avoiding a lot of the hallucinations that that people are seeing. Is is that more or less right?

>> Yeah, well, you you absolutely can go and use chat GPT, but it's sometimes going to give you useful full relevant information, but unfortunately sometimes it will pull either information that's unreliable and unverified or if it can't find the answer to what you're asking, it will hallucinate. So that's obviously one of the key limitations. Um, I'll just refer back to something you mentioned a few minutes ago as well. You were uh, you know, referred to kind of IBISWorld uh clients that integrate IBISWorld data into their systems allows them to kind of select specifically what data is relevant to what that user might need. And with chat GBT or any any, you know, broad large language model, they have access to obviously a huge volume of data and that can be good, but it doesn't necessarily mean that it's going to give you a better answer. So when you're giving it a specific curated piece of information that your business knows is relevant and verified and reliable to what you need, that's going to, you know, improve significantly the quality of the AI outputs that you're using.

>> Yeah. And I I think that's an important distinction, too, because we we've really talked about two different ways to access Ibisworld data through integration. And um, you know, the one way is, you know, using ibis world's integration with with Salesforce, which is right off the shelf. Um, you can use API to call IBISWorld data into dashboards or or other systems and and pick and choose which IBIS world industry data points are the most valuable for your organization and deliver those um to your staff in quick and easy ways. Um, the other way that we're starting to see people use our data is in these um ways that you're using queries. That could mean using co-pilot. That could mean, you know, building your own LLM, for instance, in which case you can start asking questions in ibis world is some of the data that's returning those answers to you. Um, and now all of a sudden, IBISWorld lives in an ecosystem with your internal data and maybe other external data sources and you're able to prompt your AI tool to get the information that you're looking for um in the most efficient way. So some of the questions that we see our clients starting to ask are, you know, what are the current trends in this industry and what challenges are operators facing? You know, um, how many customers do we have in this industry and how can I um bring up three pain points to a prospect that I have in that industry? So starting to use like you said, that market sizing information or the industry revenue size, number of businesses in the industry to then say, okay, what does our current client base look like? You know, who are we already working with in the space? What white space do we have? What market opportunities do we have within a particular industry? Um, maybe you're looking at an individual company or customer and wanting to say, you know, how do this company's financials compare to the industry financial averages and really start to see whether your customers are underperforming or outperforming the industry. Um, you know, these are ways that we're starting to see our clients use ibisw world in new and creative ways um in AI. So um, it's certainly, you know, something that we're at the forefront of. Um, so I think we're going to be learning a lot more about how you can leverage internal customer and operational data alongside industry data. Um, but we've gotten a lot of great feedback from the organizations that have already started deploying things like co-pilot, like large language models to to you know, essentially, you know, embed all their data in one place. So their team members aren't going to those seven or eight different uh data sources when they're doing their research.

>> Yeah, it's a major timesaver.

>> Um, as far as the AI piece, I think we covered that a little bit as well as structured data. Um, are there any other parts about IBIS World's process or um considerations that you know data folks or decision makers at these different businesses need to make when considering the type of data or the way in which they're ingesting this data to get the best results.

>> Um, I think probably the, you know, the the benefits that I've kind of referred to of the structured data for um AI probably apply similarly to uh integrations, whether whether that's through whatever system that might be that clients would use it for um, and that centers around predictability and uniformity of the data. So if if they're um integrating into whatever system because the structured data comes in a predictable structure as the name implies um API users or any any users using our data through data or AI platforms um, they can just incorporate it with a more certainty and less of a need for custom handling because it's got that uniform structure. It's it's faster to develop on their side because they know exactly how it's going to come through every time. Um, and that means fewer edge cases when they're parsing the data as well. Um, and I may have touched on this a little bit, but I'll just mention it again in case I didn't, but the um, compatibility of the structured data because it's got enforced, you know, we we enforce clear data types, whether it's integers and strings and and whatever else, those are strict and uniform as well. So it means that there's fewer data integrity issues, which makes compatibility easier. So whether clients are pulling it into a SQL database or a data warehouse or a BI tool or some of the other tools that you mentioned um, it means that that integration with those third party systems is easier. Um, and I guess the last thing is filtering. You can, you know, because it comes through in the way that it does, it allows clients to filter and sort um like like you mentioned, just pulling the necessary data that they need into the right place where they need to store it or display it for their users makes that a bit easier.

>> Great. So yeah, so I guess in a nutshell, ibisw world, we are the largest provider of industry level data and insights uh globally, which means we, you know, cover thousands of different industries and in dozens of different countries and, you know, this is data that um is not just, you know, uh created by AI or scraped by AI. There is human verified data that IBISWorld is providing um, which is giving that credibility and that accuracy to IBISWorld users. Um, Ibisworld, you know, has direct integrations with uh companies like Salesforce, Snowflake, Microsoft Copilot, for instance. So, you know, starting to integrate this data into third party systems or AI is fairly easy. Um, I think we're only going to continue to have um new and additional partnerships with with other companies. So um, regardless of the type of um, you know, centralization you're doing, whether it's through a, you know, a business intelligence tool or a CRM or an AI tool, you know, come talk to us about it because, you know, we are looking for new and exciting places to deliver our data um, because we want IBIS world industry data at the fingertips of our users. You know, we want to make sure that we are in the places where researchers are spending their time um, so moving forward, I think it's just a really exciting space for us to be Um, if you are someone that is interested in learning more about how you can integrate IBISWorld's data into your workflow, um, please just reach out to your client relationship manager. Um, they would be more than happy to put you in touch with our partnerships team who can work with you strategically to help you understand or help understand, you know, what you're trying to achieve with, uh, data integration or your AI workflow um, and really create a solution that's the best fit for you and your team. Uh, if you're not already subscribed to Ibisworld, you know, I would say go to ibisworld.com. You know, drop us a note in the contact us form. We'd be more than happy to talk to you about how um, not just, you know, data integration and AI can improve your workflow, but how industry research in general can improve your workflow. Um, so Sam, like I said, I think this is something that we're going to continue to be talking about at IBIS World is this technology is ever evolving. Um, but I really appreciate you taking the time to chat with me about it today.

>> No problem. I appreciate it. [music]