📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Why AI systems are biased | Dan Bogdanov, Liina Kamm

Cybernetica AS58:47

Transcription

So, welcome, welcome to the Science Will Make You Happy podcast. We are back again this time, uh, approaching the holiday season, and we have a new, interesting topic to bring to you. We are visiting the topic of AI again because that gets people excited still. And today, we will be talking about AI bias. And, uh, with me today again, discussing AI again, we have, uh, Lina Gum, senior researcher at Cybernetica.

>> Hi, nice to be here again.

>> Of course, uh, we got good feedback from our AI discussion for the last time. So that's why we had to, uh, come with a sequel. Also, we have new results. So,

>> Yes. So, there's more exciting things to tell you, right? So, AI, people sort of think they know what this is, despite usually the most of these definitions being wrong. But bias, bias is a more complicated word. So, usually when people try to understand AI bias, then, uh, it intuitively means that the computer is making decisions that favor a certain group or are more, um, bad, are worse for a certain group.

>> Or better for a certain group.

>> Or better for a certain group. Very much depends on the application you're building. So, the famous example that you find on social media that gets people all bothered is usually when you have a language which has no gender, like Estonian. And then you take the sentence, um, "someone is an engineer" and "someone is a teacher." And, uh, then if you take it in Estonian, there is no he or she. It's the same. It's "dema." Uh, however, when you then machine translate it to English, then the translation would be, "He is an engineer and she is a teacher." And then people get all excited about that and say that the AI is biased. Uh, is that, is that the case?

Oh, well, um, so initially, when this, uh, when these translate, uh, translation engines came around, um, it actually was this way because they were trained on, uh, on text, trained on languages. And, uh, and in those cases, it, it is the system will be biased based on what the texts mostly say. Now, actually, with this Google Translate, um, example, um, if you ask Google Translate, um, in the, uh, other order, like, "They are a teacher," and the Estonian non-gendered pronoun "dema," or "are a teacher," and "they are, uh, an engineer in Estonian," then it will translate it to, "She, he is a teacher and she is an engineer." So, I think in Google Translate, it has probably had this criticism of, uh, that it is biased, uh, um, internally. So, it seems that now it is using the, uh, from an ungendered, uh, pronoun language to a gender pronoun language. It is taking the male pronoun first, and then the female pronoun second. So, I think it comes down to you either then, uh, uh, do it in a systematic way. You basically say that first, I will use the male pronoun or the female pronoun, whichever you choose, and then second, I will use this. So, it might come out that even if you're talking about one person in one language, and then in the translation, it will have the first, uh, section as using a male pronoun, and then the second section using a female pronoun. And this is, but this is something you can do with translation engines because you can basically tell it, tell it that, uh, here is a rule that you're supposed to follow. Now, the other way of doing it is actually trying to use some kind of statistics. But this is where bias comes in.

>> Mhm.

>> Because historically, unfortunately, this is what is going on in the, uh, society. So, we have more, we have had more engineers who are men, and more women or more teachers who are women. So, naturally, uh, if we're using maybe a large language model, then the large language model will be biased in a way because mostly when you're talking about a teacher, the pronoun will be "she." If you're talking about an engineer, well, nowadays, more it is also "she," but it used to be, um, male mostly.

>> Mhm.

>> So, so you get this kind of, um, bias, um, which is not malicious bias, but it is still bias, and it is there because that's how the, that's what the data shows. But this, we don't want it to be this way. We would prefer that the system didn't make this choice for us because if it, if it doesn't know, then it shouldn't assume this.

>> Oh, that's a very good point. Yeah. So, I also saw a very similar case on LinkedIn recently where, uh, a female CEO had machine-translated text suggesting that she might be a "he" as well, in a very similar situation. So, that is a choice that you have to make. But we'll come back to those choices later on, because it's not a trivial thing if you try to engineer a system like this when you have to decide how you tackle. All right. So, but the, the history, historical thing is, uh, kind of important to understand. So, if we look, um, if you look back, yes, it has been like this. And the whole question of, should we now, uh, why do we need to take this seriously now? Is it because these AI systems might be making decisions for a lot of people? Is it because it scales so well? Or what's the reason why we should think and talk about this bias more?

So, with large language models, then they are, they are becoming more, um, they're going more into chatbots, and people are communicating with them. And, uh, whereas we know that this bias exists in society, we would like it to not be that way. We would like there to be more female engineers. We would like love it if there were more male teachers. So, but if you give this, um, if you have this kind of, um, uh, chatbot which doesn't have its own opinion, it has only data, then it will still communicate it as an opinion. And if this, this, so this kind of perpetuates the the cycle. So, it basically, if a person asks what they should be learning, they are going to be, you know, there is going to be bias in there because we cannot control it anymore. So, it comes from the data. I, I cannot say that there is one way of mitigating it. I think we will get back to the mitigation measures. But, but the fact is that even though there is historical bias, and we are making changes, then it's, it's very, very important to to see what kind of data goes into the system when training a model.

>> Mhm.

>> And there's actually during data generation, there are, um, so Sures and Gut have this, uh, kind of, um, uh, wonderful, um, uh, classification of different meth, different bias, uh, types. Um, they basically define eight different types of bias. So, there are three types of bias during data generation, and then,

>> So, when is data generation, you mean like when, uh, the way how humans have lived for all this time?

>> Data generation, in essence, is a process that kind of consists of the, so, so these, these bias basically consists of the, what the data is in itself. So, so basically, all of the historical bias that that the data set will have, or that actually the world has. And then, uh, there is, uh, bias during, uh, during the, um, during when you make the actual selection of data, and then, uh, the measurement bias.

>> Mhm.

>> So, so you have, um, so that, that you, you basically have the, the bias inherent in history. And then you have the bias when you, um, when you select people into your, uh, data set, because you might not get to all of the groups that represent, uh, the, uh, the different, um, viewpoints or different people. Because if there are fewer people, uh, in that group, then it is more, uh, less likely that they will end up in the, in your selection, in your target group, or in your, uh, training set. And some of the groups, especially the most vulnerable groups, who might not have access to different, um, technical measures, or might not, or the people collecting the data might not have access to, uh, those people at all, they might not end up in the group either. So, it is usually the, um, the, the elderly, the more, uh, poor groups, um, the, uh, ethnic minorities. So, so these groups tend to not show up in the data sets that we collect. It's also, so there are also people who are data altruists, which is a great thing, because a lot of people are just, you know, donating their data for science. But those people also have a certain kind of, um, background, a certain kind of mentality. And, and their over-representation in the groups might also bias the data set.

>> And then, uh, we, and the third one is the measurement bias, which is already when you have gathered these data. So, how do you classify them? So, if you label the data, do you, because that's also a person doing it, or a machine doing it, that you might go wrong there somehow? Or there might be actual physical measurement bias, like, uh, for instance, and this is like a weird mistake, but you, some people, uh, might, for instance, measure something in, uh, inches, and then, and other people in centimeters. And then you get a total like, and, and if the, actually, the, um, it allows for this kind of error, then you might, it might not be an actual data set, right? Or, or you might have something like, the similar kind of, um, um, of basically just measuring in a different ways.

>> Okay. So, taking us into, uh, zooming out a little bit. If I want, it used to be statistics. Before we had AI, we had statistics. And statisticians have always known about stratified samples and things like, you have to have proper representation of your key groups and so on. So, I guess, um, now that the AI revolution is trying to get all the data about everybody, then it's a lot of historical biases. Now, it's maybe there's more of it there because they try to get everything that we have. So, if, let's, let's take a simpler example, trying to develop some sort of an AI health tool. So, I take all the data available. And the historical bias would be that, um, there's in a population, there's usually similar amount of males and females. Some societies have more males because of reasons, and some may have more females. But technically, biologically, there is a sort of a balance there. Now, the question is, exactly, and the next step was, uh, after historical, there was, uh, not measuring, but what is the second one was,

>> Sample bias.

>> Sample bias. Yes. So, if you're starting a study, for example, for a drug or something, then you might just, you put out an advert, and maybe mostly men come to a trial or something. And that would be,

>> Well, that's, I mean, there you have another kind of bias, because it's far easier to do, uh, drug studies on men, because men have a, a very stable hormone cycle.

>> Mhm.

>> So, they're any time in the month you measure them, it's the same.

>> With women, the cycle is a wave. So, depending, and the, and like we're not all going like in the same wave, right? So, you might have a month of study, and then you might have women in different, uh, uh, areas of the cycle. So, it is more difficult to actually do these studies on women. And this is actually, I mean, this is a very, uh, very good bad example of bias, uh, because a lot of the drugs don't work on women because they are so much designed to men, and they, that these kind of hormonal, uh, differences have not been studied in women. So, this has been actually shown in, in, uh, in research as well. Uh, but, but this is actually where your, uh, clinical, or sorry, sample bias comes in, because yes, if you only accept men, or you accept like, I don't know, uh, 20 men and five women, then the women might become outliers, and they might be like the group that the drug didn't work on. But since it worked on most of the people, it might still be okay. So, so let's say, let's not, maybe it's an overstatement that no drugs don't work on women, but they work in a different way because of in which part of the cycle the woman is in currently.

>> And, yes, here's all the modern science on pharmacogenetics and how to understand which, I don't know, antidepressants work well on people with which kind of genomes. There's still a lot we don't know, and the protocols are getting better. But, uh, we have to still, uh, discuss this. Now, it used to be that it was the medical scientists who did a lot of the work on data, and they used to have protocols, and a lot of them. But now we're putting these, uh, data-driven solutions in the hands of everybody. We're putting them in schools, we're putting them in hospitals, we're putting them in hiring. And, uh, this is creating a potential problem where not all of these people have the same training as classical statisticians in the math departments, for example. So, now that is one of the reasons why we have to, uh, start talking about this more.

All right. So, uh, there, history is history. We can't change that. And the question is, of course, there's a whole discussion to be had on how much we should be changing history, and how much debiasing, or, uh, we can't have everything equal, most likely, in many cases.

>> Well, yes. And, and in some cases, it doesn't actually make sense. Because again, if you are making, I don't know, a treatment for ovarian cancer, then you're going, your data set is going, or sample has has to be biased towards people who actually have ovaries, right? It, you cannot, uh, you cannot use, it doesn't, it doesn't make sense to use data. So, your data set will be biased. You won't have other people in there. And, and again, with certain populations, so if you're researching elderly, doesn't make sense to include all other populations as well. So, in that sense, those models will be biased. And, and they will be biased in a positive way. They will be, and maybe biased is not the right word, but they will be targeted at a specific group.

>> But then that's where the other kind of deployment bias comes in. If you're using this kind of system targeted to a group on another group.

>> But that's, that's already moving to the,

>> Yeah, it's taking the right drug, using the right tool. So, okay, let's, uh, move on a bit. So, AI systems, we understand work like this. You feed them a lot of data, uh, about what humans have done in history, or something we have observed, or so on. And then it learns this, and it starts repeating this, or predicting similar things, generating text in a similar way that humans do, or something like that. So, but, um, where there's more places where bias can come in, after the historical one, and then the sampling bias, and then, so, there's, you start building the system. So, where do we have,

>> Yeah, so, so that's where the next, uh, next four come in, or let's say five, because the overall one is the design bias, which I will get back to. Uh, but the next four are basically, uh, so that you have learning bias, evaluation bias, aggregation bias, and deployment bias. And a lot of these, I know, sound similar because, uh, uh, they're very, they have like smaller technical differences, which in a system are important. But basically, what this means is that, uh, now you have these data, and then you will be training your machine learning model, or your AI system, or your large language model. And, uh, and with that, um, you will have this, um, uh, there will be the d, the bias in the data will influence the, uh, the system as well. So, and, and you have the extra kind of implementation details that if you're using the wrong kind of model for these data, that this model doesn't actually work on these data so well, then you will have a kind of bias coming in from that as well. And, um, and with, you might be using the, the wrong data set. Uh, so, with the, with the eval, evaluation bias, for instance, um, that's a bias where, uh, so because when we train models, right, we want to compare their, how, how good they are, how well they work, and how fast they are. So, we have these, uh, very specific data sets that, uh, that these, um, that we basically measure everything for. So, we take the, because if you have different algorithms working, um, and you, you have this one data set, and you show in the, in the same data set, what kind of results you have with this data set, then basically, you're, you can compare the different models, right?

>> But if your data set is meant for something different, and you, uh, basically try to get it to be really good in this table, then you're optimizing it for the wrong thing.

>> Mhm.

>> Because you, and then that might also introduce extra bias, because you're not doing the right thing with it. So, when, when you're a data scientist or a machine learning engineer, you have to, there might be, okay, I might pick, al, I might pick an algorithm, or, okay, so the tooling is important. There might be tools might have their own sort of biases, which might not be, um, the tool doesn't know about history. So, the tool is more of a, it's mathematics, just works like that. It's like case like this, yeah. So, okay, if I take one kind of machine learning algorithm that, uh, converts pictures into signals, u, like, okay, there's this other famous case, um, which is that, uh, soap dispensers in restrooms. And there were, there was a case that it worked for, when you put your hands under there, it gives you some soap, and it worked for Caucasian people, but not so much for,

>> Uh, people from, uh, with an African, uh, descent.

>> So, so that was also considered a racial bias, but it might have been a problem in the whole machine vision algorithm in the end.

>> Yeah, it probably, probably if it could have been, yes, machine vision, but it also could have been that the data, like when they were training, then the people who were training it were actually of Caucasian descent.

>> Mhm.

>> So, and that's why I think also when in the beginning, um, you had these, um, speech recognition as well.

>> So, so that when, when a woman spoke, [laughter] it was that the machine, uh, the voice recognition didn't understand what the woman was saying. But when the woman started speaking with a very low voice, then the m, then the voice recognition understood what the what they were think, what the person was saying. So, that, that is, but that goes, I think, more into the, not so much yet the technical part, rather than your actual representation or your training set.

>> I think that's more there, rather than in the technical part. So, the technical part is, I would say, more technical. Technical part is technical. But, but basically, you, you're more in the way of like, if you, you're using the technology wrong, in a wrong way.

>> Mhm.

>> So, I think the, the representation still goes into the, the data generation bias. And then the technical part is where you have, for instance, deployment bias, is that you're, you're using your algorithm for, for the wrong task.

>> For an application that it shouldn't be using.

>> Yes. So, these days, everybody who, if you're building an AI-based tool, then don't pick the first algorithm or machine learning model you find on Hugging Face. Uh, try to see if it is actually, uh, fitting for the use case you're building, because otherwise, you might get it seriously wrong.

>> Yes. So, there is, um, there is a, there's a lot of that new disciplines coming in, and we are getting better at that, of course. But at the same time, we're rolling these systems out in a bit of a, um, like there is no tomorrow, in some cases, uh, I would even say. So, I, um, just saw an interesting paper, and I have to, like, again, zoom out a little bit. So, how you make these ChatGPTs and these big chatbot systems these days, you have underneath these large language models, which are trained off a large collective of human knowledge in text form. Okay. And, but after that, you train these foundational models, and these are the, uh, big ones that, uh, OpenAI and Meta and, uh, Google make. And then comes the fine-tuning phase, where you actually, people start taking, uh, these, these models, and then tune it to certain applications. For example, to make a chatbot out of a foundational model, there's quite a lot of work to do. So, read a paper. The paper basically was saying that the way how these fine-tuned models are rolling out right now is, they are sort of suppressing local culture, because they are fine-tuned very much on culture that was available to these people making the fine-tuning.

>> So, um, for example, uh, um, US shopping behavior is also being deployed to people using it for education in India.

>> Yeah.

>> And this will, is it also bias? Or,

>> Yeah. Yeah. This is exactly this is deployment bias.

>> Okay.

>> So, so you're training it on one, u, data set, or for one purpose, and then you're using it for something completely different.

>> Mhm.

>> So, so this is, uh, and, and if you fine-tune, um, again, the large language models, they are large, and they are very capable. So, in those cases, you might be able to fine-tune them for a specific purpose.

>> But if you have a model that is really good at, I don't know, identifying drones from pictures, then using it for identifying trains,

>> is not going to be as effective. And this is a very, like, extreme example, but I think it's the exactly the same, uh, with that if you, if you train on, I don't know, shopping data, um, I, I guess if you train on shopping data in the US, and you use it for shopping data in somewhere else, then, then that's going to be more of representation bias, because then you don't have those, those data. But if you're using it for, um, something completely different, like, I don't know, uh, traffic management, or you're expecting that system to, to analyze traffic management, then it's going to be a problem.

>> Agreed. Didn't even so much mean like shopping data, but people's approach to consuming things, like, how do people behave in certain communities. So, uh, yeah, that was also a great example you made. I thought that, uh, the paper tried to explain that it's sort of, LLMs and chatbots are trans, uh, transmitting culture from one part of the world, world to another. And, um, again, I have been reading quite a bit on AI this year. Apparently, these, u, this newspaper article in The Guardian, about kids who are asking a lot of things from AI that they used to ask their parents.

>> Yeah.

>> So, uh, "How do I clean up vomit from a kitchen sink?" or "How do I behave so that this, uh, person of the, uh, opposite sex likes me more?" or "What should I do today to have a better career in five years?" So, these kinds of decisions are now increasingly influenced by AI. So, and these large language models and their fine-tunes. So, it's, uh, it's a kind of a, the scale of this is a little bit scary. So, there's many people using them more and more, and it's sort of transmitting a value system and the modes of behavior, which is, uh, I guess the bias in that is something to be worried. I think we can call this deployment bias. But, uh, but, but in that case, you know, because it's the system is not meant for that, uh, even though the people are asking those questions of the system, then there is always like, um, okay, well, you're using it wrong. And I've heard this, uh, comment as well, being deployed in cases where, um, you know, bad things have happened because of the use of, uh, chatbots. And then, "Oh, but in our, in our terms and conditions, we have this line where you're not supposed to use it this way." So, so the, I mean, uh, and again, this is a liability question. Uh, but it's again, the people are bound to ask questions that they're really not meant to ask from a certain chatbot. I'm pretty sure that if a person is chatting with, I don't know, a legal aid or, or like a, um, um, whatever, consultant, financial consultant, then at some point, people still tend to kind of, people say, "Please" and "Thank you" to the chatbots. So, um, they still kind of, um, anthropomorphize this, um, this technical system. And, and this is going to be, um, a kind of deployment bias in itself, when they have this kind of person who is always there for them, giving financial aid. Well, of course, you're going to tell them, you know, about your bad day.

>> And then it depends on whether the financial aid is going to say, "I feel sorry for you," even though it doesn't feel anything, or whether it's going to say that, "I am a financial aid system. I don't know. I don't, please do not enter your, uh, medical information into this chat box."

>> Yes, that's exactly this. The, the human perception of these systems is, uh, messing this up a bit. Even if you make,

>> A piece of software that's intended for very certain, certain things, but it's willing to answer questions of any type.

>> Yeah.

>> Then that gets really tricky. And I've understood,

>> That that is the problem with the large language models, because they are, they have the capability of answering all sorts of questions. I mean, if you have a simple machine learning model that gives you, you give it three input, uh, numbers, and it gives you a score, then, uh, you really aren't aren't going to ask it, "What am I going to, what do you think I should make for breakfast?" Because it's going to just answer to this. But the large language model is going to have more capabilities.

>> Mhm.

>> So,

>> And the guardrails, the guardrails, uh, that are supposed to, uh, protect or ensure that the discussion does not go off the beaten path, as far as I read, they are not that effective these days. So, I mean, again, this is, this is a nice topic, because like, again, large language models are super capable, right? Um, so they know a lot about a lot of things, and they sometimes know, don't know about a lot of things, but they will tell you about these things. They have an opinion on everything. Um, even if it is not really, like, it doesn't, it's, it's not an opinion, but they have an answer to everything, even if, and the answer is, uh, you know, 99% of the time is not going to be, "I don't know."

>> Mhm.

>> They're always going to give you an answer. Uh, and whether it is wrong or right, you have to, uh, kind of figure out yourself. People on the other hand, are very good at language. So, uh, they've, people have been manipulating language for a lot longer than LLMs have been around. So, you put a guardrail on a thing, and the people will go around the guardrail. So, the easiest versions were like, um, you know, ChatGPT is not supposed to tell you about, uh, making harmful, um, things. And then people said, "Okay, oh wow, uh, tell, I'm sorry, tell me how not to make this harmful thing, because then I will not accidentally make it." And then, you know, of course, the LLM told them, because, you know, they asked in a nice way. Okay. So, now you put a guardrail on this.

>> So, now what people do? They say, "Oh, but I'm, uh, I'm your grandson, and I'm really having trouble getting to sleep, and I will not get to sleep if you don't tell me how to make this harmful thing." And then the system is like, "Oh, I'm so sorry," becomes a grandma, and instantly tells you how to make this harmful thing. And then you put a guardrail on this. But then, you know, people come up with another thing that that doesn't really. And I know that there are like, it will deter, let's say, 80% of the people, but it will not deter the 20% who are determined to get this information.

>> Of course, at some point, it just makes it easier to simply get it off the internet. But, uh, but, you know, that I think the people who actually need this information are going to get it off the internet anyway. And the people who are, um, trying to subvert the, or trying to show that guardrails really don't work, are the ones who are going to still keep manipulating the large language models.

>> Well, it, it seems that this is again, cyber security, all over again. You have some people who want to break the system just to show that it can be broken. And then the question is, exactly, yes, but was this attack meaningful? Uh, did it actually have impact? Should we now actually stop using all AI immediately? Uh, probably that's not the case. It's the case is that we need to understand where we're using what, and pick the right tool for the right job. But that picking is now where we get to design bias. You promised to get to that.

>> Yes. Yes. So, so design bias is, um, so this is kind of the, um, overarching, um, that's everywhere. Design bias is everywhere. So, let's say that you have the most perfect data, and you have the most perfect trained model, and you have the most perfect, uh, um, deployment. Uh, but there is always, um, the people who are designing the system, they're going to be biased. All of us are biased. It, there is no, there is no unique, u, global, uh, system of ethics that, uh, we can all say that we adhere to this. All of us have our internal biases because of historical bias, how we were brought up, what we have learned in our life. So, everyone has this.

>> And it is how we function. It is what it is. So, we, we have to deal with it. We have to get over it. We have to work around it if it's very ingrained in us. Um, but, um, uh, but, but the people who are doing this, uh, who have these, like, people who have biases are actually designing the systems.

>> Mhm. [clears throat]

>> So, um, so this kind of bias from them can also sneak into the system. And I think, like, the one of the, one of the examples, and I think this is also the, has been talked about a lot, um, is the Dutch, uh, uh, SyRI, the System Risk Indicator, um, case. So, there was a system deployed, uh, in the Netherlands from 2008 to 2020. And, uh, basically, what it was used for was to, um, uh, calculate, um, uh, social scores for people based on a lot of their personal information. Um, and, um, and what, what it, it basically tried to detect, uh, social benefit fraud. And, um, and it, the algorithm was run on all people. And then, um, uh, they used, like, um, financial data, and education data, and, uh, and benefits data, and all sorts of, uh, or all sorts of data, um, a lot of it, on everyone. And, and then, um, a lot of the people were falsely accused of fraud. And a lot of the people who were falsely accused of fraud were in the vulnerable group categories. So, again, ethnic minorities, less well-off people. So, um, and, and why this comes about is, of course, that, uh, when you're talking about social benefit fraud, then, uh, that social benefits are also usually, or social, let's say, subsidies are usually given to people who are less well-off. So, naturally, the, the people who are going to commit fraud there are going to be of these groups mostly. But it doesn't mean that all of them are in these groups. So, um, but the system was, um, uh, was biased. And it's, it's looked upon as one of the worst, you know, uh, bias cases in AI. Um, but, uh, when you look into it, you find that it wasn't actually a machine learning model that was in there. It was rule-based. So, yes, under the AI Act, it does actually fall under the AI system category, because it's, uh, it's a kind of, it was a kind of decision tree, but the decision tree was designed by people.

>> Mhm.

>> And, um, you see that, I said that it was, it was around for two, from 2008 till 2020. That's 12 years. It was redesigned, I think, twice during that time. Uh, before in 20, 2020, then it was basically ordered by court, by European Court of Justice, to be taken down, because, uh, because it was basically said that it, it is illegal to use that system. Um, but this, but the questions there, so, the one of the issues is that the system was, um, was, um, uh, kind of, uh, opaque. So, you couldn't see what the questions were. Which, of course, on one hand, is a, is a good design choice, because if you know what the questions are, then you're going to optimize against this. So, you will never end up in the, in the wrongly accused state, or in the rightly accused data set either. Because again, it's like the, the criminals who will be watching out for that. Uh, so, that I think that's why the decision choice was made. But when you look at some of the questions, uh, which have become public. So, it was like, "Is your apartment, uh, on the second floor, uh, of a building where in the first floor, there is a bar?" So, do you live,

>> Like,

>> Above a bar?

>> Above a bar. So, in Tartu, in, uh, a lot of the, uh, old town buildings.

>> I wonder if any of the, have any of the dormitories have, uh,

>> Cantinas or bars underneath there. So, that would immediately make all the students.

>> Or if your car was older than like seven years, I think.

>> Or, I think some specific car, um, makers as well, were like included there. So, I mean, these questions are clearly biased.

>> Mhm.

>> And of course, this is like, not the only question. But, and, and of course, this all, I'm assuming again, that it's, um, it was based on some kind of previous information.

>> Mhm.

>> But this previous information was still designed or taken and designed by a person.

>> Mhm.

>> So, um, and then a lot of the people who were denied the money, uh, they did not have basically any income. And some of them were actually sued, u, and asked for the money back. So, it, it takes time to exonerate yourself. It takes time to prove that you didn't actually do this thing. But for that time, you don't have any money.

>> Mhm.

>> And, uh, some of them lost custody of their kids because of this. And that is a problem for the kids. And I think there were also a couple of suicides.

>> So, so with this system, you see that, um, that that there is that basically this is, I think, an example of design bias. [clears throat]

>> Because you don't, there is historical bias in there, but it is basically a designed system. So, humans decided based on historical data what rules would be used in the system to find potentially fraudulent individuals.

>> Yes. But again, um,

>> I'm not sure how again, how historically correct those were.

>> The data might have been there. But the question is then, uh, what will happen? So, in a system, human oversight is sometimes a thing. It helps in many of these cases. You have a human in the loop. They say, and the human says, "Okay, this does not look like a case. We shall not prosecute this." So, there's automation. It's, it's a question of how automated everything is. I guess.

>> Uh, so, in this case, um, I think, well, first of all, I, you, I, I love human oversight. I think it's like a perfect, you know, get, get out of jail free card, because you say that the, the machine is not making decisions, the, the human is making the decision. But if you, if your human, um, is going to be the one saying that, "Okay, I checked this decision, I checked this decision, I checked this decision," uh, then, um, how, um, how kind of, uh, vigilant is the human going to be when, um, uh, when they have been doing this for 99 times, and it's been always correct?

>> The human becomes a machine.

>> The human kind of loses a bit their agency, a bit, because they become this like, check the box machine. Uh, and if the system is correct, um, 99% of the time, then the human will start trusting the system more.

>> So, the human will do less to check all of the data. And I guess in some cases, it makes sense, because you're using the machine to actually, um, simplify your tasks.

>> Mhm.

>> But then the human is going to get complacent and not look at the things that well anymore. So, they're going to be like, "Okay, it's usually correct. Uh, I will look at it. I will glance at it. I will not go into the details or into the, um, let's say, the sources that the system provides." Which also brings [clears throat] us to the whole economic concept around AI systems, which is to remove humans mostly from the loop.

>> But, but I think with SyRI, you had another problem, because with SyRI, you also had human oversight, of course, because it was, it was not deciding for, you know, the government. It was giving you a list of addresses that you should, uh, check out. But I think there, the problem was that the system was actually run in one place, and then to the other place that was supposed to do all of the investigation later, they were only supposed to send the, uh, kind of the guilty ones. So, they were supposed to take out the, uh, the false positives, kind of. But then, I think, uh, the person who might have been checking on the algorithm side wouldn't, um, probably did not understand that they were the one who was supposed to be the kind of human oversight, and they sent it to the other, the investigative party. And the investigative party was like, "Okay, but we're only getting the ones that, uh, that are the guilty ones." So, you see that there is kind of a, a communication jump, where one person expects the other person to also check it, but the other person expects that the previous person already checked it very well. So, so you get this like human, human communication error, instead of having this, uh, kind of machine. So, you, you're, and there's always someone else who you're kind of, um, you know, um, getting support from, and then, uh, you're hoping that they did their work, or you're expecting that they did their work, and then,

>> You have this, you know,

>> A human communication error looks like the whole history of mankind.

>> Yes. Yes. I mean, it doesn't make it better, but, uh, but, uh, I wouldn't, I wouldn't say that it was only, uh, like, but, but this is again, this is where we say that, okay, we have human oversight, but then, uh, humans also might not always know what they're supposed to be doing, or who's, who, who, what, what are exactly their tasks, or who is going to be responsible. This, this is a, this is a case with AI anyway, like with all sorts of AI, like who is, who is liable, who is responsible for this decision? Is it the machine? Is it the,

>> So, meaningful, uh, feedback loops, where you, if you get outcomes, and you check the outcomes, and it turns out that some of the outcomes should have been different, then you feed them back in. So, again, I guess that gives us already two mitigations. We discussed human in the loop. We discussed, uh, meaningful feedback cycles, mechanisms to improve your AI system. But, so, u, okay, a lot of this case, this SyRI case looks very much a public sector, government. We are servicing the whole nation, where the whole country, >> population scale systems. And of course, if we are talking about very big, uh, tech companies who create software for the whole planet, pretty much, you have users everywhere. Then that's also, they have their safety teams. They have to have an answer where somebody asks, "Do you have bias?" and so on. So, but, let's now come down. So, let's take a smaller organization. If, uh, anyone today wants to build, u, an AI system, then what is there to do? How do you, uh, what would I have to do? If I'm an organization, I've decided to put out some AI-based services, so that I could check with a peaceful heart, the box, "I've done something to mitigate AI bias." So, I think, um, as with when we talked last, we were talking about having a risk-based approach, which is also a thing that the AI Act, uh, encourages you to do. And, you know, the higher risk your system is, the more compulsory it becomes. But, I always say that it's always good to, uh, do this risk management, um, even if it is not, uh, required by law. And it doesn't need to be a rigorous, uh, a very complex, systemat, systematized, yes, but, uh, a very complex risk management process. But you just follow the steps of a risk management process. So, basically, what you do is the first step is to, um, to, um, to establish the context. So, basically, you describe your system, you get together all of the information, and you write it down. And then all of the information will be in one place. So, you will have a very good place to reference to. And then you will, uh, do, um, uh, the, um, uh, the part where you identify the risk, analyze, analyze the risks, and evaluate the risks. [clears throat]

>> And then, so this is the risk assessment step. And then you will do the risk treatment step, where you decide what you do with, uh, risks.

>> Mhm.

>> So, you have context establishment, you have risk assessment, and you have risk treatment. And, uh, with bias, I think bias can be very well switched into this process. So, you have a couple of, um, you know, a couple, when I say couple, it's more like, you know, hundred, uh, other questions about, u, bias as well. So, when you are looking at different risks, you will also look at the risks that bias has. You will also evaluate the risks that bias has. So, we'll find out what is the likelihood of this happening. What is the impact when this, uh, risk materializes? And then you can decide what to do with this risk, whether to, you know, not use this kind of system at all, or not use this kind of functionality in your system. So, basically, you, basic, just cut it out, or you do something to mitigate it. You do the human oversight, or you do, um, you know, you try to get a more balanced data set, which is not simple, but you can do it. But this, there is no like, um, I cannot give you a mitigation measure that, you know, fixes bias issues for you, because it's a very system by system case.

>> Because some of the things like, a lot of the times the bias is there for one, if you were using the system for one purpose, um, in the same data set, and in another data set, there, there the bias is not there, because the data set was exactly meant for that purpose. And then this is, this is for instance, with the ovarian cancer thing, that you have a data set of women. It's, it's okay for this task, but not okay for, you know, uh, making a, a overview of the whole population, because you only have women in your data set.

>> I would imagine that, u, this overall, this generic risk analysis, risk,

>> Management.

>> Management process, it works, uh, it needs to, we understand what the bad thing is. So, we need to define who is impacted by this. So, they sort of potentially, uh, the group of people or organizations. It could also be organizations, I guess, in some services, would get hit if there is bias in the system. So, we would need to, uh, understand who can get hit if the computer says no, or computer says yes. Yeah. And then understand what happens to those people if the computer says no, but should have said yes. So, that's the sort of the thinking, uh, that shouldn't be very hard.

So the main problem I guess is uh, people don't fully the the knowledge about this AI bias is not so well common so well known. So that's something that we're fixing also with this uh show a bit. Uh, so can you make this um, can you make this simpler? So uh, small enterprise or uh, small government office can uh, take a tool and then do something and be done for like 75% of the cases.

I'm glad you asked. [laughter] No. So um uh, so this year uh, we did um, a study uh, for the Ministry of uh, Justice and Digital Affairs as part of the Equititech project. And um um, and we put out a methodology and guideline for uh, AI bias risk management and it's AI and algorithmic system. So it's exactly it applies to the Siri case. It would apply to the sir case as well. So um um, so they are also interested in detecting bias in systems and uh, and with this report so it uh, it consists of three parts.

So we have the uh, the methodology which uh, basically describes the steps that need to be done. Then there is the guideline which talks about bias in a more um, like what bias is. It also talks about the different types of bias and how the bias gets into the system and uh, and then what to do with the bias also. And then I think the part that the the uh, that is kind of a tool for smaller organizations and government uh, organizations as well. Uh, there is um, a worksheet or a workbook let's say. So right now it's in Excel form because uh, again it has a lot of pages and and Excel is very dynamic that way. But you can I think you can use it in in an open uh, software as well but I think also that the ministry is going to make it um, a kind of a web- based uh, a part of their web- based toolkit. So you can actually answer the questions there instead of having an Excel worksheet and it will do some of the things automatically because we did not include macros because you know uh, downloading something from a web and and having macros in it is going to be uh, a lot of red flags for a lot of organizations. So um, so we didn't do that. Uh, so but but yes but these things these um, reports uh, will become available very soon. Uh, and then you can basically uh, download that worksheet.

So the workbook I think it has um, it is an easy way of doing this. It has um, uh, prompting questions. It has some examples in it. Um, so you will know what kind of questions to ask and um, and what what to basically look for because I know that a lot of we can say that yes of course these this is biased against these people and this is what can happen. Uh, but unfortunately we cannot see all of the bias and uh, when we have these examples and when we have these prompting questions um, then we can actually look at it um, we can understand that yes this is is in fact bias. So we might agree with it we might just not come up with it. I think it's the same way with all risk management and risk you know risk identification because I might not think of this risk whereas um, you know um, another colleague of mine might be very um, sensitive to this kind of risk whereas and I might be so so we complement each other uh, we might both have you know unique inputs into this so so in this also same with bias you you don't have to have like a special bias person but uh, or you know special expert on bias But you can coming together um, you can also look at the questions and >> your uh, committee for deciding what is bias shouldn't be biased. [laughter] >> Yes. >> Thank you. Yes. Unbiased committee for detecting bias. >> Yes.

So this is obviously very hard thing to do. Uh, all right. So you have your people in the room. You may have a you you may have somebody who understands law. You might understands have somebody that understands societal aspects. technical aspects maybe a analyst of some sorts maybe a manager who can then make the decisions but okay you find that there is uh, bias in the system so is it um, is all know I I'm trying to ask the question is all bias bad or is all bias uh uh, does all bias warrant investments and change of systems or how do we understand what do we actually need to rush about or what's the way. [sighs]

I think um, I think there is bias that you definitely need to look at and and again it is um, so this is where risk management does help because again with with every risk you have as I said you have the likelihood and the impact. So the example is that um, some of the things are more likely to occur in certain areas or in certain places and the impact is that how bad is it if this occurs. >> Mhm. >> So you might have something that if it occurs it is really bad but it will almost never occur. >> Mhm. >> And then uh, it is like it will always occur but if it happens then nothing bad happens. So you have these kind of extremes there. uh, so it is not a clearcut of like yes to this no to this. So you have to kind of score it and you have to see which which risks you take care of and which ones you don't and this I think becomes an executive decision. Um, but but again with um, let's say with bias so some of some of the things with bias like for instance if we look at the sirrus system then um, you know deaths because the the uh, detection of fraud system was wrong I don't think that is acceptable where and this is like uh, this is the uh uh, this is like the where the death count should be zero >> if we're talking about deaths in traffic Then if we get it down to, you know, 10 per year, >> I can't even know what the right now numbers in Estonia are. They might be less. I hope they're less. >> Less than 10 per year. I don't think >> I didn't look it up. Uh, we'll have to look it up. >> So, but but but again um, I think at some point the deaths in traffic um, we still have to accept that there are going to be some casualties there. So again, you have to look at this uh, like the system and the context as well. So there is no again as I said unfortunately there is no clearcut case of like um, this is okay this is not okay. Um, it is unfortunately how it is but but again we try to take it to a minimum. We try to mitigate it as much as possible. In some systems if it is unacceptable you just cannot take the system into use. With Siri it shouldn't have been with the mistakes it made and and it was as I said it was redesigned. So they had problems and they redesigned it and they put in more rules and more more guard rails if you will and it still didn't work. So at some point you're going to actually have to kind of cut your losses and say that okay we need to do this differently because this is not acceptable anymore. >> Mhm.

And but the scoring gives you some sort of an ordering I guess that which ones are more significant and which one you should like tackle first and which ones you can tackle in the next budget period. >> Yes. And there are different methodologies for doing this scoring. So you have uh uh, you have different ways of uh, of kind of making this order of of what should be dealt with first. But I think there are there are ones that you can deal with um, you know immediately and others that you can then deal with yes afterwards. So I guess in the next budget year because you have you will have the things that are um, expensive but immediately fixable and then you have the the you know the um, the problems that will be super super expensive to like cost as cost as much as the system costs to just fix this one uh, small hole. So you will have to decide whether um, whether you will accept the risk the residual risk or whether you will not use the system until you get it fixed or whether you will just you know pay for it and >> fix the risk but there some things there you just cannot fix everything because there will always be some residual risk. The risk will never be zero. So >> uh >> it's not an easy task to fix. It's there is no easy like okay use this so you're done. And in some cases it might actually be you can fix the little things. In in United States I read there was um, a settlement in court with a company who had a hiring form on their website and this hiring form had a very uh, simple algorithm behind it. There was if your age was above a certain number then your CV was immediately rejected. So that would be age bias I guess. Yes. And then there was also a peculiarity that for women I think the age was a few years lower. So that might be another kind of bias. So uh, and they had to settle for damages to a lot of people who were discriminated against like this. So uh, don't do things like that I guess is a good take-home message. So >> yeah but I think in in the US they have a lot of liability insurance. So there you just you know you give this uh, to someone else and this is >> ah that's also a risk mitigation strategy. >> It is so you share the risk >> share the risk >> or transfer the risk >> transfer the risk to the insurance company that's also a way so in the end it becomes a matter of culture and how do you deal with new technologies as you build them. So, uh, we've made our, uh, I hopefully we've made our impact today and we've helped some people reduce the risk of uh, having uh, an unacceptable level of AI bias in their systems. So, thank you very much, Lena. It's been amazing as well. Uh, until next time, and u, may science continue to make you happy.