📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

New Methods, Old Problems: Ethics and Bias in Modern Natural Language Processing

John Snow Labs – Healthcare AI Company30:54

Transcription

Hi, my name is Ben Batorsky. I'm really excited to be giving this talk on ethics and bias in natural language processing. This is a topic that's really important to me, and I think, really interesting. There's a lot been a lot of developments in sort of research around this, and hopefully, I hope to share some of that with you and hopefully have some conversation in the chat as you as you watch this.

So, just a little bit about me. Uh, so I'm a data scientist. I've spent most of my career focused on NLP. Uh, I have my PhD in policy analysis, and I'm part of the SciAxs Health Real World Data team. So, just a little bit about SciAxs Health. It's a health information information management company, basically, uh, taking, uh, data from a number of US hospitals, um, and then sort of providing that, providing insights for researchers and for kind of development around healthcare.

A series of sort of NLP production pipelines. I'm, I'm not going to be talking particularly about the work that I'm doing, but I think it's really important for anybody who's sort of using NLP to think about these. And these are definitely things that are kind of on the table as we as we, uh, do this work.

So, uh, I'm gonna talk a little bit about history. I think history is a good place to start. Uh, I'm using "old days" like, you know, end quotes, very flexibly, just saying that this is sort of, I'm bucketing machine learning into like pre-neural net days and sort of post-neural net days. Um, and not to say that these techniques are not applicable or useful anymore, or continue to be used, or of more relevance more now than ever. But I just want to sort of talk about, like, what is the sort of recent developments that have driven a kind of, um, need for re-examination of, uh, the priorities of NLP.

And so, just on the left here, we have an example of a very kind of old machine translation system. These old systems were sort of state machines. They were parsers that would, uh, follow kind of a deterministic path based on, uh, what, uh, what it was seeing and what the state of the parser was. Um, and so these, you know, worked for a very specific set of, um, rules. If it fell into the rule set, that was actually pretty good. But otherwise, it wasn't, uh, very good and wasn't generally good enough to sort of be used in general cases. A lot of machine learning besides that was sort of required a lot of hand engineering, a lot of domain expertise, but also on the other side, a lot of transparency, right? A lot of work went into sort of this side of machine learning, or continues to go into the side of machine learning, of actually generating these, uh, these features, creating features that can then run into a model and have parameters that are learned towards a certain objective.

So, you know, this is a really important side of machine learning still. But, uh, what neural nets, one of the things that neural nets sort of offered was the ability to kind of learn, uh, representations, particularly, uh, when it comes to language, where it needs to essentially create some sort of informative representation of individual tokens in sort of a series of text. And so here on the left here, I have the kind of that earlier sort of language models. And very briefly, a language model is essentially just predicting a word based on context, and that's sort of the objective of a language model. And, you know, they came before the 2000s, but there was sort of a proliferation of them around that. Um, and so one of the, uh, sort of computationally intensive parts of it was learning these representations of the individual tokens. You can see here in the support of matrix C, where the individual tokens are sort of being represented as some, you know, vector of information. And those learned representations were super useful for that particular task, for the language model task. But if you wanted to use them just generally as sort of a representation of words, they were not as useful. And in training these language models, you sort of had to sort of train that from scratch, which took a lot of time and just generally required a lot of data to be able to create those sort of reliable representations.

So, one of the major, uh, benefits or the major, major developments in 2013, the call of paper talking about, uh, word embeddings, were a simple algorithm for creating embeddings for individual words based on their context. So, by feeding it sort of a large corpora of, uh, language data, just sort of free text, you can learn representations of individual terms that have something that have embedded in them some sense of the context of the terms. And that context was sort of shown to have some, some representation of meaning. So, the example of being able to take these representations and sort of add them together and subtract them. The, the canonical example of taking "king" subtracting the vector for "man" and adding the vector for "women" ends up with a vector that's fairly close to "queen." So, that's the idea that if there's some meanings kind of embedded into these representations. Uh, but, you know, these representations, the individual dimensions don't have a sort of readily interpretable meaning, uh, on them. So, it's a little bit different from the other kind of framework that I kind of set up in a very simple way. And of course, I'm sort of obviously getting a lot of detail here, but I want to sort of get at the point that these representations now are very useful for sort of plugging into other applications, creating features that are based on these individual terms. But they also require a lot of, uh, text to be able to create representations that are kind of sensible. Um, but, you know, they're, it's a really, uh, awesome and impressive, uh, piece of technology and method. And so they've proliferated, proliferated all kind of all over. Um, so there's medical terminology to vec, which are embeddings messed up based on medical records, which are really interesting. And then a ton of different word embeddings for different languages. Of course, still, you need language resources to be able to train a word embedding that has any sort of makes any sort of sense. And there's certain limitations, particularly of the, you know, the word to vec algorithms as proposed, that makes it a little bit less flexible for other sort of types of languages, other sort of structures. But I'm not really going to dig into that, but, um, it's worth noting.

So, now, kind of jumping ahead, everything is now like BERT, everything is now transformers. These large neural language models, uh, have become the state of the art, uh, across a bunch of different tasks. Um, and, uh, on the left here, I just have some of the, one of the translation tasks and sort of the state of the art there. All of these are neural models, and you can see the ones that are kind of starting to beat out the others are these transformers. Uh, you know, this was the original transformers paper. Um, but if we look at just recent days, these models are getting larger and larger with more and more parameters and requiring larger and larger data sets. So, that's an important thing to to note, that they have to consume more and more data. So, to be able to have one of these models, you have to have data available. And that means you have to sort of, uh, pull in, um, uh, the, the data that, that this sort of model requires. Um, but, uh, what this, the upshot of this is, is we're actually seeing, um, really, like, powerful results. So, like, this is Google Translate and showing the transition from, and this is in 2016, right? So, this is, uh, already kind of outdated. But the transition from phrase-based translation, which was before the neural models, uh, to this, uh, neural model translation technique. And you can see the performance. You know, like, the performance doesn't look like it's a huge jump that has happened. But if you actually look at the output, you can read these and see that the more that the, the phrase-based machine translation, which is like somewhat closer to that old sort of 1950s example than it is to modern technology, is a little bit, you know, you can sort of get the idea, but it's not really like readable. It's kind of a little bit garbled. But the neural models almost seem similar to the human. Of course, these are cherry-picked examples, but that's sort of what we get out of it.

But I mean, thinking about the amount of, uh, data that goes into this, we have to think about what are the sources that are being pulled in. Um, and we, in that case, we have to think of the principle of this "garbage in, garbage out." And I kind of like the point to this example. This was a Microsoft in 2016 released a Twitter, uh, bot that essentially was supposed to develop and grow based on its conversations with the internet generally. You'd think a company like Microsoft would know what the internet is composed of. And of course, the bot was immediately targeted by people that were trying to turn it into something, uh, that would say numerous problematic things, and they were successful, right? So, uh, it makes sense, right? Like, if you put a bunch of, uh, bias, problematic statements into a model, it's going to learn those, that bias, and learn those problematic statements. So, it's worth examining with these larger models and these large, these methods that were sort of, um, sort of proliferating everywhere, what are the sources that are being pulled in?

So, uh, uh, for word to vec, the original paper was trained on Google News. BERT and GPT-3 pull from Wikipedia, just these large, uh, uh, data sets of language. GPT-3 even goes more ambitious, pulls a bunch of different things from Common Crawl, has its own sort of inclusion-exclusion criteria, which have their own sort of, uh, caveats there. Common Crawl including a bunch of different web sources, things like Reddit and other, and other things that are available, right? Because they need to consume all this data to be able to achieve the performance that they're able to get to. But again, what is the content? What are the perspectives being represented in these different data sets? And that's sort of, you can get some kind of easy stats that would show you that there's maybe some problems with just kind of pulling in as much data as possible.

And so, if you think about Wikipedia, a survey of this all-language Wikipedia showed that just the vast majority of editors are male. That there is a vast under-representation of women among biographies, and that even among that small percentage of biographies, a lot of them get recommended for, uh, deletion, which would suggest that there is a general sort of, um, attitude or general sort of perspective that's being represented in Wikipedia. Not to put any judgment on what that perspective is, but it's to say that with that, it was worth being conscious of. In Google News, you know, this is not, this is kind of old news, but newsrooms tend to be dominated by men. There tends to be an over-representation of white men, particularly. Men are also more often the face of articles across different topics and outlets. So, if we're pulling in that, we have a certain perspective being represented there. Internet users, you know, there's a little bit more of a mix of like, who is being represented. But when we talk about Reddit, with Reddit, does a great job of making a lot of its data available, but who are Reddit users? Mostly male, mostly white, right? So, these are showing that there is a particular set of perspectives that are being represented in the data that's being fed to this, um, to these models.

And so, you might ask, all right, well, so what? Right? Like, what is it, what does it matter? What is the impact of something like that? And, uh, one example, one kind of powerful example of this is this, uh, this investigation that ProPublica did of this, uh, risk-scoring model that, based on a person's record, would score whether the person was likely to reoffend, you know, when they were being sentenced. Um, and they found that, uh, Black defendants were much more likely to have higher risk scores and to be predicted to reoffend, even when, uh, a comparable white defendant who, you know, did reoffend or had a history of arrest, um, uh, when they were compared to that. Um, and so you have to think about, like, what is this model trained on? We're not actually sure. Like, I can't say that this is, you know, a BERT model or something like that. And that's almost the problem, right? We don't actually know what was the data that was feeding into this model. We don't know what features they're using because it's, you know, a private company, and they're not going to release that information, which makes it even more problematic. What is the data that's being, uh, is being trained on? It's, you know, you can't say that there's particularly something intrinsic to different groups of people that makes them more likely or less likely to do these different things. There's, if you're training on historical data, you're going to pull in all the historical biases that exist there. And this, uh, ref, this can be also pulled in, uh, from GPT-2. If you use GPT-2 as sort of a language generation and you give it a prompt that has some sort of gender or race, uh, information in it, it generates pretty problematic passages to just kind of show, demonstrate that it has ingested the biases of the data that it's been trained on. Um, so this is what I'm talking about when I'm talking about these, your, like, potential impacts. And, you know, uh, whether generated text is going to affect someone's life, it's, it's not clear. But there are a number of other examples of, uh, the kind of latest models kind of failing the test, uh, of being fair or, uh, having, uh, bias being considered.

And so, I'm gonna step back a little bit, right? And talk about, like, what am I, what do I mean when I say ethics and bias? And there's a ton of different definitions, right? I'm going to focus on, like, a very, like, narrow subset, right? Because it's not a long talk. Could, you know, there's a lot of people that are talking very, uh, carefully about all of this stuff. I'm going to focus on, like, one particular aspect of it. So, ethics is like, sort of a set of moral principles. You can read more about, like, this kind of principle day III, which has a bunch of different, uh, elements to it. I'm going to focus mainly on fairness and non-discrimination and the design in favor of inclusivity. Bias, uh, just generally, is a disproportionate weight in favor or against a particular, uh, thing or idea. Um, so again, that's the sort of to the point that there's nothing, uh, particularly intrinsic about the particular group of people that makes them more likely to, you know, uh, behave a certain kind of way. That it's just, if you train on historical data, you're gonna have all of the, the, the systems that existed historically, that would sort of, um, portray that, that idea that there's different groups have have these different qualities.

So, I'm gonna talk a little bit about how to address, uh, issues of bias, how to, you know, promote, um, some, some examples of it, and then some ways of sort of addressing it. So, uh, here's just a couple of types or a handful of types of bias in machine learning and this machine learning generally. I recommend you, you take a look at this paper. It has a ton more examples, and they're really interesting and well-articulated. But some of the ones that I'm mainly going to be talking about is essentially historical bias and representation bias. So, you can think of historical bias as that, you know, sort of compass example that I showed before, uh, that because the incarceration rates of different populations are, uh, by a product of sort of the institutions, uh, that exist currently and existed in the past, uh, we're going to sort of bake that into any model that's going to be trained on historical data. Representation, uh, bias, uh, you know, I have the example here of ImageNet. You know, has certain types of people doing certain activities. So, if it sees another type of person, uh, doing an activity that's not stereotyping, it's, it's going to have a harder time, uh, uh, any model this training then is going to have a harder time recognizing that. We're going to assume a particular type of person is doing a particular type of activity. And so that's kind of a problem. Um, so those are two major ones.

And I'm going to talk about, I like to focus on one particular example. I think this is a really good example. We're in embeddings and showing the gender and racial bias, uh, that's sort of inherent and then very easy to kind of pull this out and take a look at this. Um, so on the left here, I have this sort of representation of the gender bias. On the y-axis, you have the word "she," and on the x-axis, you have the word "he." And you have sort of the projection, the similarity, uh, of, uh, these different, uh, uh, occupation terms on each of these, uh, different gendered words. And, you know, ideally, you want to see that they're the same, right? That there's nothing intrinsic about women that makes them more or less likely to be an architect. Uh, but there's something historically about that. So, you know, you can see that "homemaker" has a higher loading on "she" and a lower loading on "he." So, you can see that kind of bias there. The same is true for racial ethnic bias. Again, you'd expect to see that the top occupations that are similar to these, uh, ethnic groups, uh, would be about the same. But you can see that's not the case, right? You can see that there's a bias that's kind of baked into it.

This is kind of a busy slide, but I thought it was actually really fascinating that you can sort of trace history in word embeddings. You can actually see that the amount of the bias, you know, in, uh, towards, uh, towards the women, the, the similarity to gendered terms, is actually, uh, correlates pretty strongly with the occupational difference. And so you can see that there's more women in, uh, nursing, and the bias is very high. So, you can see that kind of correlation. And this visual on the right is kind of interesting because you can see that if you train your word embeddings on, uh, previous sort of generations, prior to kind of women's, the height of the of women's movements in the '60s and '70s, you can see that there's a lot more correlation in the adjectives being used to refer to women, uh, before those movements than after. And that's kind of interesting just to show that word embeddings are encoding the historical bias, which is just kind of fascinating. This paper is really interesting, I recommend it. Um, and you can also play around with this yourself. I have a notebook that allows you to just kind of mess with this a little bit. You can reproduce that, kind of results that we saw before pretty quickly. Vincent Varmadam from Rasa has a package that's called "What Lies," which allows you to sort of explore word embeddings and do, you know, similar explorations like this, which I find really interesting as well.

So, now transitioning a little bit, let's talk a little bit more about fairness and what I mean when I'm talking about this sort of ethical use, principle use of AI, and particularly with the focus on fairness. So, Cynthia Dwork, finding fairness, really interesting talk, talked a little bit about some definitions of fairness and some, like, easy ways to kind of like look at what is fair and what might not be fair. Group-level fairness, the statistical parity idea, where, um, you have an underlying distribution of the population, and you expect to have a similar sort of distribution, uh, in your positive negative class. Say, you're trying to predict who will repay a loan. If you see in one class that one group is overrepresented significantly, then you might want to think, like, "Okay, maybe the model is learning something I don't want it to learn." Same is true for individual-level fairness. You have two individuals, and if they're similar, they should have a similar outcome. If they don't, the model might be learning some feature, learning too much from some feature, or, you know, essentially just encoding some kind of stereotype.

So, let's take a look at that in the context of NLP, right? And so, this was an interesting, uh, uh, survey that the ACL did of, basically, all the link, like, a large set of languages and what were the resources that were available, labeled and unlabeled data. And so, you can see that there's that Group 5, and this is sort of their grouping, this they're clustering of this Group 5 is kind of where you want all languages to be, where it has a ton of data available. But really, it's only representing a very small, uh, fraction of the of languages. And about 90% of languages have basically no labeled data available. Um, so we can already see that this is kind of not representative and not particularly fair to a lot of speakers. Only about one-third of all speakers are being, uh, represented, are being well represented. You can see sort of the languages that the languages that you'd expect are sort of, sort of there. So, you can see that that's kind of problematic. And it's likely not going to change much when you see where sort of the, uh, labelers sort of exist, or the people that are doing this work sort of exist. Um, uh, much more representation of a particular part of the world, and more representation of those languages that we saw in that sort of higher group. And so this location, Mechanical Turk labelers. Um, and so what does this mean, right? Like, maybe you're not compelled by the fact that it's just not fair to not represent large swaths of the people, but it, you could also say that it's going to hurt, uh, our ability to create models that are robust across different languages or have similar performance across different languages.

This is from, uh, BERT Lang, it's sort of a list of a number of different BERT models, trained on different languages, and their performance on a set of different metrics. You can see this variation. Even though this is kind of a mismatch of different data sets, you can see that the performance varies, and that the Ember, the multilingual BERT, the model that's trained to do multiple languages, does poor, much worse than these single-language models. And so, uh, you know, it, it just goes to show that, uh, if you're interested in the performance and the, the top-performing models, um, it makes sense to have this sort of better representation. And I thought these two points, made these two quotes, made these points very strong. Um, so Sebastian Ruder's article on pushing people to work on languages beyond English, really powerful statement, that these, this technology is not accessible if it's only available for, say, English speakers with the standard action sent, or just that, that group of the one-third of the population. And also, we can't say that the models that the systems that we're designing are representative if, uh, we just have this kind of echo chamber of, uh, different of languages, languages that are particularly simply similar. So, one, we need to have better representation so that our models are more robust and see more linguistic phenomena, and we can be a little bit more sure that it's learning something, uh, about language generally, rather than just, uh, fitting to a specific set of patterns that exist in the well-represented languages.

So, now I'm going to talk a little bit about some strategies. And this is by no means comprehensive. There's a ton of work on this. Uh, I just thought that some examples might be useful. Um, uh, one, and this is particularly in prediction. So, you might want to have multi-accuracy targets. You, so break down your targets by these different groups. As we talked about before, you have some underlying, uh, distribution. You might want to see, uh, how well your model, uh, works for, uh, each different group. And so you're not just aggregating it all together. But, you know, it's kind of problematic because how do you determine which groups? A lot of the groups that have, that are at risk, that are vulnerable groups, are probably going to have a little limited representation. Think of the example of loans. A lot of groups just don't apply for loans, and so there's not a lot of data there. Affirmative action is an interesting approach where essentially you, uh, do a ranking within these different subgroups. Say, you're trying to pick, again, with these loans, you know, you pick within these different groups, you take the top 5% of the ranked probability to, uh, pay back the loan. You might do something similar to the example that we saw before with that individual-level fairness, where you have the model essentially predict the outcome, uh, on a pair of individuals that have similar traits, but, you know, you might be wanting to control for something like gender or race. Um, but all this kind of a little bit falls apart when you start thinking about intersectionality, where, uh, you know, the white people and Black people as a group are not monolithic, right? There's, you know, white women and Black women. There's, you know, white trans women and things like that. So, once you start breaking it down into those pieces, it basically almost becomes impossible to ensure fairness across all sections. Um, so, you know, there's only so much progress you can make to addressing this. Doesn't make it not worth making, uh, making the effort, but, uh, it's, it's just to show that there's this is sort of an unending, uh, pursuit.

Here are some additional methods for monitoring and addressing, uh, bias. StereoSet, uh, rates, uh, different large language models essentially according to how much a sort of stereotype they're embedding in them. It's a really useful methodology that you can take a look at. Dion from DrivenData has a really interesting ethics checklist that sort of covers, has like some questions you can ask yourself over, uh, sort of all stages of development. Interesting sources there. Um, so now I'm going to talk about a very specific example. We're going to go back to that, uh, bias word embeddings. There have been some methods proposed to essentially de-bias word embeddings, which is a really interesting set of, uh, techniques. Um, and one of those techniques is essentially, uh, creating a, uh, extracting the information, the, the sort of gender vector, right? So, you take pairs of gendered words, "she" and "he," "girl" and "boy," and you essentially subtract them for each from each other, and then you get, uh, as a result, this sort of, uh, representation of, of gender, right? And then you can see, um, you can essentially remove that, the projection of gender on things like, on occupation, things like "babysitter," "doctor," where you want to make sure that they're sort of neutral and move those to that center. And then the next step is to ensure that the distance between them and their neighborhood of gendered terms, sort of equal distance. So, you're basically imposing the structure that you want on the embedding space, which is a really interesting approach.

And once we apply it, we're done, right? We've fixed all bias. Bias is now gone forever. I'm kind of unfairly calling out this article because I think it makes a very, very strong, kind of ambitious statement that I don't particularly feel is is warranted, where it says that, "it's easier to correct bias algorithms. You just have to find the right data. You just have to apply the right data and observe, make the right observations." And those same things seem particularly hard for me. This is from an author who worked on a study to essentially de-bias this healthcare allocation algorithm. And it's an interesting study, and their methodology is really interesting. And it does seem like they removed a lot of the bias in one direction, but it seems like there's opportunities for bias to emerge in other ways that potentially they did not examine. Um, and also, you know, this is a collaboration of researchers with a private company who was already selling this algorithm. Whether or not that company took their recommendations, actually changed their algorithm, is not clear. Um, and I'll get back to that point in a minute, but I want to like talk about the word embedding is the de-biasing. And this is going to be a kind of busy slide, but to the main upshot here is that this de-biasing method may have done some work, but not, uh, totally remove the problem.

Here on the x-axis, axis, you have this original bias measure, where if it's more positive, means it's more biased towards men, and more negative means it's more biased towards women. And then on the y-axis, you have the number of male-associated neighbors, neighbors that are particularly, uh, associated with the male, um, uh, the particularly biased towards men. And you can see the original, uh, embeddings had this sort of, like, hot, like, highly, they were highly coordinated. Those two dimensions were highly correlated. The more bias it is, the higher it's on that bias dimension, the more male neighbors it has. And then the de-biased looks almost the same. Um, so the correlation, uh, is reduced to some extent. So, there was some work being done. But, uh, the main point of this paper is that the bias is coded in all of these different words. So, you can correct as many words as you want. It's still the neighborhoods of these different words have some information about the bias in them, just because you're kind of feeding it biased data. Um, and so that's not to say, like, give up, right? That's, but I think there's a lot of questions that are worth being asked. And I have some ideas for them. You know, and don't take this as, uh, you know, but particularly the be-all, end-all. There's a bunch of work that's being done here, and I have some resources at the end I can share. Um, but, uh, think about what is the right data versus the right now data. Um, you know, just because, uh, data like Reddit is makes its data very available, doesn't mean it doesn't mean that you should ignore what Reddit is embedding in terms of the, the content that its users produce. Um, Dion's checklist offers a lot of really powerful questions around sort of that scoping, uh, step. Um, thinking about what is the right method. StereoSet is a really useful thing if you're sort of considering a large language model, uh, but it's worth considering whether a large language model is particularly necessary. There's a lot of complexity and obfuscation that goes into those. So, you know, definitely it needs to be used with caution. And, uh, you might want to think about who might be helped and who might be harmed, and how you might be able to ameliorate that.

And finally, I want to kind of propose like a pretty, like, large question. Going back to that example of the COMPAS methodology, where it predicted higher risk for certain groups of people for, uh, reoffending. It's worth thinking, is that a place that we want AI in? And we do we want a system that we can only make some amount of progress towards removing bias, to be inserted into things that can really jeopardize people's, like, liberty and and freedom? Um, and so it's worth making that question. And I think that question has been raised, um, uh, recently. I, I don't have too much more time, but I just want to recommend reading the "Stochastic Parrots" paper by Bender and Gebru, which talks about these kind of big questions for NLP. And I think they're really important now to to think about. And I think it's, it's on us as practitioners to sort of think about these things.

So, that's it. Um, thank you very much. Uh, this is, I'm really excited to be giving this talk, a topic that is really important to me. Um, I have some resources here. Um, uh, hopefully, we will be having some conversations in the chat, but feel free to get in touch. And enjoy the rest of the conference. Thank you.