📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Exploratory Data Analysis Tutorial | Basics of EDA with Python | Great Learning

Great Learning4:31:49

Transcription

Exploratory data analysis, also known as EDA, is used by data scientists to analyze and investigate data sets and summarize their main characteristics. Employing data visualization methods, it helps determine how best to manipulate data sources to get the answer you need, making it easier for data scientists to discover patterns, spot and normalize, test hypotheses, or check assumptions.

EDA is primarily used to see what data can reveal beyond the formal modeling or hypothesis testing task and provides a better understanding of data set variables and the relationships between them. It can also help determine if the statistical techniques you are considering for data analysis are appropriate or not. So, without any further ado, let's get started right away. [Music]

If you haven't subscribed to our channel yet, I want to request you to hit the subscribe button and turn on the notification bell so that you don't miss out on any new updates or video releases from Great Learning. If you enjoy this video, show us some love and like this video. Knowledge increases by sharing, so make sure you share this video with your friends and colleagues. Make sure to comment on the video for any query or suggestions, and I will respond to your comments. Now, before we start our video, let me very quickly walk you through the agenda of today's course.

In today's course, we are going to learn about the basics of EDA with Python, multivariate analysis, outlier detection, and cricket world cup analysis using exploratory data analysis. Right guys, so when I say exploratory data analysis, what is the first thing that comes into your mind? What is this train of thought, uh, you know, that that you are governed with, that you understand? Right? Let me give you an example. Think of this particular situation. Right? So there's a new TV show that has come out, maybe on Netflix, Amazon Prime, YouTube, Hotstar, wherever it is that you guys watch it, and that now it is super trending in a way where all of your friends are talking about it.

We have multiple shows which did this in the past, saying, "Hey, you know, friends would ask you, saying, 'Hey, did you watch that show? It was very trending and all of that.'" So now your friends are all talking about all of these; it is in the news, it's on the TV, it's on your Instagram, it's on Facebook; it is everywhere. Right? So on the whole, right now you do not know much about the TV show apart from that it exists. What would you do? The first thing is, if you really are interested in it, if you can, you know, actually have time to watch it and all of that, you will be curious to learn more about it, to understand more about it. So you might search on Google, watch a video on YouTube, all of these things. Right? So what are you doing? You're exploring to understand what is the thing at hand here. Right? Again, this is exactly how exploratory data analysis works as well. At the end of the day, you are more and more curious about what you can find in your data. So once you go ahead and sort of tap into that potential, the immense amount of potential that exists in your data, you know, you would be very clear with how things unravel, how things wrap out and give you so much more information than what you could pick up. Right? So this is a very beautiful example, just for that.

To come towards understanding what exploratory data analysis itself is, first of all, you must begin by looking at the data to understand what it is. Now, data can be understood in multiple different ways: on a low level, on a high level, on a mid-tier level, to see what it is that you're getting for your money. Right? Here, in the case of exploratory data analysis, we're trying to investigate; we're trying to look into the data; and we're trying to see, "Hey, is there anything that I can find when I look into this data and eventually have a result which will add meaning to the entire exercise that we are doing?" That is exploratory data analysis. Right? Sometimes it is not, you know, not all the times it is the case that you will only be looking for trends and all of this in your data. Sometimes you'll be performing in-depth exploratory data analysis to find out if there are certain anomalies in your data; you know, if your data is working fine, it is doing things the way it should; is it the other way around? Are there errors in your data which is eventually hurting maybe machine learning or deep learning systems down the line, or is there anything that is going wrong in your data that you can find out by performing some analysis on it? Right? This is exploratory data analysis in one picture, guys.

Next up, you really have to understand why exploratory data analysis is so important, you know, it is the requirement of the hour and all of these things. Right? There are multiple different points; we can talk about this for hours on end, but the first important point you really have to understand is that whenever you're taking a look at data to understand what conclusions you might arrive at, to do it, to take a look at data to understand data means you're already trying to assess and analyze something. Right? That is one important point; it will get you to your goal quicker. The second thing is it is very easy to make assumptions with data, saying, "Hey, you know, the all the rows in my data are fine; there are no errors and all of that." So if you want clarity about these assumptions that you make, if you want facts which back up everything that you say and, of course, what your data shows at the end of the day, it becomes vital that you have a methodology to do it. Guess what? EDA is how you do all of that. The next thing is I mentioned to spot new trends, to know that, you know, "Hey, here is a product at hand; there is a good chance that down the line in the next two months, three months, or three years, this product might be extremely trending, so we have to work on it right now." That summer is approaching; a lot of summer wear, light clothes, sweat-absorbent clothes, and all of these will be on sale. So to know that, if you are a person who's into selling all of these, you will be upping, ramping up on your production scale because you will have such a high demand for it. Right? That is exactly, you know, my point on this particular topic. So it is not just with data; it is not just you stating out something, but your data is packed now with a good amount of proof, saying that, "Hey, this is what I say, and I have backup, and I have proof for all of that." So that is an important thing you have to know about when you're investing your time learning about exploratory data analytics.

Okay, so all of this is amazing. Can we take a look at the types of exploratory data analysis? I'm sure by now your interests are, you know, high; you want to find out more about the domain; you have to, you want to start using it practically; we're going to start using it; don't worry, at the end of this module. Before all of that, there are various types of analysis which can be performed using explorative data analysis, so it is vital that you know all of them. Three important types of exploratory data analysis that exist are univariate analysis, bivariate analysis, and multivariate analysis. Now, what happens with univariate analysis? Uni, as the name suggests, you will be analyzing only one particular variable; you will be using a good amount of data to analyze what is happening to that, you know, that column worth of data. Right? And then when you're taking a look at bivariate analysis, with bivariate analysis, as the name itself suggests, you'll be taking a look at two things; you'll be analyzing two variables at the same time to see what it is, what is the goal, you know, that you're trying to achieve. And then we have multivariate analysis. See, multivariate analysis talks about how, you know, you can go on to perform analysis based on, you know, more than two variables at a time, so it may be three, maybe four, maybe five, maybe ten. Right? Knowing that your result is based on just one column of data, two columns of data, ten columns of data, and the results you get out of it is astonishing compared to knowing that, "Hey, only one row of data can have so much impact." Right? That is why it is very important. Now, quickly take a second to see if you can think of an example for univariate and bivariate. Right? For multivariate, I can give you a good example: if you are trying to teach a machine learning algorithm to see if it can predict if a person has diabetes or not. Diabetes is one of these medical conditions where, you know, it depends on a lot of different things; it depends on age; it depends on a factor of hereditary; it depends on your blood sugar; it depends on your body mass index; it depends on, you know, if you're obese or not; it depends on your weight; it depends on your height. There are a million different variables that will that can be used to assess and say that this person might be diabetic. For example, your sugar levels, your blood sugar levels can be super, super, super high, and that is when you realize that you might have diabetes. Maybe you have a wound which is not healing; that is one of the symptoms of diabetes as well. Right? So there's many different things; that is a classic example of multivariate analysis. So you take a second to think about univariate and bivariate analysis. Right? I hope you all did that because here are some really nice examples for the same.

Multivariate analysis we already discussed, you know, you'll be trying to see if diabetes can be predicted if you have insulin at hand, if you have the age of the person, the body mass index, and all of these things. Next, you know, when you're taking a look at univariate analysis, back to one variable analysis, here if you want to understand how the person's age changes and, you know, maybe how the height changed based on the age. Right? When we were all babies, we weren't a lot tall, and you know, as you grow years and years and years, you can go pretty tall. Right? Six feet, six and a half feet, all these are achievable. So this change of height is dependent on your age. Right? You're using one variable to see if the person, how the height is changing of this person. So that's an important thing to know. And the next thing is bivariate analysis. See, with bivariate analysis, you're trying to understand variables based on two analytic criterias. For example, you know, if you look at multiple sporting events that happen. Right? For example, there's cricket, there's hockey, there is swimming, there's Batman, there's all these sports. If you want to find out what the population prefers playing right now, what the males prefer, what the females prefer playing or watching or whatever it is, you will realize that there is a difference between, you know, there's a difference between what people like as well. To understand what they like, you have to take two variables into account and have one sport and say, "Hey, you know, males usually like playing cricket; maybe females absolutely love playing hockey or bat metal," whatever it is. Right? There was just an hypothetical example, but that is the case. Next, if you're taking a look at your GRE score right now, maybe with more and more and more experience, you know, if you prepare for competitive examinations such as GRE and GMAT and all of these examinations, it could come to a point where, you know, with more experience you can do it better. Right? So here again, your score is dependent on your previous performances; your scores depend on how you learn; and your score is also dependent on your age. So there are two important criteria that you can pick up and perform bivariate analysis, even in a multivariate analysis statement as well, guys. So these are some of the important examples of exploratory data analysis that you guys should know about. And once we talk about all of these examples, we come to a very important step where you have to understand this entity called as box plot.

Box plot, again, we're going to discuss all of this in the data visualization module in a good amount of detail, but for now, you have to understand that it is very, very important you know what a box plot is. Right? For example, if you have concentrated well on the probability module and, of course, the module number two which was statistics, there I have specifically mentioned that, you know, your multiple ways of how you can split the data, work with the data, and all of that. So a box plot is basically a measure; it is a representation of a measure if I have to put it out accurately. Right now, if you have a lot of data points and you want to, for example, if you calculate the metrics of, you know, central tendency, you know, it is mean, median, and mode. Right? If you look at your box plot right now, your box plot will look something like this, where Q1, median, Q3, min, max, and outliers, all of these are individual values; each circle you see is a data point. Right? So if your data is broken down into its quartile ranges, there are again four quartiles which are calculated. Right? So there's one quartile between minimum and Q1; you have another one between Q1 and median; you have another one between median and Q3; and Q3 and max. So we're dividing data into four, you know, four different quarters and working with them because we'll be taking a look at data as a percentile and all of that. Right? So in this case, whatever data we had at hand, the minimum value was maybe one, you know, the Q1 value, the first quartile value was somewhere around seven; the median was 11 or 12; with the third quartile value, the Q3, which is somewhere around 25; and then you had a max value in your data; we said, "Hey, you know, the maximum is 30," and then you have something called as outliers. These outliers are concepts which are deviating from the actual data. For example, if you can see that your minimum is zero, maximum is 30, you might have data which is like 100, 150, 200, which are very different compared to the majority of the population of data. Right? It may be minus one, minus ten, minus 100 also; if you're not expecting that, that is called as an outlier; it may be an anomaly; it may be an error; or it might be those two out of three or two or three of the exponentially rare cases which might occur to give you that kind of result; it is possible. So it is vital that you take a note of where and how all of these elements are represented on a box plot. You have the minimum element on the left-hand side; you have the maximum element on the right side; but here you might have outliers on the left or the right also. But again, when you take a look at the divisions, there is a quartile between min and Q1, between Q1 and median, median and Q3, and Q3 and max. So whatever is the values that fall there, it is a part of that quartile in data, ladies and gentlemen. So it's a very important thing that you understand this. And the next thing that we're taking a look at, as I've mentioned in the in the session takeaway for this particular module, is the advantages of EDA. There are numerous and numerous amounts of advantages; you know, I can just have a one-hour session just to talk about the advantages of EDA, but I want to give you a good amount of highlights to tell you what they are. Right now, advantages are many, but the most important reason why exploratory data analysis is even so popular is because of the fact that you can, you know, look at data, understand that, "Hey, I want to use this data to get to this point; can I do it or not?" If the answer to that is yes, just start digging into the data, hunting for conclusions, looking for insights which your data can give you. A good example to this is maybe you are a person who sells chocolates and cakes and all of these. Right? Usually during Christmas and New Year, this sells a lot. So in that particular case, what usually happens is, you know, you you will have to stock a lot of these items. Right? You will have to, maybe in the month of November, if you sell 100 chocolates, in the month of December you might sell a thousand. So you have to have previous year's data; if you look at maybe 2019 or 2020's New Year data, there might have been 800 people, 850 people who have bought it. Right? So in that case, that's not an outlier; you use that data to know, "Hey, you know, the analysis result is that in the month of December I am selling a lot more chocolates that I am actually producing," or whatever it is. So arriving at a data-driven conclusion; it's not a guess; it's a conclusion, and it is proven by data. Right? So that's an important thing you should know about. The second thing is that whenever you're taking a look at all of these things that you can do with exploratory data analysis, one thing which is usually extremely popular in today's world is that you'll be checking out for data anomalies and the consistency of data itself, to know if your data is correct or not, to know if your data is doing what it is supposed to do, and that it is not hurting any other pipeline down the line; it becomes very, very vital that you understand data analysts, ladies and gentlemen. So that's that. And the third thing is feature engineering. Right? So feature engineering is one of these concepts where you try to understand what feature can be used to obtain what result, or if you can explore into finding a feature in a way where, you know, maybe there is another concept. In the case of diabetes, right, there is another element that we are not looking at which can have an impact on your eventual goal of a person being diabetic or not. Right? To analyze that, there might be some other feature; we are considering blood sugar, body mass index, age, weight, all of these things, amazing, but it is not that that is the entire thing. Right? There might be something else, and that something else to find out is exactly what you're going to be taking a look at with feature engineering. So I hope you have concentrated on the second point because the fourth point talks about something similar, which is data pre-processing and cleaning. Again, over the last couple of modules, you already know this; I have stressed how much is the importance of data pre-processing, data cleaning, and making sure that you feed the right amount of data to your process of how you would analyze data or, of course, your machine learning algorithm. You have to be extremely careful if you're giving the data to a machine learning algorithm to learn because it can go wrong in so many different ways that I have mentioned in the last module. Right? If you show a picture of a cat and a picture of a dog, if your computer is trained wrong, if your machine learning algorithm is trained wrong, it can interchange and see a picture of a cat and say, "Hey, that's a dog, and I'm sure," and it can see a cat and say again that, you know, "I'm a dog," so it doesn't make sense. And I gave you a very simple example, but in an actual productive environment, this can get very, very, very difficult to handle, and people will not know what's going wrong, where it's going wrong, but at the end of the day, it usually comes down to the quality and the quantity of data, too. So it's important. Next advantage here is that with exploratory data analysis, you can have the capability to handle these two concepts called as underfitting and overfitting. Now, underfitting is one of these concepts where your machine learning algorithm does not have enough access to data; you know, it is not searching your data enough where it is getting some meaningful learning done, so it is just brushing through things; it is not understanding things in detail and all of that; that is underfitting. Overfitting is the exact opposite of it; it is when it completely starts looking into your depths of data; it is looking at noise rather than useful information; it is using this noise to learn, which at the end of the day will hurt its accuracy when

It is overfitted, right? So handling underfitting and overfitting becomes extremely extremely important. Because if you do not do so, again your data is messed up when you approach it ahead. Because not only will your, uh, you know, values—the output values that you might get—can be wrong, but it also comes down to a fact where, uh, you know, your machine learning algorithms processing itself is going to take a lot more time. It's going to take a lot more time to learn, and eventually, even after it does that, the accuracy is not great. So it's it's not a win-win situation in any way possible. So it is very vital to handle underfitting and overfitting. Any of these solutions, and EDA is doing a very very good job at helping for that, right?

So guys, in the next section, we're going to be taking a look at, uh, you know, all the concepts of EDA in a practical fashion. Now that you already know all the basics, right, it will add a lot of value if you learn things practically, and I'm going to do my best to do the same. Now what happens in the case of EDA is this, especially in this data set that we're taking—is a really really fun data set to take out—so I hope that you know we do have viewers here who are fans of cricket. Right now, IPL is Indian Premier League, which takes place, uh, you know, on an yearly basis where, uh, you know, it's a very very popular event here in India. It's a cricketing event, and there's a lot of insight which even the teams are very keen to take a look at the data to understand what was the bowler's performance, batsman's performance, uh, you know, who plays really well against which bowler, which batsman. There's a million things that these guys will be analyzing and looking at. They will have entire data analytics teams together who are sitting to analyze individual players. So that is how much, uh, you know, how much complex sport has become these days. They try to predict things; they understand strategies; they try to get strategies from an AI-related algorithm as well. There's so many different things that they're doing. So we are going to use this data set, which has all the data, uh, regarding the Indian Premier League which has happened over the last couple of years.

So there are certain questions which you can ask, right? See, exploratory data analysis is you exploring; for you to explore, you must ask some questions and get answers from your data, right? There are certain data set questions that you can ask, for example, saying, "What is the size of the data set that I'm dealing with?" Well, you can find that out in Python. Next, what is the type of data that we are trying to work with in this particular data set, right? There's numerous types of data which uh one is working with here in IPL data. I think I'm going to show you all of these things practically, but of course, uh, it's not just IPL data that we're talking about here. It is you asking a question saying, "What are the types of data I'm working with? Can I find it out, and can I use it in the next couple of steps?" The answer is yes, and you can find out the types as well, right? The next thing, uh, data set question which is very important to ask is, "Is the data pre-processed?" Right? To find out data cleaning—remember when we spoke about advantages, I told you that it helps in pre-processing as well—this is exactly what we're talking about here. These are all data set related questions in a way where you can ask if the labels are correctly present to the data as well. For example, uh, you know, if you consider a player, if he is just a batsman and not a bowler at all, right? So in that case, if your data is correctly put him as a batsman rather than a bowler or any other thing, right? It might have put your captain as a coach or something like that; there will be something wrong with the data in many cases. So here is where you ask these questions to clear out that air.

The next questions are the set of questions that you're going to be asking is statistical questions, right? In the case of statistical questions, you'll be asking so many things to your data, saying, "Hey, so what are the values of central tendency, uh, you know, how is the quartile division of the chart?" You remember I showed you the box plot which contained Q1, median, Q2, min, max—all of these values. The box plot itself is a representation of the quartile division; your values will be shown; you can do that with the help of EDA. And then when you're taking a look at the third thing, right, the third thing talks about how important is it, uh, for you to understand the total count of all of the individual data elements that are present. Knowing that you'll be using 10 rows worth of data, 10 columns worth of data to, uh, you know, at the end of the day, uh, take, uh, your result and show that, "Hey, my data is my result is backed up with this amount of data," right? That's a very important question, uh, that you are going to be exploring.

So guys, without further ado, let me quickly jump in to Google Colab again. We have used Google Colab previously, uh, however, let me zoom in so that it is better, uh, you know, visible on all of your screens; uh, this should be perfect, I presume. So we're gonna be taking a look at exploratory data analysis on the IPL data set, but before all of that, I have to upload, uh, you know, the data set. So let me quickly go on to upload the data set. As soon as I hit okay, if this orange circle vanishes, it means the data set is uploaded. Perfect, I can close that, and we are here, guys. The first important, uh, thing that you have to take a look at, as as you've already checked out in the case of machine learning and all the other modules, what is the first thing that we do whenever we're working with any program? We import all the libraries and the dependencies, right? So for this particular case, to perform exploratory data analysis here, we require pandas first of all to store the data as data frames; we require numpy to make sure we're working with numerical computations using its data structures as well; and we're going to be using matplotlib and seaborn. These are two beautiful data visualization libraries which I absolutely love working with. We're going to be checking out data visualization itself, uh, you know, ahead in the next modules, but then for now, you have to know that we are going to be using it, right? So as soon as I hit the play button, these libraries are all imported and ready.

The next step after, uh, you know, importing all your libraries is to bring in your data set to the working environment. In this case, the working environment is the IPL data set which we have to bring into our Google Colab. So as soon as I want to click this, we are using read_csv method in pandas to take a look at our data set, right? Now let me zoom out a little so that you can see the data better, right? So I've basically printed the first five rows worth of data using data.head. I can print 10 as well; it's a pretty simple thing. All I have to do is pass 10 as a parameter to it, and instead of 5, I now have 10. Amazing, right? So the data that we have is something like this: there are identifications for the data rows, uh, eventually, then, uh, you know, you're taking a look at all the various, uh, cities, uh, which where the games were actually played, where the matches were actually played, uh, then you're taking a look at the dates of when individual matches were played, uh, who are the two teams who played these matches, who won the toss, what was the decision after they won the toss, what was the result, was the match forfeited, was the match won, match lost, who was the winner, who won by how many runs, and there are so many things. Let me scroll to the right; uh, you know, you can find out by how many runs they won, how many wickets they won, who was, uh, man of the match or player of the match, where was this game played, which venue, who were the first umpire, second umpires, and third umpires as well, right? So this data set is beautiful to work with; it will give you a lot of insight when you're working with it, ladies and gentlemen.

So, uh, to move ahead, uh, we have to next take a look at how big is our data set. I just printed 10 rows for you, right? So with data.shape, we will come to our realization that we're using 18 columns of data, and each of these columns have 756 rows of data. So overall, we have details about 756 IPL matches; we're going to be using that to analyze things, right? So guys, it's very important you know that this is not only, uh, data from one single IPL season, but this data from multiple IPL seasons, right? So we're going to check out all those seasons as well. So when I'm going to click data.info, it will show you just all the, uh, details about—this is a data.shape—just gives you, uh, you know, it in a concise way, but if you want more details about it, as I told you, we have 18 columns worth of data, 756 rows of data, and it will show you samples of what, uh, you know, these data is as well. See, for example, some cases, your umpire number two or umpire number three is NaN. NaN means not a number; it means that data is not present there. So you you do not have the name of umpire number three for a particular match, and for our case, it is okay, but maybe if you are trying to analyze the performance of the umpire, uh, himself or herself, then in that case, it becomes tough, right? So it's very vital that you print out to realize this. Next, let me just print out all the columns that we are trying to use: ID, uh, you know, which season, IPL season, uh, you know, which city was it, uh, uh, whether it was the match hosted, what is the date, two teams which played, who won the toss, what was the decision, what was the result, uh, you know, so basically everything that I ran you through, right? So all the columns and the labels are very very important stuff.

Next thing when we talk about, uh, is finding out if there are any NaN values where there are not supposed to be, right? See, for example, seasons are all these year dates, right? IPL 2016, IPL 2017, 18, all of these things, so these are numbers; there are no NaN values. Then city, city is not a number; you cannot say Bangalore is one, neither about this two and all of that; there you have to give a name: Bangalore is Bangalore, next Hyderabad is Hyderabad, right? Uh, so city name is, uh, will be a character. Then you take a look at a lot of things which are false because all of these are number outputs, but wherever it says true, it is trying to give you a character output, right? Now the name of the team who won a cricket match will not be a number; it will it'll maybe be that, hey, this team won by that many runs or wickets, but which team won? The answer to that question is a character or a string, right? So that's the reason why this is true, and, uh, you know, it is true for all the umpires because all of these are just names of people, right? Even man of the match or player of the match. I hope this aspect of it is clear to you.

The next thing we're going to take a look at is to describe the data in a standard way where we understand using statistical functions. So data.describe will eventually give you a lot of different things: first of all is the count of how many rows worth of data is, uh, there; we have 756 match data; what is the mean of all the IDs, you know, you can find the mean of all the average win by runs in all 756 matches; you can find the standard deviation of how far the data deviates from the mean; you can find the minimum value, minimum number of matches played, minimum number of matches won by, uh, you know, using runs; maybe the team won; there was a draw, then there's no winners, right? Then you're taking a look here at min, 25, 50, 75 percentile, and max, right? So this is the interquartile range; it's called as the quartiles; interquartile range is when you perform a calculation on it, but this is the quartile range that you're printing out; this data falls upon and in between, uh, that quartile. So that's an important thing that you are supposed to know about, right? So that's that's important.

The next thing is we start asking questions on the data set, if you remember, and one important question that we're going to ask is how many matches were played in total with all the data set, which we already know this from because I showed you the data previously; we are 756, uh, you know, matches worth of data out here, right? So that's something which you have to know. The next thing is we'll try to see how many seasons are we trying to analyze here. All we're trying to do is we're finding the seasons columns, and since there are 756 matches, you can get 756 answers here saying IPL 2018, 19, 19, 19, 18, 17. So all of these things, so we're trying to use the unique method here to only print us all the unique data. So this data set consists of data, right, from IPL 2008 all the way till IPL 2019. Again, this may be accurate to what happened in real life, or this can be a hypothetical data set; we are going to assume that it's it happened in real life, so we can just have, uh, you know, more fun with the data set, but that is how it should be as well. This is a hypothetical example of using the data set, guys.

So the next question that we can ask with that particular data is which IPL team won by scoring the maximum amount of runs. So which team won a cricket match where they scored the maximum amount of runs, right? So as soon as I, I lock is basically us, uh, you know, using, uh, a filter; we're trying to apply on the data when we're trying to find the max, uh, you know, max element that is present in that particular column. So as soon as I run this, you know, it will tell us that, hey, uh, you know, it was basically IPL 2017; the match was played in Delhi; here's the date of the match; it was played in between Mumbai Indians and Delhi Daredevils; Delhi Daredevils, uh, won the toss; they wanted to field; the result was normal; it was just win or not win; Mumbai Indians won, and they hit 146 runs which eventually gave them the win, right? So maybe the other team could not chase that amount, and they were bowled out or whatever it is, right? So these guys won by 146 runs; uh, in an advantage you have man of the match, LMP Simmons; uh, you have the venue; it was played in Feroz Shah Kotla, right? Ferocious, I think, is in Delhi; I don't watch much of cricket, but I hope that's the case. Look at this one line of code; uh, it is very simple line of code; all you're trying to do is find the maximum in a column, and you have such beautiful insights to take a look at it, right?

Next, the next question you're gonna ask is, okay, so who won by, uh, you know, just taking down and destroying all the wickets. So which team won maximum that way, right? So this is again IPL 2017; the match took place, uh, in Rajkot between between KKR and Gujarat Lions; KKR won the toss; they wanted to field; uh, and KKR eventually won the match because they bowled out every single, uh, one of Gujarat Lions', uh, you know, batsman or something like that. Again, I do not know if this data is true or not, but as of what I can see with my data set, the question I asked, I have the answer right in front of me, right? So it was played in Saurashtra Cricket Association Stadium, which again is in a Rajkot I believe, right? Again, you have details about the umpires and all of that as well; they took 10 wickets, which is again the maximum number of wickets you can take to win, right? The next question, uh, you know, we can go on to ask is which IPL team won by taking minimum wickets, where they might have not made the batsman out, but, uh, eventually their bowling was so good, uh, that they never let the opposition score, uh, runs, right? So this was a match in between Sunrisers Hyderabad and Royal Challengers Bangalore; SRH and RCB as they are called; RCB won the toss; they wanted to field; and Sunrisers Hyderabad won this match; they won it by 35 runs; this is the minimum amount of, you know, this is the runs that they did, but the minimum amount of wickets that they took was zero. So even though no wickets fell, uh, that bowling only gave RCB chance to hit 35 wickets or something like that, right? So this match was played in Hyderabad; it was played in Rajiv Gandhi International Stadium, Uppal, right? See how nice, how fantastically you're getting answers, uh, to your data here; that is what makes EDA so much fun.

Next, when we are taking a look at, uh, which season consisted of the highest number of matches ever played from 2008 to 2019, here's a beautiful graph that we are trying to, uh, print for you and showcase the same, right? So we have data: IPL 2008, 9, 10, 11, all the way until, uh, 16; we have 17 here on the left, 18 and 19. So as soon as I look at it, I wouldn't have to struggle anything at all; I can see that in the year 2013, the graph is at its highest, which means the maximum number of cricket matches were played in IPL, uh, 2013 or somewhere around 75 matches or 80 matches, around that, right? So it took me one second, uh, to look at that data and give you the output.

Next, when we take up a look at which is the most successful IPL team with, uh, the data that we have at hand here, again all of this is based on the data; I do not know if this data is true to its real-life form, right? So here it shows that, uh, you know, Mumbai Indians have been the team which has won all of these seasons; they have won the maximum number of matches till now. If you take a look at 768 matches, if you ask the question saying, okay, so who won maximum, and can you arrange that in descending order, that is exactly what we've done below: MI, you have CSK with Chennai Super Kings, you have KKR, RCB, uh, you know, Kings XI Punjab, and all of these guys who have won a smaller and smaller amount of matches compared to the one number one person on this list, who is Mumbai Indians, right? Again, it took two lines of code; all we're trying to find is the count of the matches that they've ever won, take a bar plot, and arrange it in descending order, and it is literally it, right? So see how much fun is EDA to work with, guys.

The next thing we're going to take a look at is a very interesting concept; this usually happens when you play cricket or when I see the kids on my street play cricket as well; they think that if they win a toss, there is a huge advantage to win a match; it may or may not be the case based on the pitch, based on the stadium, or whatever it is, or I do not know, right? In that particular case, we have to see what is the probability that a match was won if the toss was won—conditional probability, right? So let's take a look at that. So it so happened in 393 cases that when the toss was won, the match also was won, but in many cases where the toss was not won, or it was either, uh, you know, the basically the match was not won even though the toss was won; that is what false says; true talks about how the match was won because of if the toss was won as well. So if I want to neatly print it instead of just showing you this, if I print it nicely on a graph like this, if I plot it and say, hey, uh, you know, it is true in how many cases; it is true in 393 cases that the toss helped people win, or it is also true that in 363 of these matches that even though a person might not have won the

Task that they might have won the game, or something like that, right? So that's an important thing you have to realize right now. Basically, with this line of code, all I'm trying to do is print more numbers of rows rather than just one or two rows. Uh, you're going to see what I mean by that. Uh, so we have to find out now is the highest number of wins per team per season, right? So as soon as I go on to run this, it will give me an output, right? From the first season, arranging it in a descending order, telling me who are the people who won those seasons, right? So see, in 2008 Rajasthan Royals won 13 matches, Punjab won 10, Chennai won 9, all of these in IPL. 2009 Delhi won the maximum, Branch in 10, Mumbai Indians 11, uh, CSK 112, KKR 113. So you can see, so if I had not done this higher robot, that would just give me two or three rows and say it's done. Now I have access to the entirety of the data, and that is the thing that I've said here. I said, hey, even if I have 10,000 rows of data, you please print it. Even though I have 400 columns of data, you please print it, right? That is exactly what I am doing here, and you can see who won a maximum number of matches in that particular season of IPL, right guys?

Uh, so the next thing that I am going to print is to find out what is the task decision, uh, which happened over all of these years. What is the see tossing a coin in probability, as you checked out, is it has probability and randomness associated with it, right? So in the 700 and whatever times we tossed a coin, the uh, the umpire tossed a coin 463 times, people chose to field, and 293 times captains chose to uh bat first, right? So this shows that maybe if you field first, there's a good chance that you might win; you might analyze the opponent better, or something like that, right? So that's an important concept. The next thing we're going to take a look at is who is the person who has one player of the match or man of the match in a one matches of course. So as of what I know in cricket, once you win a match from the winning team, there is one person who is selected who becomes the mayor of the match. In basketball, it's usually called as MVP or the most valued player for a season or something like that. But here again, uh, you know, take a look at it, highest to lowest is what we are printing, right? Let me quickly scroll up since it's printed the entire thing. So Chris Gayle, uh, Christopher Henry Gayle, you know, I've been to a match where I've seen this person hit the six, and you know, I sort of got the ball; I mean, that's a different story. Uh, so you're basically this person uh has hit uh uh you know, has a record of having 21 man of the matches award. Then you have AB de Villiers who has 20, then you have uh, RJ Sharma who has 17, MS Dhoni with 17, Warner with 17. All of these data again, I do not know if they're correct or not, but as per my data, as per the questions I'm asking, I am receiving beautiful answers, right? So that is what uh, you know, adds meaning to all of these data guys, right?

Uh, and the last thing that we're going to check out is to find out where, in which city, were the maximum number of IPL matches ever played. So as soon as I got to hit this, uh, all I'm trying to do is find out all the uh, you know, find out the count of individual cities, find the maximum of it, arrange it in descending order, and my answer is there, right? So there are 101 matches which took place in Mumbai, Kolkata, Delhi, Bangalore. So these top five places have been extremely popular from IPL 2008 all the way to IPL 2019, guys. So it, ah, this is basically telling you that there is uh, you know, why are people preferring to play in Mumbai? Is it because of the matches? Is it because of how uh, later down the line people started dropping off from the tables? Whatever it is, right? You can analyze 10 different things with one result such as this. This is the beauty of EDA, guys. So this is how we can uh, you know, go on to work with explorative data analysis. It is a very simple IPL data set that we worked with, but when things get complex, uh, you know, you can go on to learn more and more and more, and uh, it is so, so, so much fun that if you ask a question, you can look at your data and pick up an answer by processing something in that data, right? That is what makes EDA so, so much fun to work with. It is uh, you know, I'm trying to portray it in a way where it's fun to work with because you guys are getting started with it, but understand this: it is extremely powerful; it is used by individuals and right guys, let's get started right now, right?

So introduction to data visualization. Ask yourself the first question here: why do you think data visualization is even required? Next, we're going to talk about data science, right? See why it is required as a section on its own, but as I have mentioned, the result of a data science output in beat any domain, machine learning, deep learning, or any kind of a subdomain in the case of data science as well, the result that you obtain will not be understood by a non-technical audience generally, right? You have a complex data set which has 10 lakh entries into it, or you have a result which only is understood by a person who knows Python, who knows machine learning, who knows deep learning, and all of these, right? Now, when you are working in an office, do you think everyone around you is a machine learning expert or a deep learning expert? The answer is no, right? Sometimes you have to take the solution to a client who will not know deep learning, or you have to take it to your stakeholders, your managers, your board members, and explain to them in a way that they understand this complex data, right? So you, as a person, should have the capability to take something really complex, textual data, and convert it into beautiful-looking uh visualizations, guys. And uh, in fact, even uh, you know, here are some really fun things that you should know. The first thing is human beings, right? As human beings, we can comprehend and understand an image three thousand times faster than reading text. How amazing is this? Now, this is one of the primary reasons why the film industry took off rapidly, uh, you know, way back uh, you know, centuries back. But at the same time, you definitely have to know, you definitely have to understand the meaning, the impact it has, right? If you can convey something 3000 times faster, why would you not want to do it? Makes sense, right? Exactly my point here. The second point is that, as I mentioned, your data, your output, or wherever it is that you have to explain, whatever point of execution, it will not be understood by a person who might not know machine learning or deep learning. In that case, it becomes your duty to make sure you convey the information in a way which they understand, right? This is where data visualization comes into uh, you know, the most important sector of data science itself. Uh, it so happens that even in real-life situations where maybe uh you're looking at a sales team, and the sales team is getting some data from an artificial intelligence model or a machine learning model, if you just send the output as it is, there's a good chance sales guys are amazing at sales, but I'm not sure if they're really well-versed with machine learning. So it is my job to make sure that my machine learning model or deep learning model is presenting an output that is understandable and comprehensible by the sales team, right? That is one example where even in real life, everywhere around you, it does happen; it is not only just this example; there are hundreds and hundreds of examples, right? So eventually, we are coming to have a discussion about the definition itself, what data visualization is, and all of that. You have to understand that this is a structured process that we use to take input as our textual data, numerical data, character data, string data, whatever it is; we take all of these data to see if we can create graphs out of it which are not only aesthetical, right? So graphs can be very boring to look at, or they can be very amazing with all colors, points, markers, and all of that, right? It is not only the aesthetics aspect of a graph that we look at in data visualization, but also the functional aspect of it: how good does it look, and how well does it work? These two questions you always have to keep in mind when you're working for a data visualization solution, right? So if you have one graph, it's one solution; it could be that one graph is enough to uh, you know, gather all, you know, compress all of this data and give one concise picture. Sometimes you would need more than one graph. If you require more than one graph, it is usually called as a report where a report consists of all of these graphs, and there are descriptions of each of these graphs, what they mean, what are the takeaways from them, and all of that, right? So it becomes very vital that you know how to take these reports, and eventually, sometimes you have to put it out on an online media, right? Sometimes you have to put it on a website, a web application, and whatnot. So it becomes key that you create a dashboard. Now, a dashboard is basically a website where you can see all of these graphs, and sometimes uh, in the case of business intelligence solutions like we discussed previously, right? If your data is changing, your graphs can be changing in real time as well. That is the beauty of data visualization, and especially in the world of uh neural networks, in the world of deep learning, where you want to track what's happening to your own uh program or code, sometimes with the help of graphs, right? It is very, very vital; it helps a lot immensely to rather not see just volumes of data and uh, you know, rather see a good amount of graphs where you can understand everything without having to spend hours and hours breaking your head about it, right? This is the importance of data visualization, but does it end at that? Is it only two or three reasons why data visualization has been so popular? The answer to that is no, because at the end of the day, it will give you multiple things which you thought which was not possible from a graph. Now, as you look at a graph, you can say, okay, so my data said this, this is what the output is, we are done. No, it helps to map cause-effect relationships, right? Now, uh, I'll give you an example which we've already discussed previously, right? So how do you predict the weather? Uh, you know, previously, if you if you went back 300 years or 200 years to predict the weather, you would have to go outside, look up at the sky; if it's dark, if it's black, if it's full of, you know, clouds which are like really dark in color, it means that there might be rain, snow, hail storm, or whatever it is. If there was a blue sky, very clear sky, it's a hot day; there's absolutely no cloud in the sky, it would not rain as much, right? So a cause and an effect relationship are built here: the cause is that the clouds themselves will cause the rain; the rain becomes the effect. To see if it rains or not on a particular day, you can use this cause and effect relationship to derive another cause and effect relationship between day and rain instead of cloud and rain, and eventually, you can build relationships like this which will give you a lot of insight, a lot of depth, right? So just convoluted layer inside layer, and uh, you know, the amount of insight you get is amazing out of it, right? The next thing we're talking about is feature engineering. See, sometimes when you go on to visualize graphs, you will realize at the end, looking at the graph, saying, okay, so if these three uh, you know, variables are pointing out in a multivariate analysis uh situation uh when they're having a good impact on my output variable, what about a fourth one? What about a fifth one? So this kind of a feature engineering thought, looking at more than, you know, more amounts of features that can add in value to your program, which you thought might not be possible, will definitely become possible if you explore the usage of data visualization in that regard, right? So you might be asking, is that it? The answer is no, because take a look at the third point: it helps in finding and mapping and removing outliers in the data. Now, I've already spoken about pre-processing multiple times in this full course video, but it so happens in many cases that sometimes there are outliers which you cannot see in a huge data set even with pre-processing methodologies. So there you might have to plot all of those on a graph and then work on removing those because there you can visually see it uh, you know, where all of your data, one side, but you definitely have points on the outside, so those are the outliers, right? So you decide if it's an outlier or not and eventually push out those values. If you do not, it will be detrimental to your machine learning model, which will eventually learn things wrong and give you a wrong output, right? I've discussed this already. Next, the fourth point, the fourth point makes sure that your data is understood by a non-technical audience. We just discussed this, and this is a very important point I want to emphasize on all of this uh because I do get asked a lot, saying, is it that important that you have to make everyone understand your data? Everyone is subjective, but a non-technical audience understanding your data is very important because if the client himself was a deep learning expert, why do you think he would come to you for work, right? So that is one thing that you always have to keep in consideration, guys.

Now, the thing that we have to understand, a very, very important aspect here, is that there are multiple uh, you know, data visualization libraries out there in Python, and of course, in other programming languages like R as well, but since we are championing Python here, and of course, the entire world is championing Python for data science, I want you guys to understand the top data visualization libraries that exist, and again, if you already know this, take a second, pause the video, note them down, or of course, put them down in the comment section saying that you know, you know all of these; it will just add a lot of value to your learning, and it just makes learning fun, right? So that we know what you guys know as well. Make sure to do that, and then, of course, when you're resuming the video, here are the top three data visualization libraries which are extremely important in Python, guys. Now, you might ask a question saying, okay, so is this the only three uh Python data visualization libraries out there? The answer to that is an absolute no, because with Python, since it's open source, even you can create a data visualization library on your own if you choose to, but these are what you see, right? Matplotlib, Seaborn, Plotly—all of these are sort of intertwined with each other in a way where uh it gives you the complete functionality which you would require in creating a good data visualization, not just one graph. If you want a hundred graphs, you can do it; if you want a hundred different graphs with a hundred different styles, and uh for it to be equally functional and understandable by a non-technical audience, it can be done with these three uh, you know, visualizations: Matplotlib, Seaborn, and Plotly. And across the world, if you look at trends, there are numerous reasons why each of these libraries are considered; they are the best at what they do. The developers have put in a lot of time and effort to make sure that when you use this library, it is easier for you uh, you know, to do it. So eventually, it adds a lot of value to your learning, guys. So let's quickly check out Matplotlib. Among Matplotlib, Seaborn, and Plotly, Matplotlib is considered to be the world's number one data visualization library if you're using Python. So it works in a variety of situations as the beauty of it. Do you want to use it in a Python script when you are scripting manually? If you want to use it perfectly, it works. Are you using any sort of shell command interfaces where you want to again print certain graphs? It works there as well. Do you want to use Jupyter Notebooks like how we've been doing for a while now? Of course, it does work in a Jupyter Notebook. If it works in a script and a shell, there's a good chance it works in a notebook, right? So that is an important thing. This kind of a diversity uh, you know, helps you learn it in an effective way. The next thing is that when you are thinking about the graphs, the different types, the different varieties of graphs and charts and whatnot that you can put out with Matplotlib, the collection is absolutely huge. The literal number of uh, you know, graphs that you can put out, the types of graphs are amazing, and each of them are very good at the single purpose, right? So you open up your dimension to a wide variety of uh, you know, wide variety of requirements that you can cater to in a very simple fashion. And the third thing is that when you when you whenever you're talking about graphics, whenever you're talking about creating an image or uh, you know, anything else of that sort, the first question a technical person is gonna ask is, okay, so how much of memory and processing power am I gonna require to do this, right? Printing textual information will take certain memory, no doubt about that, but printing an image, printing a graphical entity, will take exponentially more. So in this case, though, depending on the uh depending on how well Matplotlib is designed currently, it is actually very, very efficient in terms of memory; it is very good for processing power, and even if you can get away by using a Matplotlib on a low-power computing machine, which again, it will work absolutely fine; it might take a second or two to execute, but at the end of the day, it will work. That is the advantage; you do not require a really fancy, expensive system uh to run these visualizations, right? And the next thing we're gonna check out is Seaborn. Guess what? Seaborn is a library on its own which is eventually built on top of Matplotlib. So it extends the functionality of Matplotlib if you require uh, you know, a lot more functionality than what Matplotlib can offer natively, right? At the end of the day, Seaborn's primary reason for being another amazing uh graphic library here or a data visualization library here is that it can be used uh with a high-level interface whenever you're plotting any sort of statistical graphs and all of those sort uh so doing this will make sure your job as a person in visualization becomes easy, and at the end of the day, you have uh, you know, a good amount of time working on your logic or whatever it is rather than breaking your head about data visualization, right? Okay. Next, second point. The second point talks about how uh Seaborn works amazingly well with data from NumPy and Pandas, right? Now, we've already discussed NumPy and Pandas to a good amount of detail. Using this data now, once you have NumPy and Pandas data, you cannot say, hey, my library, my visualization library does not support that, so let's convert it into something else. All of that is not there here, as it is. Whatever data is thrown out by NumPy and whatever is the data structures, you know, series and data frames, whatever it is, it, with the Seaborn library, there's a good chance that you know you can directly use it without having to manipulate it a lot. This again saves time and efficiency, right? Of course, when you're taking a look at the third point, the third point talks about how uh, you know, you can plot certain functions over your data; you can work with your data set effectively uh in many cases uh using and learning with all of these data sets becomes easy if your library makes it even more simpler than a generic solution, right? So that is how Seaborn actually works; it simplifies an already simple solution to make sure that you guys uh can learn it all quickly and work with it effectively, right? Okay. Now, the third one that we're going to talk about is

Plotly: Plotly is basically a web-based toolkit that is very, very popular, again, in the case for data visualization. Uh, it has three very important points that you should know at this moment in time.

The first point is that it provides amazing APIs to work with. Uh, just make sure that you can tie in your program here, uh, get the result out of it without having to break your head about, uh, you know, interfacing these two entities, right? The Plotly library and whatever it is that you're trying to visualize.

So the second thing here that you definitely have to think about is how there are supports uh given a for 3D visualization and 3D plots. 3D visualization and 3D plots just take data visualization into a whole another level. But until you get to a good amount of complexity, there's a good chance that you guys might not be using it. So eventually overwhelming yourself with a 3D visualization as of now is definitely not recommended. But what you should be knowing is that with Plotly you definitely can do it. Even with Matplotlib and Seaborn to a certain extent, uh, you know, you can work with a couple of plots to make sure it looks like three dimensions or it has three dimensions itself. And, uh, you know, in a concept uh called as contoured plots, contour plots are these entities which are notoriously difficult to map, but the results they give is very beautiful to look at and highly functional. Those can be used here as well. I just gave you one example; there are many things uh that Plotly can do better than Seaborn. There are many things that Seaborn can do better than Plotly and Matplotlib, and of course Matplotlib in its own nature can do multiple things better than all the other two as well. So that's an important thing, uh, you know, you guys should know as well.

Now we are coming on to a section where we will be discussing the important visualizations that exist in the field of data science. Right? Uh, there are hundreds of uh visualizations you eventually that you can put out if you're looking to the extreme depths of it, but out of this hundred, five or six or maximum ten of them are extremely popular in a way that you definitely have to know all these guys. So I recommend you have a pen and paper ready so you can take down certain important points of all of these. And after we go on to discuss this, right, we definitely are going to check them out practically so that you know you uh know what we are dealing with as well. So theoretical learning plus practical learning is what uh we here at Great Learning stand for, and of course, uh, you know, we're gonna take it that way to make sure you guys can learn it in a simple fashion.

Right now to talk about the important visualizations for data science, there are so many as I mentioned, but I want to discuss six important visualizations with you. So we have bar plots; we have the pie charts; we have histograms; box plots; scatter plots; and of course the strip plot as well. So these six are the ones that we're gonna be checking out theoretically and practically, guys.

Now, uh, wait, let me actually go back before we discuss bar plots. Tell me if you have actually worked with any other visualization. It doesn't have to be in Python. There's a good chance that you might have used a pie chart in your school or college or sometime then, right? Histogram is very familiar for photographers because, again, while editing it that becomes a very important component. Box plots is used for statistics; we check this out in the previous module; I want to revisit it again because it's important. Scatter plot and strip plots again give you a beautiful-looking visualization, yet so much data with it as well. So guys, if you know any other visualization which are uh which you think are important, head to the comment section and do put it down there so we know again what you are looking towards as well, right guys.

Now taking a look at bar plots, right? See bar plots, as you can see on your screen on the right side right now, is what a bar plot looks like. It is a very beautifully aesthetically pleasing entity; you know, you can select all the colors; you can work with multiple things; and it is mostly used in a case where we have to, uh, you know, show that there are certain numerical changes whenever we compare it to categorical labels. Right? See a categorical label can be something such as height of a person when you're taking a look at age. Right? You can put a student's name, for example, and map their particular age as well, or consider the age of common students, like maybe just take the case of class 10 students, try to map the age of all of these. Right? So there is a categorical element being that there are people from class 10, and the numerical value here is the age that eventually that you're going to map. Ah, the right; it can be height; it can be age; it can be weight; it can be roll number; whatever it is that you want to do, bar plots will definitely help you get there. Right? Now even looking at a certain data like this, you can just say, uh, hey, here's the height, weight, and age of three different people. So, uh, using that you definitely can show, you know, the grape gray color is one person, yellow color is the other person, blue blue is another person. Of course, you can have 100 different persons uh if you choose to, but yes, that kind of a feature will make sure that you know, instead of just looking at an Excel sheet with like a thousand students with all the all the data of their age, height, and weight and all of that, if you were to look at a graph like this, you can definitely pick it up rapidly and say, okay, so this is the average height; this is the who is the tallest person; who's the shortest person; who weighs the most; who weighs the least. If that is the kind of information that you're looking at, you definitely can get access to that with bar plots.

Right? Next thing we're going to check out is pie charts. Pie charts are something I definitely remember learning in my school days way back, uh, you know, a couple of decades ago for sure, is that this is a kind of a thing that will show you there are certain divisions, right? An entire pie chart is consisting of 100 of the entity that we are in discussion, but when you take a look at a fraction of things that you have to discuss, pie charts are an amazing way to do it. Here again, we'll be comparing numerical values and categorical values and all of that. And see, for example, your screen right now; you can see three numbers: 40, 40, 20. When you add all of that, you get 200, right? So that is showing the entire completion uh in one circle, but in those circle, each color shows a representation which has a categorical label given to it. Right? Now, for example, uh, think about a situation where 40 percent of the people like pizza, another 40 of the people like, uh, you know, eating fries uh in any maybe as an evening snack or something, and then there were 20 of people uh who like to maybe, uh, you know, get a fizzy drink or get a soda or something like that. Right? When you think about evening snacks, now these are just random things that I just told you; it will apply to any different, uh, you know, situation. If you want to think about pizza—now I'm hungry, thinking about pizza—right?—you can just say that there are three people, and this is how much pizza each person ate: two people ate 40, 40 of pizza, and there was another person who ate two slices out of 10 or whatever it is. Right? So you can find any sort of situations where you have to divide in fractions, and you can represent that easily in the case of a pie chart. That is the beauty of using a pie chart, guys.

So after a bar plot and pie chart, the next thing that we're going to check out is histograms. Histograms are extremely important in the case of data science because here whenever we're trying to assess what a continuous group of variables are doing or values are doing, and we want to see their performance, their changes, or why assess why they're changing the way they are, what is the result of these continuous variables changing, all of these, right? You can find out easily with the case of an histogram. When you're using a histogram, uh, you know, usually data is in the range 0 to 255, 1 to 300, minus 500 to plus 500. So this kind of a range makes sure that your histogram doesn't go out of bounds and that it works effectively, and the visualization that you get at the end of the day will add meaning rather than confuse you. Right? That's an important thing. So talking about an example for an histogram, maybe think about how you can categorize people based on their age. Again, you can you can do it with a bar plot again, but at the end of the day, it's vital that you understand how you can do it with an histogram too, and knowing that you can do it is the first important point, guys. So you can, uh, you know, think about any other example where you would want to map a continuous value, uh, you know, for example, if you want to map the changes of uh cost of fossil fuels, maybe the change in the price of petroleum, diesel, or whatever that is. Again, that's a popular usage of histogram that I have actually seen in petrol banks near my place; they put up a chart to show that, you know, either the price went up, price went down, whatever that is, and all these, right? So this is histogram.

The next thing we're going to check out is box plot. Box plot is this very simple box with two lines that packs so much of an impact; it gives you a ton of information, uh, which is again, when you look at it as it is, you'll be like, okay, so what is it? But once you understand what a box plot means, like I've explained in the previous module, you definitely will know that it is just an expansive, uh, you know, amount of wealth of information that you're going to get out of it. Right? Uh, now as you can see, every single uh point here will uh talk about a data entity. For example, min talks about the minimum value you have in your data set; max talks about the maximum value; and of course, you might have outliers on both sides of the data as well. It can go below the minimum; that can go above the maximum; and whatnot. And then you have three important points: Q1, median, and Q3. Right? So minimum to Q1 is the first quartile; Q1 to median is the second quartile; median to Q3 is the third quartile; and of course, Q3 to the maximum value becomes the fourth quartile; and wherever you see this center, and of course all of these are not labeled all the time, there is numbers uh given in that particular case, and this the minimum maybe one, the Q1 may be five, median is somewhere around uh 11 or 12, uh Q3 can be somewhere around 25, maximum is 30, all these outliers can be 33, 34, 35, which you think make no sense with your data. Right? So it's a simple thing, but once you understand it, you will get to know everything: minimum, maximum, medium, the uh the quartile range and all of that. Once you have the min, max, Q1, median, and Q3, you can find out the interquartile range as well, and you can start, you know, reverse engineering statistics out of a chart as well. That's the beauty of box plot. So guys, it is vital that you know this.

Next we're going to talk about scatter plot. Now scatter plots are very, very vital for machine learning; they're beautiful to look at. Just take a look at your screen right now, right? Uh, it is amazing that you can have and work with visualizations like this. It is usually very popular whenever you're taking a look at two numerical data points that you want to compare, uh, you know, side by side, or usually this becomes a result; a scatter plot is the result of what happens, uh, you know, when you plot a regressor line. Right? Again, we've discussed linear regression in the case of the machine learning module. So once you plot what happens in the regression process, the eventual result is usually plotted very well and understood very well if it's done using a scatter plot, guys. That's the important aspect of things here. Uh, you might understand saying, okay, so can you give an example where, you know, what you see on the screen works in that case? Think about the number of clothes that you might wear based on the temperature around you. Right? If it's 40 degrees, 50 degrees, you would be shirtless; you'd be wearing very thin clothes, uh, clothes that absorb moisture, your sweat and whatnot; and maybe if you're sitting on a mountain or if you're hiking in the Himalayan ranges, uh, temperatures can hit minus 15, minus 20 and all of that. Right? So there you will be packed and stuffed with thermal wear, fleece jackets, outer jackets, and then, you know, just just you'll be stuffed to make sure your body can maintain temperature. Right? If you want to plot both of those: in the summer, these are the number of clothes I wear; in the winter, these are the number of clothes I wear. You know, at the end of the day, it's a beautiful uh usage of a scatter plot as well, guys. So for the first important thing, if you're you're supposed to remember is that it's important for machine learning; the second thing is that it's aesthetic, beautiful, and uh really functional at the same time. Right?

The next thing we're going to take a look at is the strip plot. So strip plots are very, very similar to the usage of scatter plots. So what happens in the case of scatter plot is that we're going to be using two numerical coordinate labels, right? Temperature in winter, temperature and summer and all of that. In the case of a strip plot, one of the values will be categorical in nature while the other one is numerical on in nature, guys. So when you're thinking about an example, think about maybe the amount of sales that happen on a particular day. Right? Sometimes on a day your sales might be amazing, or the next day the sales might have dropped a little, and each entity, each dot you see there may be an individual sale, uh sale for a week, sell for a month or whatever it is. Right? So you can see total bill is what is given on the y-axis; day is given on the x-axis. So on this day you just sold 10, 20, 30, 40, uh and you sold a product maybe which are worth these amounts; total bill was somewhere in the 40; total bill was 50. So each of these dot represent a single bill with the worth uh probably shown on the y-axis as well. Again, this is a very simple example; you can find uh numerous examples in the case of strip plot; it is used very popularly, and if you're a person who is looking towards data science, you definitely require knowledge on strip plot, guys.

So I hope all the plots that we discussed theoretically until now uh you guys are clear with them, uh uh because the next concept that we're going to look at is a little different from the graphs that we have taken look at. Right? So now all the graphs, uh, let me just go back: strip plots, scatter plots, box plots, you know, histograms, everything we just looked at is a graph; the next one we're going to look at is sort of a graph as well; it's called correlation. Now correlation is one of these beautiful uh things that will eventually uh it's it's like a matrix, but it's a colorful visualization of a matrix that will give you the the interdependency between two variables. Right? Let me show you one important use cases here; you can see pregnancies uh written, and again on the x-axis you have seen pregnancies written. So what is the relationship between pregnancies and pregnancies? It is literally equal, right? So it has to be the brightest color possible; it's the same case with all the other elements; hence the reason why you see these diagonal elements to be maximum uh correlation or white in color because you you cannot differentiate between two exactly same values and say they are different; in this case it is not possible. Right? And then if you're taking a look at the dependency of uh you know skin thickness and the pregnancies, well it does not relate a lot as uh is what we can obtain from our data because at the end of the day you can see that there is black here; black means that there is 0.0 correlation; that this data does not relate to the other data and have an effect on the output as well. Right? So this is a very simple thing; so we've already done multivariate analysis in the case by using diabetes prediction; just think about the same thing here, and that kind of a correlation, you know, it adds a lot of value; explained it there as well, and I'm repeating it enough to make sure that you guys are clear with all of these concept, guys.

So I hope uh you understood all the plots that we've used; I made sure to take uh you know examples to uh help you guys understand all of them; there was a simple uh demo visualization there as well; and all of these is eventually theoretical learning to get you prepped for the practical session that we're gonna check out. Right? So guys, let me quickly head into Google Colab and uh you know as always the Google Colab is the uh Python Jupyter Notebook that we're gonna be using throughout this course, so I'm gonna go there; we're gonna check out all these plots individually and practically and see what are all the nitty gritty things, the fine print that we have to understand to eventually use them effectively, guys. So let's quickly go to Google Colab and check them out.

Right guys, so we are in a Google Colab; we're gonna be checking out all the different plots that we just uh you know spoke about, but the one important thing that all of you guys have to know about is the order that we use to work with any sort of uh code in Python. Right? So the first step is the most important step; we are importing Seaborn, and we are importing Matplotlib. Right? So as soon as I go on to run it, uh you know Seaborn is included; my Matplotlib is imported; then we're going to take a look whenever we're plotting line plots. Right? So we're going to take a look at the fmri dataset; fmri set as a part of Seaborn uh it basically talks about the data whenever you go on to have an MRI scan; it's a very complex thing which cannot be understood by a lot of people; you have the subject uh the id given to the subject; at what time point was there; what event uh you know in which part of the brain, the uh parietal region it was taken; what is the signal that was firing from a neuron uh or you know possibly in that particular region of your brain. So it's a very complex data set, but we can simplify it by taking a look at data visualization. Right? Now these are the first five elements of what our dataset looks like; we have uh you know labels which are subject; subject is of course the person; time point of when it was measured; event uh is talking about either a stem or again we have another one which you're going to check out; then you have region, what region of the brain; and what is the signal that was obtained from that part of the brain when scanned as well. Right? So let's quickly uh you know plot a line plot by making use of time point and signal column data. Now time point is one column data; signal is another column data; so on the x axis what I want is I want time point; on the y axis what I want is this signal, right? The signal value versus uh the time point. Now as soon as I go on to use sns.lineplot where x is equal to time point, y is equal to a signal, and data is coming from the fmri data set, uh I'm just telling her, hey, on the x axis put all the values of time plot; on the y axis put all the values of signal; and eventually show me what the plot looks like. Now as soon as I run this, you can see that this is a beautiful visualization, right? Uh, on the left your signal on the y-axis, zero is there; you have minus five on the y-axis; you have the individual time points and all of these. So this adds uh instead of looking at this

And saying like, okay, so I cannot understand anything. If you take a look at this, there's a good chance you can figure out that, hey, the minimum value of the signal was maybe point minus zero seven or something. And then, uh, your maximum value touched somewhere around 0.15. This was the time point at which this particular part of the brain was doing this thing. So eventually, experts, medical experts, can understand this definitely better than me as a Python expert, but that is the that is the case, right? Helping people understand things better. Now, uh, it is in blue color; there's a nice uh, this thing running out on it, so it looks great even as it is, if viable to understand this. But is there any way we can add more value into this? Well, of course, you can add a color to it as well. You can use a different color to go on to map it. Now, sns.lineplot(x = time_point, y = signal, data = fmri). This is exactly the same that we actually use, but here we are adding another thing called as hue = event. So we're adding a hue from another aspect of data to understand what it looks like, right? So you have hue, and you have a stem here; you have the data coming in from the time points again. On the y axis, you have signal, so you are seeing a differentiation of our data where you have a hue called event, and this event data is also being published here, right? So the event says stem, uh, so you can see the changes that happen between that time point and that aspect of signal as soon as you take a look at the data set. And all we added was hue = event, right? Everything else is literally the same thing that we checked out previously.

Now you know, after adding the hue, after understanding uh, time point, signal, all of that, uh, so maybe you want to add some styling; maybe you want to change what the uh lines look like and all of that. You definitely can do it. As soon as I want to hit play, uh, you know, you will eventually see that, for example, the hue uh visualization that we're seeing, instead of just being a solid line, it's now dashed or uh dotted lines, as you can see, right? So if you want this kind of a differentiation to help you understand it better, or if you want to look at it in a different way, or if that is a requirement where a legend value is different, this is something that you are going to use, guys.

Now, building upon this, there are even more parameters that you can try uh to go on to use in this case. See, uh, we are not being shown how many data points have gone on to make this graph, right? It just looks very smooth, and sometimes maybe one or two points; you can get one, two, three, four, five, six. If you look at where the line changes a little, there's a good chance as a point there, but why struggle so much? Let us have our markers show us how the data moves, right? Now, if you can take a look at this, look at that; these are the points from one point to the other is how the eventual graph comes into being, right? Now with hue again, there's a different style; there's x and dashes; uh, in the case of stim, we have uh, you know, just uh circles and a solid line. So this is something that you definitely should know about when we are uh using lineplot, guys, to go on to use line plot, sns.lineplot, where sns is basically what we imported from seaborn, right? It's a pretty simple way to go on to understand line plots.

Now the next thing is, let us check out bar plots, but then in bar plots, as you can see, I've been importing a data set here, a very nice Pokemon data set, so let me go on to actually import it. And okay, perfect, so I already have it imported. All the data sets we're going to be using further on uh to showcase the various visualization. If you're wondering about why I didn't import fmri, fmri is a dataset which is already part of seaborn, so I didn't have to explicitly import it. But in this case, uh, you know, we are working on the Pokemon dataset. Pokemon uh, I think was an amazing TV show when we were all kids. I definitely remember uh coming back from school and uh, you know, running from school, in fact, sometimes to make sure I get home in time to actually watch it and all of that; it was a beautiful part of our childhood. So let's take a look and use that data set to see if we can produce beautiful visualizations by making use of what bar plots, right? So the first step is, as always, you have to import whatever is the libraries that you're going to use. In this case, matplotlib and seaborn were already imported here; I'm going to be using pandas. Why do I need pandas? Well, pandas is basically used to read the data set, right? So in this case, pandas has a function read_csv, which will help in bringing together and importing the data set that is done. The next step we're going to be taking a look at is to print the first five rows of what our data looks like, right? Now, perfect. Look at this: abilities of the Pokemon, uh, you know, there's an ability called as blaze, there's an ability called solar power, overgrow, chlorophyll. Or, fortunately, of course, if you know what Pokemon we're talking about, head to the comment section and let me know. Uh, there are multiple types of attacks: dragon attack, uh, electric attacks, and fairy Pokemon, uh, you have a dark Pokemon, your bug Pokemon and all that, right? So against fire Pokemon, flying Pokemon, water Pokemon, what is the value of strength uh that each of these different—so every single row worth of data you see is a Pokemon—again, against poison Pokemon, psychic Pokemon, steel, water attack, all of that, right? Uh, so what is the happiness? How happy is your Pokemon? Now, if we watch the show, we do remember one or two Pokemon which is always unhappy; they're annoyed no matter what, uh, their uh, just their, you know, their owners or their trainers do, right? So that's an important thing you have to have to know. So this is a lot of things. What is the classification? Is it a lizard Pokemon, flame Pokemon, uh, seed? How many has it been captured? How good is its defense? How much of experience does it need to eventually uh, you know, evolve or go into another form of evolution, right? Evolution was something that we checked out—your Bulbasaur, Ivysaur, uh, Venusaur; you have Charmander, Charmeleon, all these different Pokemons. And again, these are just—a little—I've just printed out the first five, but we have an entirely huge uh data set that we're talking about. Uh, what is the index in the Pokedex? Pokedex is this identification device that they use to identify a Pokemon, uh, you know, and all of that, right? Even uh, what type of attacks they have, what is their weight, what is their uh generation that they're in: Charmander, Charmeleon, Bulbasaur, all of those things, right? Uh, then you have if it's a legendary Pokemon or not; if it's zero, it means the Pokemon is not legendary; if it's one, it means that yes, it is a legendary Pokemon, right? Now, uh, to understand and ask—this is an addition to the data analytics module; you can consider it for sure because here, if there's a question that says, okay, so what is the speed of the legendary Pokemon that uh that you have to find out when you compare it to non-legendary Pokemon? Again, as soon as that bar plot, where x axis we're talking about if it's legendary or not, y axis is the average speed, data is picking up from our Pokemon data set, and when you do it, plot.show(). So here you can see two bar plots, right? First of all, x axis is talking about if this Pokemon is legendary or non-legendary, and the speed. Now, as soon as I look at this, I can find, okay, if it's a non-legendary Pokemon, its speed is somewhere around 60 kilometers per hour maybe, but if it's a legendary Pokemon, it is way, way higher; the average of all of these uh, you know, legendary Pokemons is way higher than uh what is available for the non-legendary Pokemon, right? It took me two seconds to find that out rather than look through a thousand-row data set. The next question we can ask is, okay, so instead of speed, can we check the weight? Do legendary Pokemons weigh heavier than non-legendary ones? The answer to that is an absolute yes, because as you can see, uh, non-legendary Pokemon may comparatively very less, and legendary Pokemon are huge uh in their weight and size as well, right? So this is an amazing thing; again, it took us two seconds; it's literally one line of code uh that is showing us and giving us so many insights, right? That is the beauty of data visualization.

The next thing we can check out is, let's have the same thing: if x is legendary, y is equal to speed that we previously discussed; data is coming from Pokemon here, like how we did with the line plots; we're gonna add a hue here; hue is based upon the generation when you are trying to compare, so you're trying to give a color to the various generation of Pokemon, so here that you can understand so many things. First of all, on the left, you have non-legendary Pokemons; in that, you can see if it's blue, if it's a first generation of Pokemon; if it's yellow, it's a second generation; green is the third generation. So the evolution of Pokemon and their individual speed for their evolution is uh plotted as well, and the same goes for every uh and the same goes for these legendary Pokemon as well. So there's level one legendary, level two legendary, level three legendary, right? So even when they evolve, you can find out the average uh, you know, speed at which they travel based upon the revolution. For some reason, evolutionary Pokemon, as they evolve more and more and more, their speed comes down a little, but yeah, see, again, you answered one more question even though you did not have to, or it just gave you an insight, a hidden insight, right? So that's the beauty of doing that. Now, uh, you know, we've been checking out all of these uh different colors, all of these different things. Can we use custom colors? Will be a different uh question that you're gonna have; the answer is an absolute yes. Now, if I want to plot the same speed and legend is legendary graph, you know, is legendary versus speed this graph, but if I add a color element there and say, hey, a palette—it's called a palette here in the case of seaborn—if I call it as vlag, vlag or black as it's called, is basically changing the colors into a more pink and a light blue or, you know, subtle subtle colors that you're seeing on your screen right now. So again, this is aesthetic purposes, or maybe you're a person who wants to see it everything in red; if that is the case, put color = red, and at the end of the day, instead of having a palette with the different colors, it's just going to put everything on red and show it to you if this is what you require as well. This again shows what all is the things that you can do in the case of bar plots, guys. So I hope you understood bar plots; it's a very simple thing, and since with the case of Pokemon data set, I think I hope it became more and more uh, you know, understandable by you guys and more and more fun, right? The next data set that we're going to check out and, of course, the next type of plot that we're going to check out is this scatter plot.

In the case of scatter plots, a very amazing data set to use to train all you guys is the iris data set. The iris data set talks about the varieties of petals and the sepal length of a petal, different three different types of flowers, and gives you a lot of data about all that, right? So let me actually print uh the first five rows of it and show you. Uh, see if it's a flower uh, you know, you have a setosa, then you know—in fact, let's print the first maybe 20 uh rows to see what it gives us, right? As soon as you do this, you can see uh, oh, okay, it's all set osas, for example, right now, but there are three types of species of flowers that we're going to be discussing here. So three types of species that we're going to be discussing are setosa, virginica, and versicolor as well. Each of these types of flowers has its own width in terms of petal, the length, the width of the sepal, the length of the sepal, and how it differentiates from the petal itself, right? So a lot of details can be obtained from all of that. So guys, as soon as I want to hit click, we're trying to map the sepal length and the petal length based on the data, right? As you can see, again, all the parameters are literally the same; x talks about mapping values on the x axis, y talks about mapping values on the y axis; data is from the iris data sets. As soon as you use a plot to want to do this, it is just telling you that the sepal length is this for this particular petal length. Uh, this can be actually divided into three different groups based on the three different types of flowers that the data set has, but having the capability to point out each and every single flower and say, hey, this is what its sepal length is and petal length is, is something amazing, right? Now, instead of me running through all of these manually, trying to map, break my head with the data, I look at this, and I I know easily what the sepal looks like and the petal length is like, right? Perfect. But now, as you can see, everything is just the same color; it gets confusing to know what these three types of flowers are, right? So let's actually map it based on species and give it a different color. As I told you, there are three different species: you have setosa, your versicolor, and you have virginica. Now, as you can see, uh, all the blue colors has a sepal length which is not very big; it is from somewhere around six centimeters or whatever it is; petal length is again very, very small, so the petals are tiny as uh uh, you know, very, very tiny. And then your versicolor, where, okay, the petal length is also really good; the sepal length is amazing. And then you have the biggest of the flowers in our discussion, which is virginica. Virginica are huge in terms of its sepal length and huge in terms of its petal length as well, right? So all we did is we added hue on—what did we add the hue on? What column did we add the hue on? Basically species, right? As soon as you add the hue, it is going to give you a legend, and of course, it's going to color all the individual data points so that you understand it better. Amazing.

The next thing we're going to add is, let's say if you want to differentiate between each of these flowers; if you want to add a hue based on the petal length, right? Now, if you add a hue based on the petal length, it will help you differentiate and understand it better. All these light-colored ones have a very small petal length as as and when it grows and grows and grows; when the color becomes darker, it means at the end of the day that uh, you know, your petal length is changing. And as soon as I look at it, if I look at all the light-color things, I can—okay, so the petal length here is very low, and as soon as I look at the dark ones, I'm like, okay, the petal length is amazing. So if that is an effect that you want to achieve, it is literally one word; you have to add hue, and you are done. Previously, we added hue based on species, so we got three different species, so three different colors. If you add the hue on a petal length, it is not three different petal lengths we are discussing; it is a continuous variable, right? So here it just changes it based on the uh saturation in the color to make us understand it better, guys. So this was uh uh, you know, scatter plot; I hope you guys understand it. Until now, we checked out line plot, bar plot, and scatter plot, right? The next thing we're going to check out are histograms, or are they called as dist plots. Uh, so in this case, we're going to be considering a data set; it's called as the diamonds data set. Now, if you're wondering about what diamonds data set is, let me print the data set for you to show you what it means. Now, whenever you take a look at a diamond, right, diamond is usually cut and polished; it has a carat value given to it; it has a clarity index; it has a color given to it; it has the size, shape, the price, the depth of color that you can see, or how well things are magnified on the inside; how good is the cut, because eventually the diamond's biggest thing is its cut; if it's cut amazingly well, it will reflect light uh way better and all of that, right? So as you can see, each diamond, each row of value you see here is an individual diamond; it has a carat rating; how is the cut? Is it just ideal? Okay, okay; is it premium? It is good; what is the color that it's giving? What is the clarity it has? Depth, uh, the price, the x, y, z axis values to showcase its size and all of that, right? So this is the diamonds data that we have. Now let us just take one column here; let us take the price column here and plot a histogram based on it. Again, the syntax is very similar: sns.distplot(diamonds['price']). If I print diamonds['price'], the plot that you're gonna see here is something like this, right? So you have a function here that is showing your value as well, and you're actually getting the density too. So this is basically giving you density; uh, as you can see on the left side, you have density where how dense is that material compared to the price, and you're getting a functional value to go towards that as well. But let's say we do not want the frequency, or let's say we only want to print the frequency rather than uh showcasing our density, right? So that particular case, if I just put hist = False, so this is a histogram and a bar plot that you can see on the previous uh screen, right? If I just want the frequency curve, guess what? I I just have to say hist = False. If hist = True, it is printing both the histogram and the frequency curve for us. In this particular case, we just want the frequency curve and not the histogram, so we put hist = False, and the effect is achieved, right? It looks beautiful; the price density, you just have one nice curve that is showcasing you what the frequency curve looks like. Now it looks very nice; why not add some color into it, right? So we add another uh variable here; we are on the parameter called as color, and maybe let's color it in green, right? Now, instead of having a blue-colored line, ah, you know, again, uh, if we just put here again, instead of color = green, along with that, if you put hist = False, you'll just get the frequency curve, but just to show you that it creates two different curves for the frequency chart and the bar plot and the histogram itself, uh, I just showed you that maybe you don't want uh, you know, green; you want red; perfect; you can have red as well. So there's a lot of support for colors; you can make it really, really beautiful to look at if that is what you want to as well, guys. So the next thing, uh, you know, we people entered both histogram and frequency curve; we just printed the frequency curve; what if we only want the histogram and not the frequency curve, right? So in that case, you just have to put kde = False. If you put the parameter kde = False here, you will only have the histogram printed, but no uh, you know, uh frequency curve that gets printed for you. So uh, that is one important thing that you have to understand. And the next thing—right, in fact, even before you go to the next thing, let me show you what bins—

Mean now, as you can see, this is one data of two, three, four, five, six, seven. So every single uh trough that you see on the top, it is actually a line that goes right; it goes like this, uh, into the x axis. Now, in that particular case, these are individual elements that are called as bins. Now, uh, maybe this is just printing us hundred different bins of data, which is eventually getting confusing. So I maybe only want five bins of data. Right? So if I just put bins equal to five, uh, as you can see here, uh, see how well this looks now. So instead of it breaking the data into so many things—zero to five hundred, five hundred to thousand, thousand to two thousand, all that—here it is just taking, hey, zero to three thousand is one bin; let us just group all the data and print it. Right? So you have one, two, three, four, five—just five bins. Right? If you want ten bins, of course you can do that as well. If I just change bins equal to ten and if I hit play here, well, now you have ten bins: one, three, four, five, six, seven, eight, nine, ten. Perfect. Right? So this will give you a better clarity. If you want to zoom in and zoom out to understand your data from a bird's eye view, is something that's very important that you should know about. Right?

The next thing we're going to check out is how you can plot it on a vertical axis. Now we have always been plotting a histogram on the horizontal axis. Right? We've been starting from the bottom and going top, top, top. What if we want to start it from the side and go on to do it sideways? Right? If you have to do it from the sideways, it is possible. Uh, there is another uh parameter that you are supposed to use. Uh, all the code that we have checked out until now is very, very similar; you just have to add another parameter called as vertical. Put vertical equal to true, and it will give you a bar; it will give you a histogram which is again mapped on a different scale, as you can see. Now it's instead of horizontal, everything is uh vertical. Right? So vertical equal to true means that, hey, give me my histogram and put it out in the vertical way where I would want to see it like this for some reason. Right? If you want to do it, you have the option to do it as well, guys. So this is a disc plot; this is a histogram. You saw how we can go on to print just an histogram, just a frequency curve, both of them at the same time, and a lot more things as well. Right?

The next thing we're going to check out is a box plot. Now box plots are really, really fun to work with. For this, we're going to be taking a look at the customer churn data set. So what is the customer churn data set? Let's print it out and see for ourselves. Right now, customer churn dataset talks about if you are a person who is willing to buy uh or use a mobile service or a phone service based upon a lot of factors uh that uh talk about the company itself. Right? For example, you have a customer; this customer is a female; he or she is not a senior citizen; they are a partner; are they dependent? No. They want one phone connection; they do not have services right now; they do not have phones right now. You're giving them internet; you're not giving them online security; are you giving them online backup, device protection, technical support? Are you giving TV over the phone line? Are you giving movies over the phone line? What is the contract like? Is it one month that you have to pay every single month, to month, or you just pay once for one year? And this person has for an internet and all of that. Uh, do you have to print a bill all the time, or can you, can you go paperless to save the environment? What is the payment method? Do you have to send an uh mail with the check in it, or do, can you, you know, just pay it online and be done with it? All of that. What are the monthly charges, total charges, and if the person eventually bought uh the dataset or not. Right? So all of these uh you can try to predict if a person bought or not by again performing multivator multivariate analysis here. So let us actually go on to use box plot to plot uh churn and tenure to see if a person bought something based on the tenure. Right? Now what I mean by this, as you can see, uh, is that uh the maximum number of people who bought uh the product, yes. Right? So they wanted a smaller tenure. Now if you go to three people and ask them, hey, you have to pay one amount which will cover the entire year, they might not have the budget or the finances to go on to do so. That is literally what is being showed here as well, as you can see. The tenor, if it is, if it is a smaller tenor, then uh, you know, people are likely to pick it up. If it's a larger tenure, it means that people might not be comfortable because they might not trust you; you might be new into the area; you might be a person who's giving the phone or internet services which they, no one has used before that they don't know the reviews and all of that. So these people decline just because the 10-year hay is long as well. So the person who's looking at it on the back end can understand, okay, so our tenure is long; this is the reason people don't want it; maybe let us reduce the tenure and have them on board. Right? I answered a very important question there.

The next thing is maybe let us check out based on the internet services and the monthly charges that they are quoting. Right? So if it's a regular DSL internet that they're using, the internet charges are, monthly charges are maybe the average, the median charges are somewhere around 60 dollars; minimum is 25; maximum is somewhere around 98; Q1 is somewhere around 45; Q3 somewhere around 65. Right? You remember the box plot that I showed you; this is exactly what we're discussing now to analyze it. Maybe for the people who used fiber optic connection, they are paying a lot more in monthly charges because they may be getting a better speed; they're getting more download and upload FUP, the fair usage policy and whatnot and all of that. Right? But for the people who do not want to opt in for an internet, well, they are literally paying nothing, maybe just a little bit of charges just for the phone line, but nothing for the internet. Right? So you can see from this uh saying that, okay, there are people who definitely are using both fiber and DSL, and people are not using as well. The people using the fiber optic are paying a lot more; that's an amazing thing which we just found out by taking a look at that. Right?

The next thing that we're going to take a look at is how you can uh plot based on contract. So all we are trying to change is the x axis, y axis values here, guys. Now there are people uh who prefer uh month to month uh uh you know contract here, then there's another group of people who love a longer tenure. Right? A one-year contract is going to be a longer tenure, a lot of days, and then you have another tenure which says, okay, you have to pay for two years and you'll be happy. People are there; that is the point to show this particular graph. Month to month, there is a good amount of people; the average 10-year is somewhere around 10. Uh, you know, for the average favor of one-year contract somewhere around 45, the average 10-year for two years of worth of contractor somewhere around 60. Right? So people are wanting and are willing to use either one of these to understand what is popular; you can see the median values usually change in those; nothing is at the center all the time. Right? Because it denotes the line; this important line in the middle denotes the median values. Right?

The next thing we can print is how we can change the line width. Right now, the width of this gray color line that you see, maybe it's too small for you, so let's uh brighten it up. If we just set a parameter called as line width equal to 5, 5 is basically taking it from one; the default is one. If you put it into five, this is what it would look like. If you are a person who wants to look at it better and understand it in a higher depth of clarity, this is what you would looked at. Right? So previously it was like this; the line thickness was very thin; it was at one. Now if I do this with line width equal to 5, it just becomes bigger to look at. Right?

The last thing that we're going to check out is how we can uh plot based on our customized ordering on the x axis. Right now, if you can uh see here, line width equal to five, until here everything is the same, the code that we just wrote. The next thing is order equal to one year, two year, month to month. Here is where I say, hey, I do not want you to randomly print me an order on the left; I want one-year contracts on the middle; I want two-year contracts on the right; I want month-to-month contracts, and that is how I want you to pay, or that is what, how I want you to plot a graph. Right? As you can see, it is showing that all the one-year contracts and the ten years are here; all the two-year contracts are here; all the month-to-month contracts are here. If you want to change them around, all you have to do is just change the order here, and eventually you are given uh with that different order. Right? Customizing your order on the x-axis, guys. So this is something which you can do with box plots. Box bars are amazing; they are really fun to work with, and all of these things that we just checked out, let me zoom out. Right? Line plot, bar plot, scatter plot, histogram, or what's called as the disc plot, or even it's uh the other one that we checked out is the box plot. Right? So all of these add a lot of value to your learning, ladies and gentlemen. I hope you were clear with all of these. If you're not, wherever you are stuck, make sure that you give a second and instead of just copying whatever I just showed you here, explore your own; use different colors; use different palettes; use your own data set to map different xy variables to see if you can answer a question as soon as you plot it. Usually you will have a question in mind. Right? That's how I did it in two or three times; I looked at the graph and I'm like, huh, okay, so if the evolution is decreasing as the pokemons evolve, the weight is maybe decreasing, or their height is decreasing, whatever it is. Right? So I answered a question even without asking it; that is the advantage with data visualization, guys. So make sure you spend a good amount of time with data visualization.

Now let's quickly take a look at the agenda for the session. Now when you take a look at the course agenda for multivariate analysis, we are going to be starting out by taking an introduction to understanding what data analysis is in its basic terms. Right? Because once you have an understanding of what analysis is, then you'll be very curious to find out what are some of the various types of data analysis that we have out there. I mean these types; we are going to be dividing it down to the most simplest level; we're going to be assessing; we're going to be analyzing a lot of different things with respect to the various types of it, and as soon as we finish the types of data analysis, the right next thing that we're going to do is we're going to be jumping into the heart of the topic wherein we will start out by understanding what multivariate analysis actually is. Right? Now once you have a picture of what it is, once you have slight clarity knowing, oh, okay, so this is what multivariate analysis is, you will be very curious to understand, fine, I know what it is now, but why do we use it? What is the primary objective of using multivariate analysis? Right? So we're going to take a look at all of that, and once you understand what it is, where it is used and all of that, you'll be very curious to say, hey, okay, now what are some of the various techniques that I can go on to use now? Now that I know the objective and what uh multivariate analysis is. Right? So of course we're gonna have to take a look at the techniques as well, and once you understand the techniques and everything there is to know about uh, you know, having a good introduction to multivariate analysis, you will be a person who will say, okay, so what are some of these advantages of multimeter analysis that gives it an edge over the other types of analysis? Right? We're gonna have to take a look at that, and of course we will. And lastly, as you can see, practical implementation in Python. As always, guys, anything with respect to data analysis, anything with respect to uh, you know, performing any sort of analysis, right? It will always help you if you always take a practical example to, you know, go along with your theoretical learning and to make sure that at the end of the day you are applying whatever it is uh that you guys are studying and reading and understanding. Right? That way it will help you retain it better, and with the help of a demo you will understand exactly what it is to actually go on to perform multivariate analysis. Right? So I hope all of you all are excited as me for this course; let's get started with the first item on the agenda, which is introduction to data analysis.

Now, ladies and gentlemen, data analysis is something which is fantastic because you might not know it's called data analysis, but I can guarantee you one thing: you have used it in your life a lot of different times with things which you might have never thought you would have used data analysis for. Let me give you a good example here. Now think about a situation where there's a new TV show that is coming, or let's just say season five or season six of some TV show is becoming super popular; they dropped a trailer or a teaser or something; the entire world, as you know it now that everyone's connected via social media, everyone will be talking about it; people will be going bonkers about it. Right? When you think about it, so what do you do when you see a new TV show which says people are saying, oh, this is the best TV show I've ever watched; there is nothing that comes close to this; when people are raving about it and when people are talking highly of it, what do you do? Because here is what I do. Right? Once I figure out something is extremely popular and it is like uh, you know, growing at a lightning pace, I am sure there is something in that show that is making people want to go and check it out. Right? So I'm curious now; what I, I believe is, hey, okay, so if a lot of people are talking about it, there's something that the show is doing right, so let me find out if whatever the show is doing, if it is in my liking or not. Now let's be honest; when it comes to movies and TV shows, we have our own choices. Right? Now, for example, me, I love absolutely watching stand-up comedies, or if it's any new TV show or something, I usually prefer comedy or anything on the lines of that. Right? Now you might be a person who says, hey, I am into thriller and sci-fi TV shows and all of that. So what do you do? You quickly go to Google; you read more about it. Right? You are trying to understand what it is at its core level; you want to do research; you want to do uh uh you know you want to read some reviews on uh IMDB, Rotten Tomatoes, all of these review places to see what people are telling about it; you'll watch the trailer yourself; you'll talk to a couple of friends. Now what is this entire process that you're doing? You're trying to explore more into the topic. Right? Now look at your screen; it says, what is EDA? EDA stands for exploratory data analysis, and ladies and gentlemen, to a good amount of extent, you guys were doing exploratory data analysis there because you knew what the problem was; the problem was you don't have enough information about the TV show. Now you know how to solve letters actually by watching the TV show, but before committing to watching the TV show, you realize saying, let me do a little more research; let me do a bit of analysis to see if this TV show is for me or not, and eventually you did everything that you would do and then probably you either watched it or you didn't watch it. Right? So taking a look at understanding this in the data analysis perspective that we're taking a look at, right? It's basically you having a problem; you have certain data, and you know that the data can be used to solve the problem, but as it is, if you just look at the data, it will not really give you any sort of insights; it will not give you patterns just by looking at it. Right? You have to perform some investigations on it; you have to work on it deeper. Now if I give you a single break up, you know, a couple of bags of sand and a bag of cement and tell you, hey, build a house. Right? If I just say build a house, do you think the house is going to be built? Not really; you're going to need people; you're going to need some process further to actually use all the raw materials to build this. Data analysis is that raw material where you're actually investigating. Right? That is very important for you all to understand. Now there are so many things that is hidden in data that you will not realize as soon as you look at it with the naked eye. You can look at data with your eye and say, okay, well, there's not really anything that I can tell; there's a lot of rows, lot of columns, and you know we have certain data, and you can probably guess what might be happening, guess. Right? Now I can guess that it will rain tomorrow here in Bengaluru, but what, how sure am I? I'm just guessing; I'm just seeing it; it doesn't mean it has to happen, but that is where data analysis is very different here. When you make outcomes, when you predict something based on your data, when you're analyzing data and you say, here is the result of my analysis, basically what you're saying is, hey, I know that this might happen tomorrow; here is the data to prove that it will happen tomorrow. Right? Now there you're not guessing anymore; you have data ready at your hand to back everything that you have to say. Right? Now if I say, hey, tomorrow it's going to rain because this is the precipitation factor, these are the clouds; that's how it's moving around India; you can expect a ton of clouds to be on Bangalore tomorrow, and there's a very good chance it might rain. Now when I say it this way, you'll be like, oh, okay, it might rain, correct, rather than just me saying tomorrow it'll rain. Now think about it. Now this exact small example that I gave you, uh, you know, this kind of an approach to understanding data is what is making it a key part of millions and billions and trillions of dollars worth of businesses today. Right? Everyone around you, I bet everyone, right from the grocery stores near you all the way to someone uh, you know, in a very big company, maybe like Google, Facebook, or something; these guys are using data analysis on a day-to-day basis; of course not everyone is sitting and working with algorithms, creating models and all of that, but they are in their own way performing some sort of analysis. Right? Now you, you closely observed what I just said, some sort of analysis. So what are the various uh, you know, types of analysis that one can go on to perform when we're discussing about data analysis? Right? Now, ladies and gentlemen, you have to understand that there are three very important types of data analysis out there; it is not just multivariate analysis; we have univariate analysis; we have bivariate analysis; and we have multivariate analysis as well. Now with respect to univariate analysis, it is very simple; what does uni stand for? All right, uni stands for one; evenly stands…

For single, so whenever you are assessing and analyzing a single variable which will have an impact on the outcome, that is when you you are that is when we use univariate analysis. Now when you have to analyze two variables which will tell you which will have an effect, it should have a say on the outcome, that is when you use bivariate analysis. One variable done, two variable done. What if there is a case where we have more than two variables that we will have to analyze and assess? Again, you figured it out by now, right? It's multivariate analysis, which is basically uh used whenever you have to work with two or more variables.

Now ladies and gentlemen, do pause for a second here and try to understand and see if you can find any example cases of where you might have used univariate analysis, bivariate analysis, and multivariate analysis. Right, if you're just analyzing one thing to see if it is right or wrong, that's univariate. Now if you have two objects or something that you'll have to assess to see towards a common outcome, that is bivariate analysis. Now multimedia analysis, again, there's a lot of factors that are involved, right. In fact, let us discuss these examples.

Univariate analysis: When someone says, "Hey, analyze the change in the height of a particular person," in majority of the cases, how do you think it goes, right? So when they're young, the height is pretty much very less, and as they grow in age, the age of 18, 20, 25, 26, they reach a certain age and they tend to plateau at that age for a really long time. So it's a linear growth, right? It's not like one person grew 10 feet in a day; that really doesn't happen. So it's only one variable that you have to assess to give someone the height, right? You just have to assess, uh, measure them and tell them, "Hey, this is one of the variables; this is one of the uh, you know, important things that will help you get to the outcome." Simple. Now if I say, "Hey, analyze the change in someone's age," right, age is still a single factor, right? You don't have to consider anything else apart from uh, if you know the person's age, that is all. Let's just say required for the analysis result. Analysis says that if you only want to assess and analyze a h or another stage, this is an example, uh, this is what will lead to having one variable deriving and helping you get towards your outcome. This is univariate analysis.

Now take a look at bivariate analysis, right? Bivariate analysis is fantastic because here you will be taking a look at understanding two parameters at a time. Let me tell you, uh, so think about sport preferences in our country, right? So, uh, not just in our country, it's just a hypothetical example out here. Uh, whenever you're thinking about sport preferences, it is very common that you know guys would love to play, uh, uh, you know, cricket; guys would love to play football on the national international leagues; and the women see themselves, uh, you know, mostly in archery, uh, when you think, take a look at all these track and field events, which are again extremely difficult in my opinion, right? So if you say, "Hey, what are the sport preferences for male population and female population?" Look at this: you have two uh things that you'll have to assess and not uh, you have to assess your outcome based on, right? So male population, they will have their own choices; they will have their own sport that they like. Now if someone asks you a question saying, "Hey, at that point of time in 2005, what were the uh details of, you know, the male population, what games were they playing?" Can't you answer that? Yes, you can. If the same question was asked in the case of female population, again you would have analyzed a lot of different sports to get to the conclusion to say, "Hey, females in the year 2006 were playing these mini games." Simple, right? So you can do it. Another example for bivariate analysis is the GRE score based on your ages. Now when someone tells you, "Hey, I want you to take all the GRE scores of all the students who have written the exam, divided based on their age," now you have two things that you'll have to filter something based on. First of all, you have to pick up all the GRE scores; next, you have to pick up all the age of those particular candidates who have written that examination, and eventually then you can perform some analysis. There are two inputs; there are two variables you're assessing here: the score and the age at the same time, right? So this is bivariate analysis.

When you take a look at multivariate analysis, you have got the game by now, right? So you will have to have more than two variables to actually check if a person, uh, you know, to check if an outcome is being achieved there or not. In the example that you can see here, it says diabetes prediction using insulin, ah, BFI, etc., right? In fact, we are going to take a look at a practical demonstration on the same particular topic. It's an absolutely fantastic demo because here, uh, just think about the real-life diabetes uh situation, right? Uh, there is never only one reason that will lead to diabetes, right? There's many factors: there's age, there's genetics, there's insulin, there's your diet which matters a lot, there's your exercise habits, uh, you know, your blood pressure; there's a million things; there are so many variables, and if those variables fall into place, there's a very good chance that you might have diabetes, right? So this is again a common problem in today's world, and we are going to be using the power of machine learning to actually assess and analyze and perform multivariate analysis on it because at the end of the day we will know that, "Hey, diabetes does not just depend on one factor; it depends on multiple factors." So let's have each of those factors; let's perform some analysis on them, right?

So guys, these are the various types of data analysis that we have out there. Now coming to the heart of the topic: what is multivariate analysis, right? We took a look at examples; we understand that we are going to have to have more than two variables to work on it, right? Again, I can give you more examples on multivariate analysis because, as I told you, multimedia analysis, unless you search for it in your real world, you would not know you have used it. Think about it, right? Multivariate thinks about, talks about, always making use of multiple variables at a time. The same weather prediction example I spoke to you at the start: if I'm just telling you tomorrow it's going to rain, it's an absolute guess, right? It can be the sunniest day in the country tomorrow. But what if I tell you, "Hey, tomorrow it's going to rain because I assessed these many, these many parameters," for example, the rain might depend on humidity; it might depend on pollution; it might depend on precipitation; it might depend on a ton of different other factors, right? All these numerical factors or categorical factors that will lead to that outcome of it raining uh tomorrow. Now when that is the case, you have to be uh, you have to understand for a fact that it's not just one thing that lead to it; there's multiple things, and weather prediction is actually a very good example uh for all of you all to actually understand multivariate analysis, right?

Now uh, you might be a person who might be saying at this moment of time, "Okay, I know uh what multivariate analysis is, but where is it popularly used?" Well, we're going to be discussing a section just for that, but let me quickly give you some key pointers out here, right? Now whenever you go to an analyst and say, "Hey, what are the sales figures of my company?" Now you might be a person who might have just opened a brand new chocolate uh company; you're selling chocolates, right? Uh, you will know that during this time, September or October, November, December, uh, this is when people usually tend to sell a lot of chocolates, right? Because again, it's Christmas time, it's New Year time, there's Halloween; there's always something coming around uh these particular days. So you might be uh, you might be a person saying, "Hey, I want to use the last six months' data that I have to try to see if I can predict what my sales will be during Christmas, or even better, I'm going to give you the data from last year's Christmas sales, and can you guess how many customers that I can expect this year so that I can order so many chocolates for the customers," right? So when you think about it, you can ask any sort of questions where probably 20 years back you could never answer those, right? And now you can not only just guess or guess it out or something like that, you can have data to prove everything that you're saying. Now it is very easy to say, "Hey, you know, get 10 chocolates." How do you know it's 10 chocolates, right? You cannot just say this is the factor which might be affecting the sales of any particular company, or that there's multiple aspects; there's multiple variables; there are so many dependencies uh which will lead to a situation such as this, right? So it is very, very important. Multivariate analysis, you have to understand, ladies and gentlemen, is a very important part of this domain that we call as exploratory data analysis, or EDA. So the next time whenever someone speaks about EDA, I understand that multivariate analysis is a very important, has a very important role to play in that, right?

Now that we understood the various objectives of multimedia analysis, I think it is high time we dive right into the heart of the topic to actually understand what are some of the various techniques that we have in today's world to help us with multivariate analysis, right? Now ladies and gentlemen, there's many, many techniques of how one can go on to perform multivariate analysis, and as a data analyst, it is very helpful for you to know what each and every one of these techniques is. Now a common question I get whenever I'm teaching these techniques is, "Do we have to know all of this to an extreme amount of detail?" Not really, because whenever you you are given a problem statement, you will be using various techniques out there, let's be honest, but it doesn't add any sort of value for you to be a thorough expert in 100 different techniques when uh you might be just using three or four, right? Give it a thought. So whenever you require any sort of technique, know what it is, know where it is used, know this; it's very important for you to understand how to use the technique to solve the problem; that is where all the uh magic lies in, right? So let's quickly brush through; let's quickly summarize some of these really popular techniques that we have in today's world uh where multivariate analysis is popularly performed, right?

The first one on your list is multiple regression analysis. Multiple regression analysis, or MRA as it's called, is again one of the most commonly uh use multivariate techniques out there. Why? Because it is uh, it is very effective even from a non-technical perspective, and at the end of the day when you're performing multiple regression, right, it just solves all of your problem because it will help you assess the relationship that exists between one single dependent variable and multiple independent variables. Are there there is one single dependent variable; there's multiple independent variables around it. So if there's any sort of relationship that exists, this analysis method will show that out and bold to you, right? So when you have this kind of an ability where you know you have one dependent variable, multiple independent variables here, you can determine a linear relationship very, very easily, and eventually, to let's say we are trying to use this technique to predict tomorrow's weather, it's going to be very easy. What is the goal here? The weather. What are the parameters that will help for the weather? Uh, precipitation, pollution, uh, humidity, all of those factors, right? So you're building a linear relationship between that, and once you go on to actually do it well, you can predict tomorrow's weather, right? Simple. That is multiple regression analysis.

Now coming to talk about discriminant analysis. Discriminant analysis is again another very popular multivariate analysis technique, but here uh what we what do we do? Things are slightly different here because what we're doing is we're trying to classify all of our observations, and we're just creating multiple tiny groups, subgroups, bunches, whatever you want to call it, right? So we just take all the observations and we split it out into different groups. Now we can have a function that is that will basically classify all these observations that we have, right? Now once you have observations, you divided into groups, you will know roughly, vaguely what each of the groups or how each of those groups will behave, and then writing a function on it to make sure that you can you can sit and classify the observations for sure; you're no longer guessing; you have a structured way of actually performing that assessment; you can do it, right? So it'll give you this ability to understand what variables will have more impact on the function at hand. Now again, think about the diabetes example that I've been talking about, right? How do you know which has the biggest factor? Does having really high blood pressure cause diabetes? Does having uh really bad blood sugar levels cause it? You know, you cannot say you have a bad age or something; you can say maybe you have a bad body mass index, right? Personal fitness, maybe if that isn't up to the mark, and if you see if you are ob is a very obese or something like that, how do you know which of these has the highest impact, right? Using discriminant analysis, you actually have the ability to say, "Hey, this uh particular feature, this particular point of analysis might have that much, that much of an impact on your solution," and that is very, very important for you to analyze as an important technique we have today, right?

And then coming to another multivariate analysis technique which is super popular, which is basically called MANOVA. MANOVA, multivariate analysis of variance, is again a super popular technique which is used to assess the relationship between—note this—this is used to assess the relationship between many categorical independent variables and multiple numerical dependent variables. Independent variables are categorical, but the independent variables are numerical, right? Now when when this when is this the case? Give it a thought, right? I actually want you guys to analyze when is the situation where you will have a categorically independent variable and a numerical dependent variable, right? Uh, to give you uh to give you a hint to take a look at it, it says very popularly used to assess the market trends. We know how the market works, right? It's just charts up and down, but when you have to analyze how much something is varying with respect to another point, you're gonna have to take a look at the variance, and once you have to perform analysis on the variance, and if you have to perform multivariate analysis of variance, it's still called as MANOVA, right? It's a it's a very simple generalization of what the technique is, what the technique does, but this point where it says this is the relationship, it will help you assess this relationship between uh many categorical independent variables and metric-based dependent variables. This is something that you guys should understand, right?

Coming on to the next technique which is factor analysis, right? Now factor analysis is also another very, very popular technique. Now you might be uh, you know, taking a look at the trend here; I'm saying that this technique is very, very popular because data analysis is such a huge domain out there; each of these techniques have become popular because they are very good at solving certain problems for multiple domains out there, right? Now factor analysis is very different from all the various other techniques that we have today because see, in the other techniques that you will have, you will, we are trying to find relationships between variables, right? Dependent, independent, categorical, numerical, all of that, but here what we are trying to do is we are trying to understand the structure that is hidden in the data rather than saying, "Hey, I want a relationship, and uh, you know, when I say I want a relationship, basically I mean the..."

Data point: Uh, you know, finding relationships between each other right now that is very important for you to understand. And I bet you have heard about factor analysis, right? PCA, principle component analysis. I'm sure you might have read about it, heard about it, or something. In fact, what you should know at this moment of time is that PCA is a very, very good example of actually performing factor analysis. A point where you say, hey, I want to understand the structure of it; I want to understand why the data, uh, you know, why the variables behave the way it is, and how is it linking to each other. When you are hunting for that structure rather than just looking for a for looking for the relationship between the variables, you know what to use now, right? Factor analysis.

And then we have another type of multivariate analysis technique here, or it's basically called as cluster analysis. Now, cluster—what does the name cluster remind you of, right? Because in this particular case, when we are performing cluster analysis, we take a very, very large data set and we actually break the large data set into smaller chunks, smaller groups of individual elements, right? When we are performing cluster analysis, that's what we do: We take large data, we cut it out into small pieces, and we analyze, uh, those small pieces and we try to understand, is there any sort of similarity when we're trying to cluster? Is there any other characteristic which will group it into clusters, right? I'll give you an example. Think about apples, right? Uh, if I—the fruit apple, not the bran—now, whenever you're thinking about the apple fruit, there are so many ways; there's a there's a raw apple, there is a ripe apple, right? The raw apple is usually green in color; the ripe apple is usually red in color. Now, what if I say, hey, can you can you divide this? Can you cluster these apples based on if they have if they're ripened or not? What will you look at? You look at the colors here: This is green, unripe; this is red, right? You're going to do that, right? So it's a characteristic that you're trying to assess. Similarly, when you have so many different things, especially market segmentation analysis, where you will have a lot of these categories where you know your data is huge, it's absolutely huge, but you can, uh, try to group it down, try to simplify it and understand, hey, these are similar points of data that I can club; these are some of the characteristics which are making the data points look similar, so let me club these as well, right? And then later you perform analysis on it. How simple is that? But you have to understand one thing, ladies and gentlemen: Factor analysis and cluster analysis, as we are discussing right now, these are very, very popular, very powerful, and, uh, people have started to realize that it is way it pulls way more weight than people initially thought. So these are extremely popular, uh, for today.

The next thing that we're going to take a look at is some of the various advantages of multivariate analysis that we can have today. Now, to talk about the advantages of multivariate analysis, it's a small problem because there's a ton of advantages that you can talk about, right? This huge number of advantages which I can go on and on and on, but to actually simplify all of my thoughts into five, six very good points—that's what you see on your screen right now. The biggest advantage of multivariate analysis is, for a fact, that it will help you arrive at accurate, data-driven conclusions and insights. There's two ways of coming to a conclusion: One, you can guess it; the second one, you can have a conclusion and you can say, I have data to prove that this is the conclusion. Which would you pick if you were a client? If you were an individual? If you're an enterprise user? Which would you pick? You would, of course, want data-driven insights, data-driven conclusions. It will help you arrive at that, right? Not just data-driven conclusions, but accurate, super accurate, data-driven conclusions, right?

Now, the next point where multivariate analysis actually helps is to find out if there's anything wrong with the data, to find out if there's any anomalies in the data, and also to understand to auto understand what, how level of what is the level of consistency that you have with respect to data out there, right? Now, this is very, very important for all of you all to understand because people usually think that, hey, multivariate analysis: You have multiple variables, you analyze it, you give some conclusions, you're done. Not really. It'll help you in a lot more places. Think about the third point: Feature engineering. Feature engineering is actually a concept where you will try to analyze what parameters have what sort of impact on your outcome, very similar to multivariate analysis, or in fact, it's an initial step to multivariate analysis if you take a look at it from another perspective as well. So telling that, hey, these are the features which will help me: Age is a feature that will help in diabetes; genetics is also a feature that will help in diabetes, right? So you have multiple features that you're gonna say that this will, uh, have, uh, you know, an outcome; this will have an effect on the outcome, right? That helps again. Feature engineering is extremely critical, critical for machine learning, so give it a thought of how how advantageous this is, right?

And then data cleaning. Data cleaning is very, very important because, at the end of the day, again, if you consider a machine learning application, if you, uh, you know, if you just provide unclean data to a machine learning algorithm, it is not going to realize; it doesn't know what a car looks like; it doesn't know what a ball looks like; it has absolutely no idea, right? It understands only zeros and ones, and through structure it builds and understands and has a certain meaning, but if your data is unclean, uh, let's just say you're trying to collect the patient data; someone has asked, what is your blood group, right? So some person, by mistake, instead of typing in their blood group, they have written their date of birth, or they have written their mobile number, right? Think about it; it can happen; it will happen, right? We humans, after all. So when something like that happens and when you know that, uh, yeah, you know, the data data is incorrect, and as soon as you feed it to your machine learning algorithm, it doesn't know it's wrong; what it's going to do is going to use that to learn incorrectly, right? It's not its fault at the end of the day, but that will happen. So with respect to using multivariate analysis, where you can clean and pre-process your data better, helps a million folds for all the processes which come down the line after analysis, right? Superb. Take a look at the last point. The last point is handling underfitting and overfitting. Underfitting and overfitting are the absolute—it'll almost feel like, hey, these are small problems, but these are the biggest impacts on the outcome of whatever is being done, right? Underfitting is when your machine learning model cannot really, uh, go into the depth of analyzing and assessing, uh, what's going on with the data, so it really cannot dig deep enough to find out any structure, pattern, or anything. But overfitting is reverse: Overfitting, it has dug so much into your, uh, problem; it has dug so much into your data that instead of looking at useful data, it started looking at noise; it has it started picking up things which are which might be of no use, uh, to your outcome. And what does that do? That both in underfitting and overfitting, it ensures that the entity, your overall performance is coming down; your machine learning model is not going to be performing well. Why? Because it's doing all of these things, right? One, in one case, it has absolutely no idea what sort of data to pick up and work on, where in the other, uh, kind of case where it's working so hard at a granular level that it does not know; it does not able to differentiate between noise and it does not able to differentiate and it's not able to pick up anything useful in both the cases, right? So to avoid all of that, to help handle it, again, you have multivariate analysis.

Okay, guys, fantastic. We've come to a point where we can actually take a look at that fantastic diabetes demonstration that I just mentioned to you all, and let me quickly jump to Google Colab. Google Colab is nothing but a Jupyter notebook that is hosted on the Google Cloud Platform. Now, you guys can code this in any other, uh, you know, any other choice of IDE that you want; your PyCharm, you can use the regular Jupyter notebook; it's basically a simple Jupyter notebook, but it's on the Google Cloud, right? So let's quickly get into that. All right, guys, as you can see, I've opened up my Google Colab out here, and all we're trying to do is we're trying to perform the diabetes detection; we're trying to understand how we can get a machine learning model to figure out what are the various factors, what are the various points of analysis that will eventually lead to the result, right? Now, ladies and gentlemen, the goal here—one thing you have to understand is that the goal here is not to show you how easy—of course, one point you have to understand that to achieve this particular demonstration, it's very, very easy, but I understand that some of you all might not be, uh, you know, might not be up to speed with respect to these libraries such as NumPy, Pandas, Matplotlib, Scikit-learn. So let me quickly give you a glance of what's happening with respect to each point of code, but you actually try to analyze and assess the, uh, MBA, the multivariate, uh, analysis aspect of it, right? Well, the first step of any analysis out there, uh, we're gonna be have we're gonna have to import all the libraries, right? NumPy is called as numerical Python; it's basically used to perform all these numerical related activities when we are working with Python. Pandas will give us two fantastic data types; it's called as, uh, the Pandas Series and the Pandas DataFrames, which are again very, very important for us to work with. And then we have Matplotlib and Scikit-learn. Matplotlib and Scikit-learn are critical libraries when it comes to data visualization. If you have to plot anything, uh, with respect to using the graph charts and all of that which we are going to be using in this small demo, so you're going to require that. And then we have Scikit-learn; Scikit-learn also stands for Scikit-learn, which you might have heard; it's a very, very important library for machine learning, right? Now, as soon as I hit this play button—it's very simple to actually work with the Google Colab here. Now, as soon as I hit the play button, it's going to take a second, and as soon as it stops, basically it has executed this, right? Now all these libraries are ready for me to use.

Now, the next step is for me to actually talk to you about the, uh, data set, right? To show you what the data looks like. But if I just hit play on it, it's gonna give me an error. Now it's going to say, hey, where is that file that you're asking me to open? Well, it's actually in my local computer here. Let me open up these files. Now I have to upload it to the Google Cloud so that it understands, right? Now this entire Python environment is on the Google Cloud. Now that I've uploaded it, I'm going to close it; I'm going to hit play again, and as soon as I hit play again—done. We have actually uploaded our data set; our data set is ready. So let's quickly take a look at what the data set is all about. Now I've now printed the first 10 patients' worth of data. Now every row you see here is one patient's data, okay? So we have multiple, uh, ones; we're going to see how much we have in a while. Now, what are these various factors that cause, uh, that cause diabetes, right? One is your glucose level, uh, the blood pressure levels, what is your skin thickness, what is your insulin level, what is your body mass index, what is your age, and eventually the last column will be the outcome column, right? Now, if the outcome is zero, uh, it means this person does not have diabetes; if the outcome is one, it means that person has been diagnosed with diabetes, right? So we already know; we know that, hey, this is the detail of the person, and if the person has diabetes or not, we know. But what we are going to do when we are supplying the data to our machine learning model is we are going to just remove the last column and say, hey, now you predict if there is any sort of, uh, use your multivariate analysis result to see if you can find out correctly if the person has diabetes or not, right? So that's about it. Uh, BMI, again, when you think about it, I think 25 is the normal value, and I think about 28 or 30 something is where you are called as an obese person; anything less than 25, uh, you know, if it is like around 20 or something, your malnutrition is something along the lines of that, right? You guys know this.

Now, the important part that you have to be asking me right now is, hey, okay, so you just printed out it for the data for 10 people; how many people do we have in total to analyze all those data? As you already know, to perform multivariate analysis, we require a large amount of data, correct? Yes. In this case, as you can see, for each of the patients, uh, for each of the columns that we have, it's just showing a parameter which is showing a variable; we have 768 entries. What this tells me is that there are 768 patients' worth of data that I'm looking at right now, okay? So if there's 768 people who are, uh, you know, going on to giving me some data points out there, what are some of the things that I can tell as of now with respect to a statistical description? Well, the first, uh, the row here, it says count; count is basically telling me how many entries there are with respect to that column. As you can see, 768 everywhere, so we don't have any empty, uh, columns or rows out here. And then it is giving me the mean of it; mean is basically the average, right? Now, in the entire 768 people, you can see that the body mass index average is actually 32, which means that overall there are—if you consider this, uh, the data set in general, this hypothetical data set—the body mass index is slightly higher than what we would prefer, which is another 25, 26, 27, right? You can take a look at all that; you can take a look at the standard deviation; you can find the minimum values; you can find out what is the value at the 25th percentile, 50th quartile, 75th quartile, and eventually even the maximum value as well, right? So there are various things that you can figure out from this. Looking at it, as soon as you take a look at it, you'll be like, okay, so this is what is happening, uh, with respect to all of these, uh, individual entries, right? This is, uh, taking your entire data center instead of walking you through each of those 768 entries; this is me giving you a statistical aspect of it, say, hey, look at this with respect to one command, right? `data.describe` is all that we use; if the respective data are described, done; we just used it, and we realized that we can find out so many things about our data set, right?

Now, the next thing is the part where we'll actually be performing some pre-processing with the data. Now, thankfully, in my case, I do not need to perform any sort of pre-processing here because all of these values—wherever says false, it means that there's only numerical values that are found—is basically telling is it not a number? Is it anything apart from a number? So wherever it says false, it means it's a number, correct? Now, again, as I told you, right, when I ask someone, hey, what is your body mass index? Please import it into this column; they might have written O positive blood group, right? Of course, they don't do it on purpose, but it is done by mistake. And if it says true here in many of the point of cases, it means that there is a wrong entry there that you have to go and fix it up. Now, in my case, as I told you, this data set, this hypothetical data set, is already cleared; I don't have to perform any sort of preprocessing on it. The next thing that I can jump directly is to understand this concept called correlation. Correlation is a beautiful concept, especially for multivariate analysis, because here it will tell you how much of dependency exists between two variables, right? First of all, take a look at the heat map on the right; if it is, of course, as positive correlation, there's negative correlation, all of that. To simplify it for your case, whether it's just going to be talking about positive correlation here. Now, whenever there is this chart that starts from zero point zero all the way to one, you can see the color lightens. When the color lightens in the legend, it's basically telling you that whatever, if it is at that color, that is the amount of positive dependency, the correlation that exists between, uh, these two variables, right? For example, pregnancies and age, right? Age is your pregnancy is here, so this is the box that we're talking about; you can see that this is light in color compared to all the other boxes, right? So pregnancy and age has a good amount of correlation there, and very, uh, uh, you know, you could figure it out just by looking at your data. Now you might be looking at the diagonal elements and say, hey, those are white; does it mean it's exactly equal? Yes, it is exactly equal; I'll tell you why it is white because if I have to compare the value of glucose—now this is white—what are we comparing it with? Let me scroll down; I'm comparing it with glucose itself. Glucose compared to glucose, for any person, if it is like 300 versus 300, isn't it exactly the same thing? Yes. Similarly, for blood pressure, skin thickness, insulin, BMI, all of these different parameters, you have to take a look at it where in each and every one of those, when you compare it with the exact same thing, it's always going to be 100 equal, right? So all the diagonal elements will always be, uh, you know, be white in color, right? So you can you can analyze various points here; for example, take a look at insulin and skin thickness. Insulin has some sort of relationship with skin thickness because this is a lighter color compared to all the shades around it, right? So like this, you can actually sit for a couple of minutes, try to analyze and try to write down what are some of these, uh, correlations which are which will be very important for you to later—not only understand these are the variables that have the outcome with respect to correlation—you can say these are the variables that have so and so outcome, right? You can give, attach a numerical metric to it to say, hey, uh, you know, pregnancies are 80 related to or age factors or something like that; you can always go on to perform, uh, all of that with respect to a simple correlation matrix.

Now we know that we have 768 people; how many people, uh, how many people do we have who have no diabetes, right? As soon as a couple of lines of code, three or four lines of very simple logic will tell me that there are 500 people who don't have diabetes; another piece of code tells me that there are—of course, if there are 500 people who don't have diabetes—it is 268 people who have diabetes. Now you might be saying, hey, wait a minute; I didn't understand how did you figure it out? Let me scroll up to the data set, and you can see this last column, one and zero, one and zero, right? So all I have to do is I have to filter everything here; if I just have to pick up everything that is one and count how many are there, I have to put everything which is zero on the other side, count how many are there; that will give me the total number of people who have diabetes and who do not have diabetes, right? That's all I did. But if you still want a simpler explanation, as always, with the help of data visualization, you can take a look at it. If the outcome is zero, it means the person does not have diabetes; if the outcome is one, it means the person has diabetes. Now you can see that it it is a sharp mark at 500, means 500 people don't have diabetes, and here are somewhere around 268 people, roughly around 268 to 70.

People have diabetes right now. You can just look at this graph, and within two seconds, you can figure it out. Uh, if you're a person who says, "Oh, I don't really like to, like, write all of this code just to find that out," two seconds to create a, uh, you know, try to create a plot, and you will have all the details in front of you, right?

The next couple of steps are pretty simple now. Uh, multivariate analysis is where—now now is where—machine learning takes over a bit. The next thing that we need to do is we need to break the data apart into two pieces: one is the training data set, and the other one is the testing data set. Why is this done? Well, this is done to make sure that, you know, we have brand new data that our machine learning algorithm can test later on. It's just like a college examination: you read using your textbooks, you close your textbooks, and eventually go write your examination from your memory. Why is that done? Let's start to see how we have learned what it is that you have exactly learned from those books; we want you to reproduce your learning, right? Similarly, if we divide stuff out into two pieces, and if we hide one piece for a later, uh, verification stage, we can understand, and we know that it has learned it better, right? Because if I give it the training data right now and later send the same training data again, it will give me a 100% result; it will say, "Hey, everything is correct." That's not what we want, right? It's like having your textbook next to your examination; we don't want that.

So basically, what we're trying to do is when we are providing the data to our machine learning algorithm, that last column which said outcome, where it had the details of if a person has diabetes or not, we're gonna drop it; we're gonna remove it straight away, and we're going to break the data into two pieces: one, one chunk, one basket consists of 80% of data, while the other basket consists of 20% of data. Now, for all the people who are not used to machine learning, if you're a person who's thinking, "Oh, I'm going to require hundreds of lines to perform machine learning," not really. As soon as I click this, uh, particular snippet, you are—these two lines that you see—`l = logistic regression` and model fitting is a place where we actually provided it with respect to the training data, and as soon as this is done, the model has actually finished learning everything. The next thing is where I give it that test when I tell it, "Hey, here is new data; tell me if you can predict this correctly," and with again one line of code, that is done. And after that, you have one more—all of these are basically statistical concepts—it's the measures of accuracy here. Uh, there are concepts such as true positives, false positive, true negative, false negative; those will actually be used to assess and gauge the performance of your machine learning algorithm, but for simplicity's sake, let's have an accuracy score. Uh, it says the accuracy score is 74.67%. All that it is trying to tell us is this machine learning algorithm could figure out what are the various factors that's going to lead to analysis that's going to lead to diabetes, and we make sure to take a look at it manually with respect to performing MBA over using correlation.

Once we did that, we figured out how many people have diabetes, how many people don't have diabetes; we created a small graph; we used logistic regression; and at the end, we printed out an accuracy score to say, "Hey, by not only just performing odd multimedia analysis, I have also used the result of multivariate analysis to feed it into a machine learning algorithm and get the machine learning algorithm to predict if something—if it can find out if the patient has diabetes or not." How fantastic is this, right? We started out by thinking, "Hey, this demo will only be talking about multivariate analysis." We also made sure to complete, however, a bigger view at the picture to tell you—not to tell you exactly—not just how a multivariate analysis works, but again, whenever you think about where it is used, this is where it is commonly used; it is attached with either machine learning or deep learning to drive insights based on those results. It's very important for you to understand, and as I just told you, right, the goal here for you to is actually to understand the structure, to understand how we try to attack this particular problem and solve it. And, of course, there might be some of you who are watching this course right now who might be saying, "Hey, I don't know why you've written this exact line like this." Yes, you're going to require a bit of introduction to Python to understand how Python syntax is written to actually go on to figure all of this out in detail, but that—that's not the goal here. The goal here was to show you, uh, a multivariate analysis and also give you an extended view about using machine learning there as well.

Well, guys, with this, we have come to the end of the course. Thank you so much for watching. We started this one out by taking a look at understanding what data analysis basically is. Once we figured out what data analysis is, we took a look at the various types of data analysis we could perform. Once we understood the various types of analysis, again with very good examples, we honed in on multimedia analysis. We started understanding what is multivariate analysis, where it is used, what are some of its popular applications, what are its very good techniques that are used everywhere around us, what are the advantages of it, and we also made sure to take a look at practically implementing the concept to showcase to you guys that, you know, uh, in a practical situation, this is how you would be performing multivariate analysis. Now, there's so many other use cases; there's so many—it's—it's literally—or an almost infinite outcome where you can find a problem statement and perform multivariate analysis on it, right? Because usually that is the case; that is how it happens in the real world, and the real world has realized it, and that's the reason why we are seeing data analysis to be blowing up like we have never seen before, because people realize now that having data, performing analysis like this will not only help them clear out any, uh, you know, something might have gone wrong in the past and you're trying to fix it, but it is also giving you an ability to forecast what's happening in the future, and not a lot of domains give you that, and that is the power of data analysis, right? Okay, guys, superb. Thank you so much for watching. My name is Anirudh Rao. I hope you're cured with everything that I have covered in this particular course. As always, see you on the next one, right? Have a good day. Cheers, guys.

In this particular part, we are going to talk about actually—or outlier detection. What is anomaly? What is actually meant by outlier detection? We will talk about everything, so let's see and let's go to objectives and let's see what are the things we are going to talk about in this particular part. So first, we will start with what is statistics, right? So this angular detection or outlier detection is nothing but the basic part of statistics you need to know. Then we will talk about what is population. Then from there, we will have a look at what is parameter and sample. These are nothing but basics of statistics. Then I know you all have an idea about what is mean, median, and mode. Again, for your recap, we will talk about that. Then we will see what is normal distribution. Then types of analysis and statistics we have. Then we will start with what is an outlier, what is interquartile range—in short, which is known as IQR—then we will see what are the, you know, interquartile range like upper range or lower limits; what it is, we will talk about, like, you know, in depth for interquartile range. Then we will do a demo how to handle outlier in the data set, right? So let's start with our—what is statistics? So the first question is: what is statistics? So what do you think? What is statistics, right? So statistics is a part of integrated applied mathematics which deals with data, right? It helps to collect data and analyze them properly. Okay, so if we know math, it's an integrated applied mathematics, right? So basically, it helps us to understand our data in a more better way. It gives us the idea that, okay, maybe your data is coming from this; it's like, you know, these are the insights you can actually have a look at; that's why we go for statistics, right? In a, you know, very basic part, as a data scientist, you need to have an idea about what is stats, right? How can you understand your data using statistics concepts? So it helps to collect data and analyze them properly. From there, with the help of statistics, we can read the data and organize them in order to get the hidden information from them. The main reason behind it to use statistics that gives us really some, you know, uh, and like, you know, we cannot even think of by looking at the data, but actually the data has those insights or the information, so they are nothing but our hidden information. So this is what statistics stands for, and that's why we try to use all the concepts from stats in data science. Two main statistics concepts are used to process the complex data to get the insights from them using mathematical computations, right? So it's basically a game; it helps you to process your, like, you know, it's a complex data, and it gives you a really good insight from them just by using stats. So let's start with, like, what are the terms we have in statistics? So this is the slide we talked about. You can see stats, what—what is stats, why do we need it, and why we say while it comes to data science or artificial intelligence, you need to know the basics of stats, right? You—I think you got all the—I like—idea about why do we need stats. So please let me know in the comment section if you feel like—like you are not yet convinced that why do we need stats, right? Then we will talk about what is population. The term population in statistics is used to refer to the total state of observation. Okay, let me tell you what actually I mean by say total state of observation. So certain state of observation is nothing but when you have a data set, right? Or you can think about the example of the whole population of the world. Yes, that's what population stands for: the whole population, the whole total number of human beings we have in the world, right? So suppose again we want to study a diabetes data, say, to understand the symptoms and the other factors in the whole data set is referred to as population, right? Now, if I just go and talk about what is parameter, so parameters are referred to as characteristics which describe the population, right? It helps us to describe our, like, data set, right? So this is what stands for parameter. Parameters are like average or percentage would help you to describe the entire population. Suppose you have a diabetic status, age, and you want to understand what is the average age of being diabetic; I know how can you do the average. So this average is going to represent your whole data set, right? So you can understand what is parameter. Mean and standard deviation are two common parameters of population, right? Now, example I have given you already: average of age, average age for being diabetic is the parameter for whole diabetic data population. Now we will jump into what is sample. So if you have a lot—like, you know, we have the whole population in the world. Now, if you just want to talk about West Bengal, or if you just want to talk about India, right, then the India will be your sample from the whole population, from whole world population. This is what sample stands for. Sample is basically a small part of the portion of the large population. Now, suppose from the whole diabetes data set you picked 100 rows of information to do the analysis; that 100 rows of information will be referred as sample, right?

Now let's move further; we will talk about what is mean. The term mean is referred to as an average value of the whole population that we already know. Now, if we talk about median, so mean is the middle value of the data when your data is sorted in the map. Yes, to get the mean—like, you know, median value—you have to make sure that your data is sorted. Now we will talk about mode. Mode stands for the most occurring element in the data set. So mean, median, mode is a very basic thing, right? We start like learning those things when we are kids, so still if you have any doubt regarding what we have talked about—why do we need statistics, what is population, what is sample—then we talked about what is mean, median, and mode. Now we will talk about what is normal distribution. The normal distribution is a probability function which describes you how the values of the variable are distributed. Now we will say why do we need it. This basically gives you an idea about your data points, who are your data points. Suppose you are talking about diabetes dataset; all the rows you have about the description or the information about the person, these are your data points. Now, this normal distribution is basically comes from probability function, and it helps us to describe how our values of a variable are distributed properly to give us the idea what is the range of our data set, right? This is what normal distribution stands for. Now, properties of normal distribution: the mean, median, and mode all are equal, right? The curve is symmetric at the center. Then this is also referred as Gaussian or Gauss distribution, right? The parameters of normal distribution: mean, standard deviation—like we have the parameters of normal distribution that is mean and the standard deviation, right? Now we will talk about the types of analysis in statistics. Basically, we have two types of analysis: one is descriptive statistics, and second one is inferential statistics. What is descriptive statistics? It helps to describe the data in the mathematical or a graphical way. Now, when it comes to inferential statistics, inferential statistics speaks to data into samples and applies probability to arrive at the conclusion, right? This is what descriptive statistics and influential statistics stands for, right? I hope you understand what is actually meant by descriptive statistics and inferential statistics. So, as the name suggests, again, as I recap, describe the data; you need to have descriptive statistics. Now, if you want to use, you know, speak to your data to samples—I hope you understand what is sample and how can you apply probability to arrive at the conclusion—we use inferential statistics, right? Now let's move further; we will talk about what is an outlier. Outliers in the data set are referred to as the unusual values which can distort and violate statistical analysis. Now you will ask me what is outlier. Maybe in the class there will be always an outlier. What is an outlier? Suppose we have students who are mediocre, or we have a student who is really good, like the first boy or the first girl, right? So they are nothing but the outlier, right? So this is what outlier stands for. Now, outliers are basically experimental errors in the data set, or like in the case of, you know, data, it's experimental error, but in the case of—what example I get—it's not an error, right? Now, some outliers are good for data set to detect anomaly, like detecting for transaction. Now you will ask me how can we detect a fraud transaction. Let me tell you, so basically detecting fraud transaction is like, suppose you have a data set, well everyone has done the credit card fraud analysis, or everyone has the idea—like, you know, description or the information about their credit card—suddenly you'll see if there is a sudden spike in the, you know, transaction. When a person is used to, you know, buy something between 10,000 to 20,000, it's—he is going to buy something in 10 lakhs; it's nothing but the outlier, and there is a possibility it's a fraud transaction. I cannot say it's not a pro transaction; it's a—so it can be an immense transaction, but yes, but there is a possibility as the outlet is there, we would need to give a good look on that, what is happening. So this is how we actually use our outliers. Now, its effects are mean and the standard deviation of the data; that's obvious, and the data and the most of the machine learning techniques does not perform good with outliers. If you see outliers is not needed, you can just cut it off because it's—it's gonna affect your model, and it's gonna affect your data, mean, standard deviation, everything, right? Now we will talk about what is interquartile range. Interquartile range divides the data set into quartiles to measure the variability and the spread of the data set, right? So what is that? It's basically helps you to divide your data set with four quartiles to measure that how your, like, data is, you know, spread; what is the variability; what is your data points to give you the every possible way information about your data. Stick the data set into four equal parts, but in sorted manner, right? So Q1, Q2, Q3, Q4, right? We have—we will have four parts. So Q1, Q2, Q3 are called first, second, and third quartile. Now, what is Q1? Q1 is the 25th percentile of the data set. What is Q2? 50th percentile of the data set, and Q3 is 75th percentile of the data set. This is what IQR stands for. Now, what is the formula for IQR? That's nothing but third quartile minus first quartile. Now we will talk about what you—what are upper and lower limits in interquartile range. Yes, you have to have an upper limit and lower limit. Now let's have a look at what is upper limit and lower limit. Now, the—we have, like, upper limit and lower limits in the interquartile are basically the ranges where your data points are lying. Suppose it starts from 0 to 15, so low upper limit is 15, and your limit is 10—sorry, your limit is zero, right? So the formula to find the lower limit is basically Q1 minus 1.5 into IQR. Now, to formula to find the upper limit will be Q3 plus 1.5 into IQR. And what are the concepts we have learned about? What is statistics? What is sample? What is population? What is mean, median, mode? What is IQR? We will do everything, right? And how can you find outlier as well, right? So let's start with the data set demo, and do let me know if you have any doubt in the comment section.

In this case, we are going to use a dataset from Kaggle which is known as Pima Indian Diabetes data set, right? So this is your data set link; if you just click over here, you will redirect to the data set, right? Let me do that. Okay, so again it's a free data set from Kaggle where you have the information about diabetes patient. Can you see? This is how you will redirect, and you will directly download the dataset from here. It's from Kaggle, right? Can you see? This is what the dataset is, right? Now we will talk about what is this dataset all about; what are the thing we have. The Pima Indian diabetes database is the—is from National Institute of Diabetes and Digestive and Kidney Disease. It consists of nine variables; we will talk about that, out of which outcome is the target variable. The objective of these data says is to aid in painting if a person is likely to have diabetes or not. What are the different types of diabetes we have? The most common type of diabetes are type 1, type 2, and gestational diabetes, right? Now, what is type 1 diabetes? If you have type 1 diabetes, your body does not need insulin. So we all know about diabetes that it doesn't stop making insulin; your immune system attacks and destroys the cells in your pancreas that make insulin. Now, type 1 diabetes is usually diagnosed in children and young adults, although it can appear at any age. People with type 1 diabetes need to take insulin every day to stay alive, right? Now, it comes to type 2 diabetes. If you have type 2 diabetes, your body does not make or use insulin well. You can develop type 2 diabetes at any age, even during childhood; however, this type of diabetes occurs most often in middle age or older people. Type 2 is the most common type of diabetes. Now, if we talk about gestational diabetes, gestational diabetes develops in some women when they are pregnant. Most of the time this type of diabetes goes away after the baby is born; however, you have gestational diabetes; you have a greater chance of doubling developing type 2 diabetes later in life. Sometimes diabetes diagnosed during pregnancy is actually type 2 diabetes. Other

Types of diabetes less commonly include monogenic diabetes, which is inherited, and slightly fibrous-related diabetes. So these are the types of diabetes. Maybe you think why I'm talking about diabetes? Just to give you a heads up if you do not know what diabetes is.

So we have talked about different types of diabetes. Now we will see what the data description we have in our data set is, what the nine variables we have are. We have pregnancy—the number of times the patient got pregnant. Then we have glucose—that plasma glucose concentration. Then we have the column called blood pressure; it's like, you know, we all know that is BP, right? Then we have skin thickness—that triceps skin folds thickness, right? Then we have insulin—that two-hour serum insulin. Now we have BMI, that is body mass index—weight in kg per height in, like, to the power of two. This is how we calculate BMI. Now we talk about diabetes pedigree function, which is nothing but a diabetes predictor function. Then we have age and outcome, that will be your class variable—0 or 1, right? So these are the total columns we have. Now we will say what this outcome stands for. This outcome is basically: if your outcome is zero, this person is not diabetic; if the person is diabetic, then the class variable will be one, right? So this is what the different types of columns we have.

Now, if we go further and if we have a look at what the libraries we need to import are, then we have talked about pandas, NumPy, Matplotlib, Seaborn. We are going to import these four libraries, right? We have written `import pandas as pd`, `import numpy as np`. And if you know why we use `as`, please do let me know in the comment section. I would love to hear from you, as we have talked about this particular `as` so many times. So what actually I mean by `import pandas as pd`, right? Please do let me know. Now, uh, like after you like write this comment, then again start the video and let's see what you have written—it's right or wrong. And guys, if you like the video, do not forget to like it; do not forget to subscribe to Great Learning, so you will get updated with all the new content we are going to come up with, right? And please do let me know in the comment section if you like the video; if you like the video, what are the things you want to learn from us as well, right?

So now we will do what? We will load the data set. How we do the loading of the data set? We can write `data = pd.read_csv`. What is a CSV file? Comma-separated value, right? Yes, this is what CSV stands for. And then we have written `diabetes.csv`, right? So you have your data, and if you just want to read them using pandas, you just need to use `.read_csv`. So my data is a comma-separated value, right? So CSV stands for comma-separated value. Now if you ask me how can I upload data in Google Colab? I have talked about this so many times, but for your reference again, just click on this particular box here; you will get an option for upload to session storage. It will redirect to your local system, and you just need to upload the data set from there. So this is how I have done it. Now here, `pd.read_csv`; you have to give the path. How can I get the path of the data set? Now, if I just click over here, you get an option for copy path. Let me just copy the path. Now just copy-paste it here, right? This is how this copy path option works, and this is how you can read the data set using `pd.read_csv`, and you can upload the data set in Google App, right? I hope you do not have any doubt so far about what we have talked about. Give me a quick confirmation on that as well, right?

Now we will see the basic part of pandas, and then we will jump into all the, like, you know, all the techniques we have learned, and we will see how can we implement them using Python. Now we have written `data.head`. Why do we write `data.head`? I have talked about this again, like, for so many times, why, like, while we were talking about `pd`, like pandas, right? Okay, I have asked you guys a question: why do we use `as pd`? `as pd` is nothing but giving the alias to your library, right? So if I don't want to use, uh, like pandas all the time, I can easily call them as `pd`, right? It's a nickname to your library. You can see it in layman's terms. Now, if we just write `data.head`, it will return you the first five rows of your data. Let me just execute that, right? So in default, it will return you the first five rows, right? Now, if you say, "No, I want to see the first 10 rows," how can you do that? You just need to do `data.head`, then you can specify the number—the first 10 rows or the first 15 rows you want to see. You can—you will get a total of that many—0 to 14—that many data sets, that many rows, right? So this is why we use `.head`. Now comes `data.tail`. Why do we need to use `data.tail`? `data.tail` will give you the last five rows of your data set. I hope you get the difference between `head` and `tail`, right? So then, as the name suggests, basically they will give you the first five rows of your data. Now let me just execute that. Yes, this is what you get—the last; you have 768 total columns, you have total rows you have. Now I want to see all the nine columns we have talked about—only the columns. How to do that? You have to write `pregnant`; like, you have to write `data.columns`. So the `columns` will return you all the columns you have in your data set. You have columns called pregnancies, glucose, blood pressure, skin thickness, insulin, BMI, diabetic predictive function, age, outcome, right? So this is what you get after you write `data.columns`. Now I just want to talk about what my data shape is. Total columns I said nine; what, like, how many rows we have? In order to get that, you have to write `data.shape`. So after you write `data.shape`, you will get the total shape of your data—that is 768 comma; so total rows you have 768, and columns you have nine, all right? So we have already talked about this `data.head`. Now we will see how can we, uh, like, you know, treat the missing value detection and treatment. In this part, we are going to do the missing value detection and treatment. So how can we do that? So in this case, how can we understand we do have missing values? The following values in the data sets are considered to be missing values. Suppose you have blank values or you have `None` or `Null`, or some continuous columns might have zeros to indicate missing data. This is where we can understand this particular thing that it's having a missing value. Now how can you get that? Let's start by checking the count of the records in each column of the data set. So if the count of the record is lesser than the total number of records, then save like sub-sub-768, we can conclude that there are missing records, right? We have written `data.info`, right? Can you see what I get? The count of the records for all the columns is 768. This indicates there are no blank columns in the data set. Yes, you do not have any blank columns. Now if I just want to understand, do we have any null values? How to understand? We can use a function called `isna` function. This is a function used to understand: do you have any null values in your data set, right? Let me just execute that. What it's going to return? It's going to return true or false, right? If you don't have any null values, then it will give you false, but if you do have, it will give you true, right? Suppose you have written `data.isna().any` for each and every column; it's going to give you: do you have any null values for this particular column? So in the case of pregnancy, you said no; it said basically no for glucose; it said no for blood pressure, skin thickness; for any of the columns we do not have any null values. Now if you just want to get the summation of the null values, right? So there is an option we have: `data.isna().sum`. It will give you the total number of values, like total number of null values you have in every column. So you can see for all the columns you have 0, 0, 0; that means you do not have any null values. Since all the predicted columns are continuous in nature, there might be a chance that zeros in these columns indicate missing values. Let's say: do you have any zeros? If I just write `data.describe`, that—what is the describe function? A describe function, from the, like, you know, it's going to give you what the total count of the data is, what the mean of your data set is, what the standard deviation is, then what the minimum value is, then your interquartile range, and the maximum value. Now if you look closely, you can see the minimum value for pregnancy is zero; that can happen. The glucose is zero—will that be possible that a person has zero glucose? No. A person can have zero blood pressure? No, it's not going to happen. Skin thickness zero? Insulin zero? BMI zero? It's not possible. Even so, you actually have null values or that is nothing but representing as zero. Let's see what I have written here: From the above description of the data, we can see that columns pregnancy, glucose, blood pressure, skin thickness, insulin, and BMI have a minimum value of zero. It makes sense to have zero pregnancies, but it does not make sense for other mentioned variables to have a minimum value of zero. So we can conclude that glucose, blood pressure, skin thickness, insulin, and BMI have missing data; that zeros in the column should have been replaced with the median. We always replace with the median since the median is least affected by outliers, as we have already talked about outliers. Now if we just do `data['glucose']`, so basically I just want the column for glucose. I have written that `data['glucose']`; you will get all the data from the Google, uh, like from the glucose column, right? Let me just execute that. Okay. Now if I just write, "Replacing the zeros with NaN," we need to import one particular part from NumPy that is known as `import numpy`. The records that have zeros in the column glucose, blood pressure, skin thickness, and insulin, and BMI will be replaced with NaN, to understand how many NaN values we have which are hiding behind zero, right? That's why I have written `data['glucose'] = data['glucose'].replace(0, np.nan)`. So if you want to replace any value in the data set, you have to use the `replace` function. Now I have 0; I want to replace 0 with None, right? So this is why we use `np.nan`. Now if I write `data['blood pressure']`, the same way I have written `replace(0, np.nan)` for skin thickness, for insulin, for BMI, right? You can see this is how we can do this. This is replacing 0 with `np.nan`. Now if I write `data.head`, so this `data.head` will give you that, like, suppose you have zero; you can see in insulin, normally you can see for the first three records has NaN values; for skin thickness has NaN values. So basically these were replaced; like the zeros were replaced with NaN. Now if I just write `data.describe`, you can see there is a minimum value for glucose is 44, for blood pressure is 24, so we do not have any zero values because as they are already replaced with NaN. So this is how you can find out your missing value, but finding out missing values is really problematic, and also you need to keep your mind working, as I said: working with data is easy, but understanding data is not easy. You have to first understand the data; you have to make sure what the information you are getting from the data is worthy of getting, right? So these are the things you have to keep in mind while you are talking about the exploratory data analysis, which is learning more. Again, our module name is anomaly detection or outlier detection, right? So I have written `data.isnull().sum`. So you can see the total number of null values we have; we will get the summation of that. So for the case of glucose, we have five null values; sorry, for the case of glucose, we have five null values; and in the case of blood pressure, we have 35 null values; in the case of skin thickness, oh my god, 227; and for insulin we have 374. You can see how many NaN values, like, you know, how many missing values we had and which were just hiding behind zeros. For BMI we have 11, right? As inferred, columns glucose, blood pressure, skin thickness, insulin, and BMI have missing data. Glucose, blood pressure, and BMI have less numbers of missing data, while in the case of skin thickness and insulin have high or very high amounts of missing data. Now removing this missing data from the data set will result in information loss, and it is not advisable to remove those records since we have only 768 records. Hence, we will impute these missing values with the median of their respective columns. Since the median is the least affected by outliers, we will replace all the NaN values with the median, right? This is how we are going to handle missing values, right? Please do let me know if you have any doubt about missing values, about outliers, about whatever we are talking about, right? So I will happily help you, and I'll help you to solve your doubts as well. Now if you just write `data.median`, this median value or the, like, we know how median works, right? So just by calling the function `median`, it will give you the median for each and every column you have. So for pregnancy, median is three; glucose is 120; for blood pressure is 72; for skin thickness we have 29; for insulin you have 125; for BMI you have 32; and for diabetes pedigree function you have 0.8; and for age you have 29, right? So this is how your median function is going to work in your whole data set. Now what we need to do? We need to import, or we need to not import, we need to fill our NaN values with `data.median`. How can you do that? Imputing missing values with their respective columns' median. And to fill NaN values, we need to use a function called `.fillna`. So we have written `data.fillna`. What you need to fill? `data.median`. To your values will be `data.median`, and `inplace=True`, all right? Let me just print `data.head`. So can you see your skin thickness is replaced; your insulin that we got that NaN value—this 125.0, 125.0; this first three columns are replaced with the median. So you will not have any data loss. It is actually recommended not to delete the data or not to cut the data. If you have a huge data set, maybe you can go for it, while you have just 768 data. This is not recommended to remove the data with NaN values except like this; you can replace NaN values with the median, right? Now we are going to check if the missing values have been imputed or not. How to do that? We just need to again use `data.isnull().sum`. Why do you use `.isnull()`? We use `.isnull()` and `.sum` to check how many null values we have in our data set. Can you see now we just have zero null values, right? In the next part, we are going to talk about how can we handle outlier detection and how can we treat them, right? I hope you have all understood about missing values. If you do not know, in the normal way that you have missing values, then now you have to go deep, and you have to investigate a little more to understand: do you really have missing values or you don't? If you have missing values, then go and remove them if you have a big data set. If you do not have a big data set, then replace them with the median. And after that, you are ready to use your data set, right? So this is what missing value treatment works for, and how can you find it out. Then in the next part we will talk about how we will find out outlier detection and how can we treat them.

Now, in this part we are going to talk about outlier detection. I know you all have an idea about outlier detection, right? Now we will talk about outliers. Outliers are extreme values that deviate from the other observations in the data. They may indicate a variability in the measurement or experimental error or novelty. Box plots are a great way of detecting outliers. We already talked about box plots, right? So yes, it will help you to find out the outliers. Once the outliers have been detected, they can be imputed with the 5th or 95th percentile. We will talk about that—how to do that. So how can we do that using box plots? You have to write `plt.figure`. So make your figure size. Now you will go for `plt.subplot`. In one plot, there will be a subplot you want, right? So you have written 4, 4, 1. Now we are going to use Seaborn, and I have given Seaborn as sns. `sns.boxplot` and `data['pregnancy']`. So for the pregnancy column, you just want to plot the box plot. Now in the case of glucose, you will again write `plt.subplot(4, 4, 2)`. Can you see the numbers are changing? That is where you want to place this particular box plot. Now I have written `sns.boxplot`, then `data['glucose']`. Then move further; we have `plt.subplot(4, 4, 3)`, and `sns.boxplot(data['blood pressure'])`, right? So for again we will going to write `plt.subplot(4, 4, 4)`, your data, where you want to put your box plot, and I have written `data['skin thickness']`. The same way we have done for `data['insulin']`, we have done for BMI, diabetes pedigree function, and age, right? So it's going to give us the box plot. Now if you don't know box plots, then I should suggest please check the previous part of the video where we have talked about box plots. Now in the case of pregnancy, can you see there are three dots? So basically it's showing there is an outlier. In the case of glucose, you do not have any outliers. In the case of blood pressure, you again have outliers. In the case of skin thickness, obviously you have outliers. As it has outliers, diabetes pedigree function has outliers; insulin has lots of outliers; and BMI also has outliers. Now if I go for the observation, apart from glucose, all the other attributes show the presence of outliers. These lower-level and upper-level outliers will be replaced by the 5th and 95th percentiles respectively, right? Now we have talked about IQR. What is the interquartile range, guys? Let me know in the comment section again, for your, what's up, though, like, you know, for your heads up, that the interquartile range, that IQR, is a measure of variability based on dividing a data set into quartiles. What else? It divides the rank-ordered data set into four equal parts that we have talked about, which the parts are known as quartiles. So the values that divide each part are called the first, second, third quartiles, and they are denoted by Q1, Q2, Q3 respectively. Now what we are going to do? We are going to, you know, that we have talked about interquartile range. Now what we are going to do? We are going to write `numpy.clip` function. So `numpy.clip` function will help you to keep our data as we have outliers, right? So what is, like, how to use the `clip` function? So if this `clip` function, you need to give the limit, the values in an array syntax—`numpy.clip`. And the parameters you have: the array containing elements to clip; you want to clip; then the minimum value; if None, clipping is not performed on the lower interval; h not more than one of `a_min` and `a_max` may be None; and `a_max` is nothing but your maximum value, right? Now if I just need to clip, what I need to write? I have written `data['pregnancy']`;

Want to work on the pregnancy column? Then I have written data of pregnancies.clip. I want to clip my data now. What I need to do for the lower limit? I am going to give data of pregnancy.quad I 0.05. Then we have written that upper will be theta of pregnancy quartile 0.95. As we have talked about here, right? That uh the values divide each and okay, that not so yeah. This lower and upper level will be replaced by 5th and 95th percentile respectively. So your fifth quartile like fifth percentile and 95th percentile while it's comes to upper and then while it comes to lower, right? This is lower, right? So this is how we can do that. Now let me do that. Okay. Now I have written data of blood pressure and data of blood pressure.keep. Again, you are going to keep the blood pressure column data for skin thickness. Yes, you are going to again clip skin thickness. For insulin, you will keep insulin. You will clean BMI, diabetes medical function and age, right? Wherever we have seen this particular uh, you know, outlier, then what we need to do? We again need to plot our block box plot to understand do we have again any outlier in our data set? If we have, how can we, you know, replace them again with the clip function to keep our data to remove outliers, right? So we are going to write plt.subplot the same way one then sin is a blocks plot and then for pregnancy, glucose, blood pressure, skin thickness, insulin, BMI, diabetes prodigy function and age, right? Can you see for the case of pregnancy again you do not have any outlier. Glucose, blood pressure, skin thickness, okay, age type respiratory function and BMI are fine. Uh, for skin thickness also you have a little outlier, right? So skin thickness and insulin is not worked out, right? As we can see there are still outliers in the column called thickness and insulin. Let's try manipulating them with the percentile value. Now we will try to do heat and trial. Now we have written that again we are going to click and quartile will be 0.07 to 0.93. Let's see how it's going to work. And for insulin also we have written 0.21 the lower one and the upper one we have written 0.80. Now if I just plot both of them, okay, insulin has more night or insulin didn't work, but skin thickness has worked on. Now let's manipulate insulin little more. I have right quantile quartile equals to 25 for the lower and the upper will be 0.75 percent. Now if I just plot it, can you see your insulin is also worked and there is no outlier for your insulin. So this is how we can actually help to remove outliers from the data. The outliers of data scheme thickness were treated by minor changes in the percentile, but the outer layers of insulin require a major changes in the percentile. This might result in too much data manipulation which might uh like, you know, um again it's not good for your model. Attribute insulin might have to be removed from the dataset. It will give you good value because as you have to manipulate your insulin for so many times. So this is what missing value and outlier analysis stands for.

Now after this part we will talk about data visualization. How can you visualize our data using matplotlib and cbon? Now we are going to see how can we do the data visualization. But what are the things we have done so far? We have talked about how can you understand data using pandas. Then we jump into how can I see that my data is already uh, you know, uploaded. Then we have used pd to automate CSV. When we go for understanding missing value and basic terms to understand your data like describe, info and all. Then we have talked about missing value. How can I find it out? Missing value maybe in the simple case you will not get any missing value. You have to deep go and edit drive. I need to see do you really have any missing value or not? If you don't have any missing value, then it's fine. If you have missing value, you have to make sure that you are like those missing valve are treated. Then we go for outlier analysis. So it's basically trying to understand we have any error in our data set or any anomaly. If we have outlier, we will try to replace them using under quartile range, right? Now we will do the data visualization. Now we are going to use fns.count plot. So we just try to understand what is the data set we have where how many people are diabetic, how many people are non-diabetic, all right? In that case we are going to use outcome column as I said outcome is zero means the person is not diabetic where your outcome is one then it's showing that your person like the person is actually diabetic. So you have written ds in s.count plot. This sms.count plot is basically give you the total count of your data for zero and one, the different class we have. Now in the case of uh like count you can see zero is high, one is low. So zero means the more person you have is non-diabetic and 50 percent of the person you have diabetic, right? So this is why deutschen is dot count plus stands for. We have already talked about count plot, right? Why don't you guys see what are the things we can do using this particular data set where we can use count plot, right? Now we will see the observation again from the above plot. We can infer that the majority of the data consists of non-diabetic patient. Let's try to understand the percentage distribution of diabetic versus non-viability in the data set. Now what you have written, you have written total equals to you have written total equals to float and then length of your data. A little bit of programming we have done here. Now you have written ax equals to sms.com plot and your x equals to outcome and data equals to a data set you have. For p in a x.patches height equals to p.get height x.text p.get_x + p.get_width divided by two 8 + 3, right? Why do we have done those things to get the percentage of the data? You can see the 65 percent of the data contains records belonging to non-diabetic patient and the 35 percent of the data is saying you have a diabetic patient. The data set has a class imbalance and might have to be treated in future during the model building stages. While you try to build a model maybe logistic regression uh we have like we will talk about linear regression analysis. We will talk about machine learning as well, but yes, if you want to model you have to make sure you are going to use the class imbalance technique. Not not now you need to don't know about it, but yeah, these are the things it's as like, you know, suggested.

Now we will see if like we will do a correlation. What is correlation? So let's now plot a correlation plot that is like this plot will help us understand if there is a multiple linearity in the data set or not. So what I have written, I have written f, comma ax. Again we are going to understand this figure size that plt.fixed size fixed size equals to 20 comma t. Then for correlation you have to just write data.correlation what type pearson. Now you have written sms.heatmap and then your correlation, right? So you will get what you will get this correlation one. Can you see correlation with pregnancy and pregnancy it will be always one. That's why diagonally all this the index you have that is one, right? This is something like that you know or read part with the more hype like morally highly correlated, right? Now from the above correlation it can be infer that there is no high multi quality in the data set. The correlation plot shows the relation between the parameters glucose, Hbmi and pregnancies are the most correlated parameters with the outcome. Insulin and diabetes predictable function have little correlation with outcome. Blood pressure and skin thickness have tiny correlation with outcome. There is a little correlation between age and pregnancies, insulin and skin thickness, BMI and skin thickness, insulin and glucose, right up. Now we are going to work on pair plot, right? And this is what we are going to do and that's all for our data visualization. What is fair plot? Pear pod is a pair wise relationship in the data set. If you want to see pairwise relationship, the pair plot function creates a grid of access such that each variable in the data will by shared in the y-axis across the single row in the x-axis across a single column. Now to get the pair plot I have written sms.pair plot data u equals to outcome and uh like this what how how we are going to do a pair plot. But like you know working on pair plot it takes time why because player pod is making pair for each and every columns you have, right? Let's see how it's going to work. We will see about pear plot. Let's see how it's going to work on pear plot. So I hope you guys have understood all the technique we have done so far. If you have any doubt do not forget to let me know in the comment. I will happily help you to solve your doubts. And if you like the video do not forget to like it and subscribe great learning. So can you see this is what the pair plot we get. Now what I want to understand from they have a player fort. What are the things we can conclude? We can infer that the most of the particular variables are weak predictors of outcome. The kernel density plots the diagonal suggests that the distribution for diabetic and non-diabetic are very similar and are overlapping each other significantly. Hence they won't be able to differentiate between a diabetic patient and a non-diabetic patient. The scatter plot also suggests very poorly correlated data with not hidden patterns or relationship. Hence model on this data might not be able to find out any hidden patterns or might identify nonsense pattern patterns that do not make sense. The plot shows that there is some relationship between parameters. Outcome is added as hue. That variable. What is hue? The variable in data to map the plot aspects to different colors. We say that blue and orange dots are overlapped. Pregnancies and age have some kind of a linear line and blood pressure and age have the linear little relation. Most of the aged people have blood pressure. Insulin and glucose have some relation. This is what we can actually conclude from the player plot we have done, right? I hope you understand how to do a animally and outlier detection. How can you make sure that you are making sure that whatever the expertise data analysis you are doing in your data it's giving you enough insights to work with your data, right? This is what we are have done about this particular diabetes data set.

Now in this part we are going to do one more exploratory data analysis on different data dataset. We have done its fruity data analysis outlier detection in uh deputies dataset. Now we are going to use auto mobile dataset where you will get the dataset. You will get the dataset on www.k. Just click on the link you will redirect to this particular data set. It's a free data set and you will get this data set on kegel. Now if you want to download this particular data set just click on download and it will automatically start downloading, right? So this is pretty much about the data set where you will get the data. Now let's talk about the data set description, right? So in 1985 model input car and truck specifications and 1985 words automate your book so personal auto manuals insurance service office 160 water street new york ny one double zero three eight insurance coalition report insurance institute of highway safety watergate 600 washington dc20037. All right. So this data set consists of three types of entities, right? What are those entities? So basically entities is nothing but the columns we have and this because the specification of an auto in terms of barrier characteristics b it's assigned insurance restoration. It's normalized losses in use as compared to other curves. The second rating correspondence to the degree to which the auto is more risky than its sprites indicates. So cars are initially assigned a risk factor symbol associated with its price. Then if it's more risky or less this symbol is adjusted by moving it up or down the scale. Uh then uh the okay actuarians call the process symboling. A value of plus three indicates that the auto is risky and minus three that is probably pretty safe. The third factor is the relative average lots payment per insured um vehicle. Yeah, this value is normalized for all autos within a particular size classification two-door small station wagons sports or specific uh specialty um etc and represents the average uh like you know represent the average loss per car per year, right? So I know this is little complicated while we are talking about the dataset description. So let's talk about more in depth in the dataset. So if you do not know how to you know upload the data you have to just click on this particular button and from there there is option for upload the state station storage. We have talked about before has been just click over here it will redirect to you to your local system, right? And this is how you can upload your data in the session storage. My data set is auto mobile_data.csv. Now I have imported few packages. I have imported pandas, numpy, c pawn and matplotlib. So a quick are like you know a question for all of you. Why do we use seaborn and matplotlib? Please let me know in the comment section. And if you like the video do not forget to like it and if you uh like want to get updated with the new content then do not forget to subscribe great learning, right? So how can you load the data set? You have to use a function called .read_csv. We have talked about that before as well, but as a recap just to remind you. So we use .read_csv function to read csv file. So I have written auto mobile_data.csv, right? Now okay I forgot to install this importing this libraries. Okay. Now if I just write it yes it gets executed. Now if you want to see the you know first five rules of the data how can we do that? We just need to write data.head, right? Let me do that. Now if you want to see the columns in your data set how can you do that? You have to write data.columns, right? I have written data.columns. Now if you want to see the shape of your data set how can you do that? You have to write data.shape. So it will give you the total shape of your data. How many rows you have, how many columns you have. Let me do that, right? Now uh I want to see the first 10 rows of my data. What I need to do? I have to write data.head and under the parenthesis you have to write 10. Otherwise in default it will give you 5 rows. So here it gives me 10 rows, right? Now if I want to see the last of the data, I mean the from the last if I want to see the rows I have to write data.tail. What it will return? It will again return the five rows from the last from your data set, right? Total data how much you have? Two zero five data. So as it comes from zero total two zero four three two one zero you got the last five rows of your data. Now if you want to check do you have any null value or not? What we need to do? We have to use a function called .isnull and why we are using any please let me know in the comment section because I have already explained that, right? I have written data.isnull.any. So basically it's going to return yeah. So if you have any vernal value in your data set it will show here. So for all the cases it's false. So you do not have any null value. Now why do we use data.described? It's basically describe few things about your dataset. So total count of your data, what is the mean value for each and every column you have, then what is the standard deviation, then minimum value you have in the data set there. If you just want to you know uh like part your data that IQR range, right? First first second part third part fourth part. So first quartile, second quartile and so on, right? So 25 percent, 50 percent, 75 percent and then you have max value in your data, right? So this is what a data.describe function describe us about the data. Okay. Now you can see from the above snapshot of the data we can see the sum columns have missing values. What are the missing values? So no I cannot see here. Okay, right. So yeah, can you see here this is the um like maybe you did not get in each null value but you have the question marks, right? Can you see? Okay, if you can see please do let me know in the comment section. So yeah, can you see we have this we have to remove this. Okay. So for is now maybe it's not taking as null value but we do have few values that's unappropriate while we are going to do this data analysis. Now we have to what we have to remove them as a missing value detection and treatment, right? So let's start by checking the info, right? The count of the records in each columns of the data. If the count of the records is less than the total number of records so that we can conclude there are blank records. So no we do not have any bank records, right? But the records are representing this these are the actually blank records, right? Now since some of the predictor columns are continuous in nature there might be a chance that zeros in the columns indicates the missing value that we have done earlier as well. Now if I just try to write data.describe.t then you can see here the same thing, all right? So here we are saying none of the continuous seem to have a missing value represented by 0. We can also see there are total of 11 continuous variables but the described table has only 10 continuous variable. This is because of column normalized losses has missing values, right? So here how can we do that? So here we can see we are going to use numpy. So we have written from numpy import nan and data of normalized losses equals to data of normalized losses.replace that question mark, comma I want to replace the quotient mark with np.nan value. If I just do that and then if I just do the data.head you can see we do have nan values. These are the way the different way you can understand you do have missing value and you have to handle it with care, all right? Now you have to write data. if you want to see again. Now I want to see something that data.isnull.sum. Now I will get the actual sum of the null values. So for normalized losses we have 41 null values, right? So normalize has 41 missing data points and we will replace this missing values with the median of normalized losses because median is least affected by outliers, right? So if I just write data.fillna data.median, comma in place equals to true, right? So I want to fill my nan values with the data of median. Now if you want to see what is median differently let me show you. I have shown in in the previous project as well. Let me just show that. Okay, data.median it will give all the median value for each and every column you have. So for normalized losses 115 it will automatically replace those with the one one five. Let me show you. Yes, can you see all the nan values are replaced with median. This is how you can actually replace your non nan values and how can you like you know treat this missing values. Now if I try to see theta.isnull.sum then it will give me zero that we do not have actually any null values. Now if I I just don't want to go for data.info, right? So now what we need to do? So you can see over here that notice that the data type of normalized losses is object. We have to change its data type to numeric float. Yes, because we don't want to have a data type of object. We want to make it a numeric value. How can you do that? You have to write data of normalized losses equals to pd.to_numeric. If you want to change

Something to numeric, and then I have given data down normalized losses, comma down, class equals to float. Right now, if I just write data.info(), it will show that normalized losses is floor 34 too, right? It's a non-numeric value; we have done it all right. Uh, okay, what is this? Okay, right.

So now, if I just do the info(), it's done, right? Now that we have treated the missing value, let's move to the outlier detection and treatment. Now we will see how can we do the outlier detection. So outlier detection is basically: we can find it out by using box plots. A great way to detecting outliers. Once the outliers have been detected, they can be imputed with the 5th and 95th percentiles. Right, so I have written theta.columns, right, and then we have just done the outline detection using box plots, right? We have written plt.figure(figsize=(20, 15)). Then we have written plt.subplot(3, 4, 1). So what subplot you want, like you know, this is the like portion you are giving while you have a plot, and then you have the subplot, right? And your figure size will be 20, 15. Then you have given the normalized losses, symboling, what will base length, wheat, all the columns you have. You have given to see what are the columns has the uh outlier. You can see this at the outlier for highway mpg. If you have completion ratio, you have engine size, we the yeah, otg mpg, you have the outlier. Okay, few of the columns have the outlier, right? So you can see how the outliers, this dots, its are nothing but the outliers you have. Now, like from the above box plot, we can infer that out of 11 continuous variables, eight of them have outliers. These outliers will be imputed with the, right, uh, imputed with the fifth and the ninety-fifth percent. And how we did? We have used a clip function. So again, for your like, as a sweet question to all of you guys, what is clip function? Why do we need to use clip? Please do let me know in the comment section. For that, we can understand you are actually listening to us. If you like the video, do not forget to like it, and do not forget to subscribe to Great Learning.

Now we have written data of normalized data equals to data.normalized_losses. Uh, okay, clip(lower=data.quantile(0.05), upper=data.quantile(0.95)). Let me just do the for all the columns we have, and then we are going to again plot the box plot. Can you see all the columns have no outliers? So you have successfully uh handled the outlier we have in the dataset. Now you can see, you can go forward with the data. Now we will do few data visualizations. What are those data visualizations? We try to understand the pair plot. How can we do that? We have written sns.pairplot(data), and then it will give you the pair plot. Pair plot is nothing but giving the relation between different columns you have, right? So let it get uh printed. Let's see. Let's wait for a few minutes, and in the meanwhile, if you have any doubt, please do let me know in the comment section. For that, I can help you out to solve your doubts as well, right? Uh, let's see. Okay, it's coming. Seaborn and pair plot. We are using from Seaborn, so that's why it's sns, and sns.pairplot as the dataset is little bit uh, you know, big, that's why uh doing the pair plot, it's taking little much of time. Okay, so it takes a little bit of time, but you can see this is what it's showing that you get the pair plot, right? Okay. So some of the kernel density estimate plots show more than one peak, indicating the presence of clusters in the dataset. So basically, maybe you have clusters in the dataset, right? Now, if I just want to go for decorating data, like understanding the correlation between different columns we have, we use data.corr(method='pearson'), and then we have used the sns.heatmap. So this heat map is nothing but your correlation between different uh columns you have. So with symboling, symboling, it's one, you can see, and this is the like, you know, it's more of a rate. This is the more of like your columns are correlated, right? So the correlation plot shows presence of multicollinearity in the data. So you can see there are multicollinearity we have in the dataset, right? And if you just want to go for a few uh understanding few visualizations, how to do that? So I just write data.symboling.hist(). I want to understand the histogram, then I have given the title as Insurance Risk Rating of Basel Number of Vehicles Risk Rating. This is the risk rating according to the vehicles, right, for the from the above histogram. We can infer that a major part of distribution lies between the range of 0.5 to 1.5. We can also infer that a large number of cars in this dataset are safe. Obviously, cars are mostly safe.

Now, if I just want to understand the frequency chart, uh, then we'll type that data.fuel_type.value_counts().plot(kind='bar'). Right, and we have written plt.title('Frequency Chart and Fuel Type Number of Vehicles'), and fuel type, right? Let's see number of vehicles, and we'll type yes. You can see this is the number of vehicles, and this is the like we have gas is more, diesel is less, right? So the above bar chart, we can infer that the majority of the cars recorded in these datasets run on gas, right? Now we will do a few things about the dataset, the data pre-processing. This dataset has 15 categorical variables, and most of them have more than two categories. We cannot run a regression model on the text data. So we actually are going to talk about regression models. You have to keep tuned to know about regression models, and in the meantime, let's see how can you do the data processing. So in order to deal with this challenge, let's learn about label encoding. Label encoding is the process of converting categorical data into numerical data. Let's see how this is done in the example. We will be working with the variable body style, which has five categories, namely convertible, hatchback, sedan, wagon, and hardtop. Now what we have done? We have just done data.body_style.head(20). So I just want to see the 20 data from my this particular column. Now we are going to do a label encoding. How to do a label encoding? So basically, for each and every like vehicle, you will get an idea, like you will get a numeric value. Suppose you have the person is diabetic or not, you have zero and one. So you have five, one, five for different uh parts, like five different types, so you will have zero, two, zero, one, two, three, four. This is what stands for label encoding. So I have written, `from sklearn import preprocessing`, and `from scikit-learn.preprocessing import LabelEncoder`. Then I have just called the LabelEncoder, and just I have given my `label_encoder.fit_transform(data.body_style)`. Then, if you just write .head(20), you can get that what are the thing you have for a particular. Suppose you have for convertible, it's converted to zero, then hatchback converted to one. So for all the hatchbacks you will get one, for all the convertibles you will get zero, for all the sedans you will get two. This is how we did the label encoding, right? So after running the label encoding code, we can see the variable body style has numeric values ranging from zero to four. The problem with the label encoding is that it introduces an order between the categories, that zero greater than one, one greater than two, two greater than three, and four. This might confuse the model into thinking that convertible is greater than hatchback. So to deal with this problem, let's understand the concepts of one-hot encoding. In one-hot encoding, categorical columns have been label encoded, split into multiple columns, and the values are replaced with zeros and ones. One marks the presence of a value, and zero its absence, right? Let's look at the example. So I have done this particular label uh encoding, and then now I'm going to go for one-hot encoding. So again, to import one-hot encoding, I am going to import it from scikit-learn, and then I have called OneHotEncoder(handle_unknown='ignore'), and then I have given my dataset to transform. Now, if I just want to show you what's happened, can you see? You get one, zero, zero, zero. So you have total four classes, right? Total five plus zero, one, two, three, four. So if your class comes under zero, it will be one; if your class comes under two, it will give one, and other than everything will be zero. So there will be no like zero greater than one, one greater than two. Then these concepts will not be there; that's why we went for one-hot encoding, right? So this was pretty much about the data pre-processing. We have done the data visualization, how can you handle outliers, how can you handle like um different multi-collinearity, and also like how can you find out multiple linearity in the dataset, and how to do the outlier analysis as well using box plots, right? So we have talked about everything in both of the datasets. I hope you will like it.

I would like to start off today's session. So we'll be starting with this library called as pandas, which stands for panel data. And if we want to do any sort of data manipulation or data wrangling on top of a table, then this needs to be our go-to library. So as it is stated over here, so pandas would provide a single-dimensional and multi-dimensional data structures. The single-dimensional data structure is known as the series object, and the multi-dimensional data structure is known as the DataFrame. So we'll start off the single-dimensional labeled array, which is the series object. So if we have to work with the pandas library, we'll have to start off by importing the pandas library. So we'll type `import pandas as pd`. After we import it, we'll have to create the series object. So we'll have `pd.Series()`, and inside this we will be passing in a list. So here, as you see, I am passing in a list of values starting from 1, going on till 5, and that I am storing it in this new object called as s1, and I print it out. So if you see, I have all of these numbers, and we have the indices for these numbers over here. And in Python, you have to remember that the index value starts from zero. So element one is present at index zero, element two is present at index one, element three is present at index two, and so on. And after I create the series object, when I look at the type of this, you will see that, so `type(s1)`, you will see that this is `pandas.core.series.Series`, which is basically a series object. So I'll actually quickly head on to Jupyter Notebook. Let me open up Jupyter Notebook over here. So this is how Jupyter looks like. I'll just open up a new Python file over here, and I'll start off by importing pandas. I'll have `import pandas as pd`. Then, going ahead, I'll just wait for a couple of seconds for this library to be loaded. So when you're loading Jupyter Notebook for the first time, it might just take a bit more time. And though meanwhile, I'll just write the next set of commands of `pd.Series()`, and inside this I'm passing in a list of values. So I'll have one, two, three, four, and five, and this I will go ahead and store in an object called as s1. Now, if I print out s1, you will see that this is the result which I get. Now, after this, if I would want to check the type of this, so inside the `type()` method, I will be passing in the object s1, and you would see that this is what I get. So we have created the series object. Now, after this, so as you see, we have the indices over here, and if I would want to change the index values, so by default the index values which you get are numeric, and the initial value starts from zero. But if I don't want it like that, if I want to give some other index values, then in that case, I can just add this attribute called as `index`, and I will pass in a list of new indices which will be replacing these old set of indices. So here what I'm doing is I'm giving these new set of indices which are a, b, c, d, and e. So I'll just show you guys how it's done. I'll copy the same command, I'll paste it over here, and what I'll be doing right now is I'll add this attribute called as `index`, and inside this I'll be passing in new set of indices. So I'll have a, b, c, d, and e. So one, two, three, four, five, and I have five labels over here. Then I'll print out s1. Let me print this out, and as you guys see, this is what I get. So I have the series object where I have all of these values, and these are the indices corresponding to these values. So earlier we had created a series object from a list. Now let's see how to create a series object from a dictionary. A dictionary, you can consider it to be a data structure which comprises of key-value pairs. So as you see over here, inside this we are passing in this dictionary where we have the set of key-value pairs. So here we have a 10, b 20, and c 30. So a, b, and c are the keys, 10, 20, and 30 are the values. Let me head back over here. So `pd.Series()`. I'll have this. Let me give the first key-value pair. Then I'll have b 20. After this, I'll have c 30, and this I'll just go ahead and store in a new object called as s2. Now, if I print out s2 over here, you will see that the keys become the index, and the values over here become the values of the series object as well. So this is what happens if you try to create a series object from a dictionary. Now, once this is done, we will again try to add or change the index values. So as you see here, we have a, b, and c as the initial index values. Now what I'm trying to do over here is I am trying to add one more index value, but the problem is initially we have only three keys and three values above, but I'm adding this parameter called as `index`, and over here I'm providing four keys. So, and also if you notice, I am changing the order of the indices. So what happens over here is initially the sequence will be a, b, and c, but since I'm changing the order, I'll have first b 20, c 30. I don't have a key called as d over here when I'm creating the series object, so that is why I will have NaN, which basically means a null value over here. Then I'll finally have a, and the value for a will be 10. And when I print it out, this is the result which I get. So let me copy this again. I'll paste it over here, and this time again I'll be using the `index` attribute, and over here let me just pass in a couple of index values. So I'll have b, then I'll have d, then I'll have a, then I'll have c, and let me print out s2 over here, and as you see, I have changed the sequence of the elements which are present. So now that we have created a series object, if I'd want to extract individual elements which are present in it, so let's say if I have this normal series object where I have indices like, I'll print out s1 again, and let's say if I'd want to extract the second element, so the second element is presented index number one. So what I'll do here is I'll write s1, and inside parentheses I'll give one, and as you see, I'm able to extract the element number two from this. If I'd want to extract the last element, so for the last element, the index value would have to be given minus one. So seems like we have an error actually over here. So maybe instead of minus one, what we'll do is we'll just give four over here because that will also work. So as you see, when I give the index value of four, I'm able to extract the last element from this. So this is all about your single-dimensional data structure which is a series object. So a DataFrame gives you rows and columns, so you have data in both the dimensions. Now a DataFrame, if you have worked with SQL or Excel, then you basically, you can consider this to be a table format. And if you want to represent or work with table formats in Python, the analog format will be a DataFrame. So here we are creating a DataFrame from our dictionary. So first again we'd have to import the pandas library, and after we import the pandas library, here I am creating this DataFrame by using `pd.DataFrame()`, and here you have to keep in mind that the D and F are capital O. So let's see if you give either the D or F as small, or both of them are small, then you will get an incorrect result, or you'll basically get an error. So `pd.DataFrame()`, and inside this I'm passing these key-value pairs. The first key is name, and I'm giving the list of names for that. Then the next key is marks, and I'm passing in a list of marks over here. So now if you look at this dictionary, you will see that the keys become the column names, and the values become the records over here. So name is the first column, and Bob, Sam, and Annie become the records. Then we have marks, which becomes the second column, and these list of values become the records as well. So now if I go back, what I'd have to do is since I actually have pandas imported, I don't have to import it again. So I'll just directly write down `pd.DataFrame()`, and inside this I'll just pass in a dictionary. So first I'll give the name of the column which is name, and then I'll pass in a list of values. So actually name, um maybe I'll give the first name as Raj, the second name as Howard, the third name as Sheldon, and I'll have the second key over here. So the second key would become marks, and then I'd have to pass in a list of values for this. So let's say Raj has scored 70 marks, then we have Howard who has scored 80 marks, then we have Sheldon who has got 90 marks. Now I'll just click on run, and you'd see that this is what I get. Let me actually store this in a new object, and I'll call that object to be df. I'll print out df over here, and this is the result which I get. So this is how a DataFrame works. Let me actually show you the type of this. So I'll have `type(df)`, and I will hit on run, and you would see that this tells you it has `pandas.core.frame.DataFrame`, which essentially means that this is a DataFrame object. So this is how you can create a DataFrame. Now let's see some methods on top of this DataFrame. So for this, now what we've done is we've actually created a DataFrame, but mostly in Python, if you would have to work with some inbuilt CSV files or inbuilt tables, you can directly read them using this pandas library. So for this, um now whenever we talk about a CSV file, what is a CSV file? CSV stands for comma-separated value. So all of the Excel files which you have, they're mostly stored uh with the extension .csv, and we'll be um, you know, we'll be loading them up with the help of this pandas library. So I have this CSV file in my Jupyter Notebook called as iris. So I'll start off by loading that particular file. So I'll write down `pd.read_csv()`, and inside this I'll have iris.csv, and I will store this in an object called as iris. So now once I store this, I look at the head of this. So `head()` of something which will give me the first five records of the DataFrame, and I'll hit on run. So now this method will help me to load a CSV file. I've loaded it, and I've stored it in the iris object. Then the `head()` method will help me to look at the first five records of this DataFrame.

The first five records and this is what I'm doing. So this iris data frame comprised of these columns: so we've got sepal length, sepal width, petal length, petal wetland, species. And this basically denotes the different species of the um iris flower. So you've got the setosa, virginica, and voici color species for this iris flower.

Now, similar to the head, so initially the head method gives you five records by default. Now, instead of these five records, if I would want to look at the top 10 records, I'll just given the value 10 inside the method, and you will see that I have extracted the top 10 records from this.

Now, a method which is analogous to the head method is the tail method. And here what I'll be doing is I'll write down iris.tail, and inside this I will just pass in um, or I can just directly hit run, and you would see that I've got the last five records. If you look at the index values over here, the index value starts from 145 and it ends at 149. So we've got the last five records. So we've got the last five records from the iris data frame. And similarly, if I'd want to get the last 10 records, I'll just pass in the value 10 over here, and you'd see that I've extracted the last 10 records.

Then, if I want some basic information about the different numeric columns which are present, so what I'll do over here is I will write down iris.describe. And when I hit on run, you'll see that I've got all of these information about these different numerical columns. So let's say if I want to find out the mean sepal length, you can directly find it over here. So the mean sample length is 5.8. Similarly, the minimum sepal width is 2, the maximum petal length is 6.9, the standard deviation for the petal width column is 0.76. So this is some interesting uh, you know, there are some interesting pointers which you can directly find out by using the describe method.

Then, if I want to know how many rows and how many columns are present in this particular data frame, I can use the shape method. So I'll write down iris.shape, and as you see, I get the values 150, 5, which will tell you that there are 150 records and there are five columns. So this is a brief info about the pandas data frame, pandas library.

So with help of matplotlib, you can create all of these different geometries. So you can create a line plot, scatter plot, bar plot, a histogram, and so many of the other geometries. So first what we are doing over here is to create the data. We'll be importing another library which is numpy. So numpy is something with the help of which we can we can do a lot of numerical computation. So we'll be importing that. So we'll write down import numpy as np.

Then after that, what we are doing over here is we are importing the matplot library. So I'll write down from matplotlib import pyplot as plt. So sure, pyplot is a sub module which is present in matplotlib, and I am giving this an alias of plt. Now, once I load the required libraries, I will create the data over here. So first I will have x = np.arange, and np.arange will help you to create a numpy array where we have values within a particular range, and I'm setting the range limits over here. So the range limits are 1 and 11, and you'd have to keep in mind that the initial limit is inclusive and the final limit is exclusive, which means that I'll have value starting from 1 and going on till 10; it will not include 11 because 11 is exclusive. So sure, I'm creating the data for the x-axis which I'll have over here. Then, to create the data for y-axis, I'll just multiply this with 2: 2 * x. Now 1 becomes 2, 2 becomes 4, 3 becomes 6, and so on. And then I'll just go ahead and plot this out. So here I will have plt.plot. Inside this I will pass in x and y, and I'll just write down .show. And as you guys see, oh sure, I will have a straight line for this.

So let me load all of the required libraries. So I'll write down import numpy as np; from matplotlib import pyplot as plt. Again, let me just wait for a couple of seconds. So we have loaded these two. I'd have to create the data now. So here x = np.arange, and inside this I'll pass in 1, and I'll print out x over here. Then I'll have, let's say y, and y will just be two times of x. But this is what I'll have over here. I'll also print down y. And now I have values for x and y, I can directly create the plot. So I'll have plt.show. Inside this I will pass in x and y. I'll have to show this out plt.show. Um, this is actually plt.plot, and then I'll have plt.show. Now let me run this, and this is what we get.

So now the problem is we've created this plot, but it's extremely bland; we don't have much information for this plot. So I will add a title, x-axis label, and y-axis label for this. So I'll have plt.title, and I'll add the title as line plot. I'll add an x-axis label, so I'll have plt.xlabel, and I'll set the x label to be x axis. And I'll have plt.ylabel, and I'll set this to be equal to y-axis. And as you see, I've assigned a title, the x-axis label, and also the y-axis label. So this is how we are creating this line plot.

Let's directly head on to our um case study for today's session. So we've got this cricket data set where we have all of the information about the different matches which were played till 2019 world cup. So we will be analyzing all of those data. So Narsimha is asking where is the data set from. So we've taken this data set from Kaggle. So if you want any of these open source data sets, you can take them from Kaggle, and that is where we uh take most of the data sets which we use for our sessions.

Right. So coming to this case city, I'll just quickly import required libraries. Now, once I import these, I have a certain, you know, have different data sets which I'll be working with. So I've got all of these different CSV files: world cup layers, ground averages, batsman data, paolo data, odi match totals, and so on. So now the one main data set which I'll be working with is this. So I've got this odi_match_totals.csv, and I'll be loading this by using pd.read_csv, and I'll store it in this object called as odi_total.

So now that I've loaded the required library, what I'll do over here is um, let's say I also have this data set called as world_cup_players.csv, and I'll be loading this as well. I'll start off by having a glance at all of the players which are who are present over here. So if you see uh this players.head comprise of these three columns: I've got player id and country. So player, we've got the player name, we've got the unique id for the player, and the country for which um, you know, for which this particular player hails from. Then, as you see over here, initially head only gives you the top five records. Let's say if I given 10 over 2, I'll get the list of these 10 players. So these are the 10 players who uh, you know, constantly played. So this all of these uh data sets, they have data starting from 2013 going until 2019 world cup. So throughout these years, these were the list of cricketers who have consistently played for Afghanistan team.

Now, let's say if I dub, instead of head, if I given tail, let's see what do we get over here. So we've got West Indies, and all of these other players have played for western east. Have got Fabian Allen, Sharon Gabriel, Hitmeyer, Shay Hope, Evan Lewis, and so on. So information about all of these players. And now uh from this particular list, so as you see over here, so you've got id, player, and country. If I want the list of all of the Indian players who are present in this particular data frame, I can do that with the help of this particular command. So what I'll do is first I'll given the name of the data frame which is players. Then inside parenthesis, I will given the column name. So here, as you see, the column name is country, and what I'm doing is I'm giving a condition; I'm making sure that whatever value is present in this column that is equal to India. So now let me actually insert a cell above this and show you how this works. I'll have players. Inside this I'll have country, and when I set this value to be equal to India, you'll see that I have a bunch of true and false values. Now, wherever you have a false value, it would mean that the condition is not satisfied, and wherever you have a true value, that would mean that the condition is satisfied. Now, these true and false values, what I'm doing is I'm just passing them through the data frame again. So when I have players, I cut this condition out and paste this back into the data frame, this is what I'll be getting. So as you see, I've got the list of all of the Indian players. You've got Virat Kohli, who is the captain of Indian cricket team, Rohit Sharma, who's the vice captain of Indian cricket team, and we've got all of these players who were part of the uh part of the 2019 world cup. So we've got all of this information over here. Now what I'm doing is I'm just taking this command and storing it in this new object called as indian_players. I'll just have this. And now after we do this, what I'm going to do is I'm going to work with the odi_total data frame, odi_total um object which I've created over here. So I'll have a glance at the first five records from this odi_total data frame, and I have all of these different columns over here. So this odi_total has different records representing, you know, each match which happened between different countries, and you've got all of these different things over here. So this what you see, this is a match which was played between Pakistan and India, and the score which Pakistan made was 250. So and who won the match? So if you have 1 over here, it would mean that since in the country column you have Pakistan, this particular match was won by Pakistan. Where was this match played? It was played in Kolkata. What was the date? So date was third January 2013. You've got the country idcf, basically called country ids, representing each of these countries. So for Pakistan it is 7, for India it is 6, Sri Lanka it is 8, and so on. And now I want to see all of the matches which were played by India. So again I'll be doing the same thing. This is the name of the data frame, odi_total. I'll pass in this particular column which is country. I'll set the value to be equal to India, and I'll store it back to the same object, India again. Now I will run this and I will print this out, and as you see over here, these are the different uh records which I have. And if you look at this particular column, if you look at the country column over here, you would see that all of the values are equal to India. So now if I look at this particular, if I'll just explain how it is happening over here. So this particular match was again, so this was played between India and Pakistan, and India had lost this particular match. So if you look at this and this, so this represents the same thing: India, Pakistan. You have odi id, so this is 3315, 3315. So as you see, Pakistan had made a score of 250, and India were only able to make a score of 165, and they had lost the match over sure. And then we've got this particular match, so this was between India and Pakistan again, so this was won by um India, and even though India scored only 167 runs, India was able to win the match. Then if we look at, let's see this particular match, so you've got, so let's see though, you know, the last three matches maybe before the world cup, India was playing Australia. So India had lost all of those three matches over here. So sure, even though India made 281, so Australia had made 314, India was only able to make 281, and thus the India lost the match. Socio India set up a target of 358 and also lost the match here. Australia set a target of 273, India were only able to make 237, and hence lost the match. So and these matches, if you see the last three matches were played in Ranchi, Mohali, and Delhi. So this is the sort of inference which you get from this.

Now, from this entire data frame, if I'd want to extract only those records but India has won the match, you would see that we have this result column. So I'd have to extract only those records where the value in this result column is equal to 1, and that is what I'm doing over here. So india_result = 1, and I will store it in this object called as india_when. Then I will print it out, and now if you look at this particular column, you would see that all of the values are equal to 1. So now what I'd want to do is I'd want to understand against which opponent have we won the most number of matches. You've got this opposition column over here. So I'd want to get a frequency count, since this is a categorical column. So to understand these categories or to get the frequency of these categories, we'd have this method called as value_counts, and that is what we'll be using over here. So india_when, in this particular um, so in the parenthesis we'll just pass in the name of the column which is opposition. Then I will use the value_counts method, and I will hit on run. So you would see that during this time frame between 2013 and 2019, India has won 15 matches against Sri Lanka, 13 matches against West Indies, and 12 matches against Zimbabwe. Then, you know, rest you have Australia, 12 matches; South Africa, 10 matches; England, 10 matches over here.

So now if I would want to make a bar plot for this, so um we have already imported the pyplot library, and we've seen how pyplot works. So first I'm going to set the dimensions of the image which I'll be creating. I'll have plt.figure. Inside this I will set this parameter called as figsize, and I'll set the dimensions to be equal to 8, 5. Now after this, to make a bar plot, I'm just using plt.bar, and I'll explain this command to you over here. So this takes in all of these parameters. The first parameter is india_opposition.value_counts[0:5].keys(). So till here it's the same thing. So this is the command. Now if I want to make a bar plot for only the top five entries over here, what I'll do is I'll have a parenthesis over here, and I'll just given um, I'll set this to be equal to 0:5. And now if I want only the categories over here, I will use the keys method. So when I given the keys method, I will only get the categories, and I'll not get the values. And this is what I am passing in as the first parameter to this bar or to this plt.bar. Then the second parameter what I'm passing in is the values. So I'll just remove this .keys method, and this when I pass on this, I'll get only 15, 13, 12, 10, 10, and that is what I'll be passing over here. Then I am setting the color over here, so color = green, and this I will uh, so when I just have color = g, this is what I'll have over here, plt.show. And as you see, I am able to make a simple bar plot out of this. So India has won the most number of matches against Sri Lanka, followed by West Indies, followed by Zimbabwe, Australia, and South Africa. So I'll just show you where I've got this data set. So just write down cricket world cup data set Kaggle, and click over here, and let me just open this particular link. Russia, let me see if this is the right link. Firm, this is not right link. Let me just go back again. Right, this must be the link, if I'm not wrong. Yes, right. So uh just type in cricket world cup data Kaggle, and this is the uh these are all the data sets which I'm using over here. Right now, the data set which I'm working with is this odi_match_totals.csv. So you guys can go ahead and um download this dataset from over here.

So now that we've done this, we don't want to do a similar analysis for the matches India has lost. So for that, what I'll do is I'll have india_result, and this time I would set the result to be equal to lost, and this I will store it in this object called as india_loss. And again, what I'll do is I would want to check the value counts, india_loss_opposition.value_counts, and this tells me that during this time period of 2013 to 2019, India has lost the most number of matches against Australia. So India has lost 13 matches against Australia, 8 matches against New Zealand, 8 matches again, sorry, 8 matches against England, 8 matches against New Zealand. So somehow it would seem that India was able to perform better against the lower ranked teams, and India did not really perform that well against the higher ranked teams, because if you look at this, India has a lot of wins against Sri Lanka, West Indies, and Zimbabwe during this time period, which maybe Europe do India has won the matches, but then again they were quite low ranked during this time period. And if you look at this over here, India has lost matches against Australia, England, and New Zealand, which tells you that maybe they were, you know, the quality of uh um the batting or the bowling from the Indian team was not really up to the mark, and that is why maybe they did not win the 2019 world cup. So here I'm going to do a similar plot. First what I'll do is I will um set the figure size, so plt.figure, figsize, and I'll set it to be equal to 8, 5. Then again I'm making a bar plot out of this, so india_loss_opposition.value_counts, and I would want to get uh the I would want to make this bar plot for only the top three or themes over here. So I'll have 0 to 3, and since I want only the categorical data, I'll have Australia, England, and New Zealand. Then if I would want the values, I will pass in the same command without the .keys method. Then again I'm showing the color. So now let me just change the color; instead of green, let me actually make it red, and I am printing out the plot over here. You would see that India has lost the most number of matches against Australia, then you have England and New Zealand closely falling behind.

Now, if I would want to make a pie chart for this, so pie chart will basically give me a distribution of, you know, what percentage of matches have we lost against different teams. So again I'm setting a figure size over here, so plt.figure, figsize = 7, 7. And when I am making this pie chart, I will have plt.pie. And when it comes to a pie chart, the parameters what you pass in as a sort of opposite of what you do with a bar plot. So your first will be passing in the numerical entries, then we'll be passing in the categorical values. So first we will have um, you know, india_loss_opposition.value_counts, overseas. So first we will get all of the numerical entries, then we'll have the labels. So labels are basically the categorical values. So here india_loss_opposition.value_counts.keys. Then to set the percentage values, I will have this new attribute called as autopct, and I am setting percent 0.1f%. So this 0.1 will mean that the percentage will be till the first decimal value. And I'll just go ahead and run on this, and you would see that if we look at the overall losses, all of them, 26.5 of the losses have come against Australia, 16.3 percent losses have equally come against England and New Zealand, 12.

Two percent losses against South Africa, ten point two percent against West Indies, and so on. So this is how you can create all of these beautiful, uh, visualizations.

Now, as we had analyzed this for Team India, similarly, we don't want to do an analysis for Australia as well. So now, from this ODI total data frame, I'm setting the country value to be equal to Australia, and I will store it back in Australia over here. And now I don't do a similar analysis where, up, you know, Australia team has won. So here, as you see, country is Australia, and you have the result column over here. I'd want to analyze all of those records where the value of result is equal to one. So Australia result is equal to one; I'll store it in this object called as Australia. When then I do want to see the value counts, you would see that Australia has won 14 matches against England, 13 matches against India, and 13 matches against Pakistan.

Now, I will again make a bar plot for the top three teams against which Australia has won the most number of matches. So the command is pretty much the same: plt.bar. First, we will pass in the categorical entries; next, we will pass in the numeric values, and this is what we have. So most number of matches against England, India, and Pakistan. Similar sort of analysis where Australia has lost the matches. So Australia result is equal to lost, and this I'll store in Australia_loss, and I'll look at the value counts method over show, and you would see that so if you look at this, so Australia has won the most number of matches against England, and Australia has also lost the most number of matches against England. Similarly, the second most against which Australia has won is India; again, the second most against which Australia has, um, you know, lost, as India. So third position, uh, Australia had won 13 matches against Pakistan. Over here, you would see that Australia has lost 11 matches against South Africa. Um, but if you look at this, so it would seem that Pakistan was never really able to perform well against Australia over here. So in this time period, uh, Australia had won 14 matches. So Australia had won 13 matches against Pakistan, and Pakistan had won only one match against Australia during this time period. So this is something you will, um, get to know over here. So this is an analysis for the Australian team, and now, uh, we'll do a similar sort of analysis for England.

So from the ODI total, I am setting the country value to be equal to England, and I will store that in this object called as England. So here, this is what I have. If you look at this country column, you will have England; um, you'll have, um, result, which will tell you if England has won this match or not; you will have the opposition column over here, and I am, um, setting a condition over here: the result is equal to one, and I'll store it in this object called as England_win. Then I'll look at the value count so that I know against which team has England won the most number of matches. So England has won the most number of matches against Australia, then followed by West Indies, then followed by New Zealand. And I'll do a similar sort of analysis where England has lost the matches. You would see that England has again—so if you see this, England has won the most number of matches against Australia and also lost the most number of matches against Australia. Then, uh, England has lost 11 matches against Sri Lanka and 10 matches against India. It would seem that Englanders, you know, was not really able to perform that well in the, uh, you know, uh, South Asian subcontinent. So they weren't really able to play that well against India and Sri Lanka.

So this is all of the analysis which we have done with respect to today's session. Now let's quickly summarize the video. In today's course, we discussed exploratory data analysis; we discussed the types, advantage, and also demonstration of EDA in Python; and we also understood data visualization libraries and graphs. Then we understood how to do multivariate analysis, after which we understood the different methods of outlier detection. Finally, we did our analysis on cricket World Cup data set. If you haven't subscribed to our channel yet, I want to request you to hit the subscribe button and turn on the notification bell so that you don't miss out on any new update or video releases from Great Learning. If you enjoyed this video, show us some love and like this video. Knowledge increases by sharing, so make sure you share this video with your friends and colleagues. Make sure to comment on the video for any query or suggestions, and I will respond to your comments.